I’m a research scientist working on efficient inference for large language models and on search and retrieval. I’m currently a Senior Research Scientist at Apple Machine Learning Research in Paris. Before that, I was a Principal AI Research Scientist at Autodesk AI Lab and a Staff Research Scientist at ServiceNow Research, both in London. I received my Ph.D. from INRS (Université du Québec) in Montreal in 2021.

Research interests

My current work focuses on making language models cheaper to serve, and on search and retrieval:

  • KV-cache compression and reuse: learning context-adaptive compression across depth, precision and rank, sharing caches across layers, approximating attention over long cached contexts with small neural networks, and translating caches between models.
  • Discrete diffusion language models: learned unmasking policies, trajectory-aware training, and multimodal masked diffusion over text, images and audio.
  • Search and retrieval: amortized maximum inner product search, dense retrieval trained on graded relevance labels, and question answering over unseen reference documents.

Earlier, I worked on out-of-distribution detection, robustness to distribution shifts, metric learning for verification, and the training of generative adversarial networks.

See the publications page for papers.

Experience

  • Apple Machine Learning Research, Paris. Senior Research Scientist, Aug 2025 – present.
  • Autodesk AI Lab, London. Principal AI Research Scientist, Sep 2024 – Aug 2025.
  • ServiceNow Research, London. Staff Research Scientist, Dec 2021 – Sep 2024.
  • Borealis AI, Montreal. Research Intern, May 2021 – Oct 2021.
  • Google, Montreal. Student Researcher, Sep 2020 – Apr 2021.

Education

  • Ph.D., INRS (Université du Québec), Montreal, 2017 – 2021.
  • M.Sc. in Computer Engineering, University of Pernambuco, Recife, 2015 – 2016.
  • Bachelor in Mechanical Engineering, University of Pernambuco, Recife, 2007 – 2012.

Service, teaching and talks

  • Area Chair for ICLR (2024 – 2027), ICML (2025 – 2026) and NeurIPS (2026). Reviewer for machine learning conferences since 2020.
  • Mentor at the WiML workshop at NeurIPS 2024 (Interpretability and Explainability).
  • 3-day course on language model training and evaluation at the Bertinoro Spring School 2025, University of Bologna.
  • Invited talk on language models and the challenges of applying them to AEC domain-specific languages, 2300: Technology and Practice, Yale School of Architecture, February 2025.
  • RepLiQA: a robust benchmark for question answering (recording).
  • Closing the gap between machine learning research and practice via versatile and robust predictors, ServiceNow, July 2021.