MAI-Thinking-1
Microsoft AI’s reasoning model, trained from scratch for math, coding, and tool use.
My focus at Microsoft AI · Foundation-model safety
AI RESEARCH & ENGINEERING
AI safety.
RL, reasoning & agents.
I’m a Member of Technical Staff at Microsoft AI, working on safety for foundation models. Previously, I worked on reinforcement learning for reasoning at Meta Superintelligence Labs (MSL) and on conversational AI at FAIR (Facebook AI Research).
You’ll find my publications under Da Ju.
Foundation-model safety at Microsoft AI.
Reinforcement learning and reasoning for Llama 4.
Travel-planning agents and conversational learning.
SELECTED WORK
My work spans foundation-model safety, reinforcement learning, agentic planning, and conversational AI.
Microsoft AI’s reasoning model, trained from scratch for math, coding, and tool use.
My focus at Microsoft AI · Foundation-model safety
Meta’s natively multimodal model family. I worked on reinforcement learning and reasoning for Llama 4.
Model development · RL & reasoning
I worked on an agentic travel-planning system that turns natural-language requests into plans with guarantees.
Agentic research · Language → plans
An attention architecture that brings recurrent processing to sequences. We explored looped-transformer ideas in our 2021 preprint.
Oral presentation · Preprint 2021 · Attention → recurrence
I’m a joint first author on BlenderBot 3, a deployed chatbot designed to learn from conversations and engage responsibly.
Joint first author · Conversation → learning
THE RESEARCH LIBRARY
Peer-reviewed papers, preprints, and team reports. Checked against Google Scholar on September 5, 2026; duplicate versions are combined.
22 works
Try another topic, title, or year.
Years refer to the published version where available; earlier preprints are noted. Collective reports retain their team authorship. Code and dataset links point to verified releases. Some repositories are archived; their licenses and setup notes apply.
A LITTLE BACKGROUND
My research spans deployed dialogue systems such as BlenderBot 3, recurrent architectures such as Staircase Attention, and agentic travel planning with To the Globe. More recently, I worked on RL and reasoning for Llama 4; I now focus on foundation-model safety at Microsoft AI.
At FAIR, I worked with teams led by Gabriel Synnaeve, Jason Weston, and Yuandong Tian. Before that, I studied machine learning, human–computer interaction, and data science in Paris.
More on LinkedInMember of Technical Staff · AI safety
Research Scientist · LLaMA reasoning
Reinforcement learning for reasoning models.
Research Engineer · FAIR
Télécom ParisTech · Machine Learning & HCI
UPMC / Paris VI · Data Science
Information Science & Engineering
“High-resolution LED Headlamps Testing System with Web GUI,” completed at L-LAB in Lippstadt, Germany. The thesis remains the intellectual property of HELLA.
Read the thesis (PDF)A PERSONAL PROJECT · OUTSIDE THE LAB
NYC What To Do brings together events, exhibitions, dining, and interesting things around the city—with filters, a map, and a place to plan your day.
Explore the NYC guideSAY HELLO