I train multimodal models to act in the real world. Start from the capability we want, build the tasks, environments, rewards, and synthetic data that teach it, then train with RL.
At
Meta Superintelligence Labs I lead synthetic data and RL for Muse, a real-time omni model for personal agents, and built the offline RL system behind Vibes. Before that, I worked on Visual Intelligence at
Apple and foundation models at
MIT-IBM Watson. I got my Ph.D. in Computer Science at
MIT, and B.A. in Physics at
UC Berkeley. Outside work I read, write, bike, and hike.
- Open questions
-
- What does an agent need to perceive and act continuously, rather than turn by turn? Muse ↓
- How far can a task definition alone take you: environment, reward, and data? Synthetic Data RL ↓
- How small can a model be and still plan on-device? Octo-planner ↓
- Talk to me
- I'm at COLM 2026, Oct 6–9, and want to compare notes on two things: rewards for real-time models that act while they perceive, and how far a task definition alone can carry an RL environment. If that's your problem too, email me. I reply.
Research
Muse · Meta Connect 2026
Muse Realtime Voice carries the conversation as a stream of speech tokens, and Muse Realtime Avatar turns the same stream into a live, expressive character for as long as the conversation runs. I lead synthetic data and RL, and work on the distillation that makes it fast enough to serve in real time.
Keynote
Vibes · 2025
Built the agentic long-form video workflow and the offline RL system behind Vibes, Meta AI's personalized video feed, and led its human-evaluation data.
Blog
Apple · 2025
Built the datasets for on-device visual question answering on iPhone with privacy-preserving VLMs, and post-trained Apple's foundation model for proactive question answering on egocentric video. Shipped in Apple Intelligence at WWDC 2025.
Blog
Synthetic Data RL · 2025
Given a target capability, synthesize the tasks, the environment, and the reward, then train with RL. No labeled dataset required; the task definition is the input.
PaperAPI Pack
JetMoE · 2024
JetMoE-8B, an open-source mixture-of-experts model trained for a fraction of the usual cost that outperforms Llama2-13B, with a synthetic corpus for pre-training and post-training.
PaperCode
Octo-planner · 2024
Efficient on-device task planning and decomposition, trained on a synthetic planning dataset with a multi-LoRA post-training pipeline.
Paper
All publications →
News
-
Oct 2026 · Upcoming
At COLM 2026 in San Francisco, Oct 6–9. Email me to meet at the poster sessions or the Lifelong Agents workshop.
Read more →
-
Sep 2026
Muse Realtime Voice and Avatar announced at the Meta Connect 2026 keynote: one streaming system for conversation, speech, and a live, expressive avatar. I lead synthetic data and RL for Muse.
Read more →
-
Jul 2026
Muse Spark 1.1 released, a multimodal reasoning model for agentic tasks, alongside the Meta Model API public preview.
Read more →
All news →
Essays on AI
All essays →