I post-train multimodal models that act in real environments. Given a capability we want, I build the tasks, environments, rewards, and synthetic data that teach it, then train with RL. Right now I care most about RL environments and post-training for computer-use and real-time multimodal agents.
Now at Meta Superintelligence Labs on real-time omni models for personal agents. Previously Visual Intelligence at Apple and MIT-IBM Watson on synthetic data. Ph.D. in Computer Science at MIT; B.A. in Physics at UC Berkeley. I read, write, and hike.
Research
Omni models · 2025–26
Muse: Real-time perception and action
An omni model that perceives and responds in real time instead of turn by turn, and can be interrupted mid-action. I lead efforts in synthetic data and RL, and work on the distillation that makes it fast and natural enough to talk to.
Vibes · 2025
DPO and offline RL for a generated-video feed
Built the agentic long-form video generation workflow and the DPO and offline RL system behind Vibes, Meta AI's personalized video feed, and led the human-evaluation data behind it.
Built the datasets for on-device visual question answering on iPhone with privacy-preserving VLMs, and post-trained Apple's foundation model for proactive question answering on egocentric video. Shipped in Apple Intelligence at WWDC 2025.
Given a target capability, synthesize the tasks, the environment, and the reward, then train with RL. No labeled dataset required; the task definition is the input.
JetMoE-8B, an open-source mixture-of-experts model trained for a fraction of the usual cost that outperforms Llama2-13B, with a synthetic corpus for pre-training and post-training.
Hierarchical multi-agent systems import human coordination constraints into a substrate that doesn't have them. The evidence, the mechanism, and the alternative.