I post-train multimodal models that act in real environments. Given a capability we want, I build the tasks, environments, rewards, and synthetic data that teach it, then train with RL. Right now I care most about RL environments and post-training for computer-use and real-time multimodal agents.

Now at Meta Superintelligence Labs on real-time omni models for personal agents. Previously Visual Intelligence at Apple and MIT-IBM Watson on synthetic data. Ph.D. in Computer Science at MIT; B.A. in Physics at UC Berkeley. I read, write, and hike.

Research

Omni models · 2025–26

Muse: Real-time perception and action

An omni model that perceives and responds in real time instead of turn by turn, and can be interrupted mid-action. I lead efforts in synthetic data and RL, and work on the distillation that makes it fast and natural enough to talk to.

Vibes: AI video creation and remix in the Meta AI app
Vibes · 2025

DPO and offline RL for a generated-video feed

Built the agentic long-form video generation workflow and the DPO and offline RL system behind Vibes, Meta AI's personalized video feed, and led the human-evaluation data behind it.

Announcement

Apple Intelligence features across iPhone, Mac, and iPad
Apple · 2025

Visual Lookup and StreamingQA on device

Built the datasets for on-device visual question answering on iPhone with privacy-preserving VLMs, and post-trained Apple's foundation model for proactive question answering on egocentric video. Shipped in Apple Intelligence at WWDC 2025.

WWDC 2025

Synthetic Data RL pipeline: data synthesis, difficulty adaptation, selection and RL
Synthetic Data RL · 2025

Environments and rewards from task definitions

Given a target capability, synthesize the tasks, the environment, and the reward, then train with RL. No labeled dataset required; the task definition is the input.

PaperAPI Pack

JetMoE-8B benchmark comparison against Llama2-13B
JetMoE · 2024

Open MoE at Llama2-13B quality

JetMoE-8B, an open-source mixture-of-experts model trained for a fraction of the usual cost that outperforms Llama2-13B, with a synthetic corpus for pre-training and post-training.

PaperCode

Octo-Planner decomposing a user request into web search, video search, and email actions on a phone
Octo-Planner · 2024

On-device planner-action agents

Efficient on-device task planning and decomposition, trained on a synthetic planning dataset with a multi-LoRA post-training pipeline.

Paper

News

Recent Essays