Real-time avatars for live conversation
A live, expressive avatar that talks with you in real time. It's audio-driven, streams as you talk, and is fast enough to hold a real conversation. Announced in Mark Zuckerberg's Meta Connect keynote.
Training real-time multimodal agents. Research Scientist, Meta Superintelligence Labs · 2025–
I train real-time multimodal agents that talk, listen, and act while the interaction is still going. I build the tasks, synthetic data, and rewards, then train with RL.
At
Meta Superintelligence Labs I co-lead post-training and lead synthetic data for Muse Realtime Avatar, and developed offline RL for Vibes’ personalized feed. Before that, I worked on Visual Intelligence at
Apple and foundation models at
MIT-IBM Watson. I got my Ph.D. in Computer Science at
MIT, and B.A. in Physics at
UC Berkeley. Outside work I read, write, bike, and hike.
A live, expressive avatar that talks with you in real time. It's audio-driven, streams as you talk, and is fast enough to hold a real conversation. Announced in Mark Zuckerberg's Meta Connect keynote.
A feed of AI-generated short videos in the Meta AI app, where anyone can create and remix. Personalized with preference learning and offline RL.
Ask about anything your iPhone sees, answered on device by Apple's foundation model. Shipped in Apple Intelligence at WWDC 2025.
Given a target capability, synthesize the tasks, the environment, and the reward, then train with RL. No labeled dataset required; the task definition is the input.
JetMoE-8B, an open-source mixture-of-experts model trained for a fraction of the usual cost on a curated mix of public real and synthetic data. Its chat model beats Llama2-13B-Chat on MT-Bench.
A small on-device model that breaks a request into steps for an action model to carry out, so a phone agent can plan without the cloud.
All publications → · Google Scholar →
At COLM 2026 in San Francisco, Oct 6–9. Email me to meet at the poster sessions or the Lifelong Agents workshop.
Read more →Muse Realtime Voice and Avatar announced at the Meta Connect 2026 keynote: one streaming system for conversation, speech, and a live, expressive avatar. I co-lead post-training and lead synthetic data for Muse Realtime Avatar.
Read more →Muse Spark 1.1 released, a multimodal reasoning model for agentic tasks, alongside the Meta Model API public preview.
Read more →A clip model never faces consequences. A real-time model does: any frame can change what the user does next. That makes it a policy, not a generator.
When the headline metric becomes output per joule, a technology has become a utility. Intelligence just did, and the money will go where it always goes: upstairs.
Machines calculate. Humans commit. The gap is not intelligence — it is what you are willing to wager once the evidence runs out.