Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken DialogueExplained for Beginners
Chengqian Ma, Wei Tao, Haoyu Zhang +1 more
Abstract
An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first generates the spoken response and then runs a motion model over the finished audio, which requires a second full inference pass and precludes any joint optimisation between the two. We present Motion-Omni, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech. Joint training is not optional here: with the speech pathway frozen, motion remains misaligned with the audio, and co-adapting the LLM, Speech Generator and Motion Generator under both objectives is what recovers alignment while retaining spoken-dialogue ability. Supervision comes from a scalable, model-agnostic pipeline that pseudo-labels consistent-voice speech responses with a replaceable motion teacher, yielding 422,856 quality-ranked pairs (1,402 hours). We further release SwDA-500 and, to our knowledge, the first public evaluation protocol for stochastic open-ended full-body spoken dialogue, matching audio across motion systems while unifying rendering, automatic metrics, human evaluation, and latency measurement. Instantiated with a Qwen2.5-7B-Instruct backbone, Motion-Omni-Q7 matches the same-audio teacher cascade to within 2% on reference-free motion metrics while responding 5.4 x faster (RTF=0.78, faster than real time), surpasses all non-teacher cascades on beat correlation and diversity, and reaches a 2.62% word error rate, the lowest among the omni-modal systems compared.
Motion-Omni: When avatars move while they speak
The Problem
Most conversational avatars are built from two separate pipelines that never talk to each other. A spoken dialogue model—essentially a next-token predictor trained on voice—decides what to say, but it outputs plain audio with zero body movement. Separately, co-speech motion models take that finished audio and animate a 3D avatar’s hands, torso, and face. The standard way to combine them is a cascade: first generate the response, then run the audio through a motion model. But this approach has two structural costs. First, it requires two full inference passes—one for the words, one for the motion—making the system sluggish. Second, because the motion model never sees the dialogue model’s internal states, there’s no way to jointly optimize them. The motion can never influence what the model says, and the speech model can never benefit from motion-aware representations.
Motion-Omni eliminates both costs. It is an end-to-end framework in which a spoken dialogue model natively outputs both speech and full-body motion, generated directly from the hidden states that produce the words. Motion is not a post-processing step; it is a first-class output of the same forward pass that generates the transcript.
How It Works (The Technical Mechanics)
At the heart of Motion-Omni is a four-part system that processes user input and autoregressively produces a spoken response together with synchronised facial expression, hand motion, upper-body movement, and lower-body motion.
The framework has four components:
- A frozen Whisper-large-v3 speech encoder converts the incoming audio into continuous representations.
- A speech projector that downsamples these features and feeds them into the LLM.
- A Qwen2.5-7B-Instruct LLM backbone that processes the combined speech and text tokens to generate a response.
- A Speech Generator that autoregressively emits discrete speech units (at 12.5 Hz) from the LLM’s hidden states.
- A Motion Generator that produces the body motion.
The critical technical innovation is how the Motion Generator conditions on the Speech Generator. It uses a Token-as-Query Gated Fusion (TQGF) block that lets motion tokens “query” the LLM’s hidden states. Think of it like this: the Speech Generator is producing a sequence of words, and at each moment, it has a rich internal representation of what has been said so far, the prosody, and the dialogue context. The Motion Generator taps into these representations using a gated fusion mechanism—it asks, “What is being said right now?” and uses that to decide how the avatar should move. The gate acts coordinate-by-coordinate, selecting only the motion-relevant signals and ignoring speaker timbre or other irrelevant details.
Crucially, the Motion Generator does not wait for the audio to finish. It consumes interpolated speech-unit embeddings, bridging the rate mismatch between the 12.5 Hz speech units and the 30 Hz motion frames. This means motion is generated directly from the representations that produce the speech, without needing a separate audio-to-motion inference stage.
The architecture uses a dual-input conditioning: the Motion Generator’s keys and values come from the Speech Generator’s last-layer hidden states (carrying semantic intent and prosody), while its queries are the embeddings of the emitted speech tokens. A part-specific MLP head then produces per-frame logits over a 256-entry codebook for each body part—face, hand, upper body, and lower body.
Key Results & Benchmarks
Motion-Omni was trained in a four-stage curriculum using 422,856 quality-ranked speech-motion pairs (approximately 1,402 hours of data). The resulting model, Motion-Omni-Q7, instantiated with a Qwen2.5-7B-Instruct backbone, shows impressive results.
On the SwDA-500 benchmark—a set of 500 open-ended dialogue prompts—the model matches the performance of a strong cascade (teacher audio + motion model) on reference-free motion metrics while being dramatically faster. Specifically, it stays within 2% of the teacher cascade’s motion quality scores while responding 5.4 times faster (Real-Time Factor RTF = 0.78, meaning it runs faster than real time). Its word error rate on speech is 2.62%, the lowest among all omni-modal systems compared.
Perhaps most notably, Motion-Omni-Q7 surpasses all non-teacher cascades on beat correlation (measuring how well movements sync with the rhythm of speech) and diversity (ensuring the motions aren’t repetitive). It achieves the highest beat correlation score (7.59) and the highest diversity score (13.67) among systems that don’t run the motion teacher at inference time. On lip-synchronisation and facial geometry, it also ranks first or second across the board.
Why It Matters (Key Takeaways)
Motion-Omni represents a significant shift in how we build conversational AI. Here are the key takeaways:
- True end-to-end capability: For the first time, a conversational model natively produces both speech and full-body co-speech motion in a single forward pass. This removes the latency penalty of the traditional cascade approach.
- Speed matters: At RTF = 0.78, Motion-Omni-Q7 responds faster than real time. In practical terms, if a user speaks for 10 seconds, the avatar can generate its full response—including natural-sounding body language—in about 7.8 seconds. This makes real-time, interactive avatars feasible on a single GPU.
- Alignment through co-adaptation: The paper demonstrates that simply freezing the speech pathway and training the motion head doesn’t work; motion ends up misaligned. It is only through jointly optimizing the LLM, Speech Generator, and Motion Generator under both speech and motion objectives that alignment is recovered while retaining spoken-dialogue ability.
- A new evaluation standard: The release of SwDA-500 and the accompanying evaluation protocol addresses a long-standing gap. Previous benchmarks either assumed fixed input speech or only evaluated facial animation. SwDA-500 allows fair comparison of systems that generate their own spontaneous responses, unifying rendering, automatic metrics, human evaluation, and latency measurement.
Motion-Omni proves that we don’t need to sacrifice motion quality for speed or capability. By training a model to natively “think with its body” alongside its words, we get avatars that feel more alive, more synchronised, and significantly faster—without the computational overhead of running two separate models.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →