EchoWM: Open and Enterable Omnimodal World ModelsExplained for Beginners
Songchun Zhang, Yaowei Li, Junhao Zhuang +19 more
Abstract
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.
1. The Problem
Generative AI has spectacularly expanded what video, audio, and text can look like. We now have models that can render a minute of high-definition video from a short text prompt, or compose a believable piece of music from a mood tag. But if you’ve tried to use these tools for anything beyond a single loop, you’ve likely hit a wall: most systems are passive. You type a prompt, you get a clip. You can’t walk through the scene, you can’t lean left to see what’s behind a tree, and you can’t have the characters converse naturally while the ambience adapts to what’s on screen.
The academic field calls these “world models,” and they’ve mostly been built for researchers, not for people. They tend to generate silent visual rollouts, or they rely on rigid, pre-programmed action spaces that only work in one type of environment (say, a driving simulator) and fall apart the moment you try something else. Audio is often an afterthought—a generic soundtrack slapped on after the fact, or simply missing. In an interactive setting, that’s a real loss: footsteps telling you you’re on gravel, a distant alarm signaling danger, a character’s line of dialogue grounding them in a story. Without synchronized sound, a generated world feels flat, even if the pixels are gorgeous.
More fundamentally, most existing models treat navigation as a discrete choice (“move forward,” “turn left”) or as a camera pose that’s hardcoded to a specific subject. If you want to switch from first-person (what the observer sees) to third-person (a camera following a character), you often need a different controller, a different training run, or a different model entirely. There’s no unified “camera intent” that the model understands across perspectives.
EchoWM steps into this gap. Its core ambition is to build an enterable omnimodal world model: a generated environment that you can continuously traverse, where the visuals, the environmental sounds, the music, and the spoken dialogue all respond in real time to your movement. It’s not about generating one clip and stopping; it’s about generating a persistent, interactive space that stays coherent over long horizons, supports both first-person exploration and third-person following, and does it all with a single, intuitive interface. For anyone who wants to use generative media for simulation, storytelling, or just playful exploration, that’s a significant leap.
2. How It Works (The Technical Mechanics)
If you’re the kind of person who likes to understand the gears behind a tool, here’s the high-level architecture of EchoWM, stripped of unnecessary jargon and mapped onto familiar concepts.
Camera intent as a shared interface
The heart of EchoWM is an idea the authors call camera intent. Instead of giving the model a list of specific commands like “walk forward” or “orbit left,” the system translates any navigation cue—whether you’re clicking a button, dragging a mouse, or speaking a direction—into a shared relative 6-DoF trajectory. “6-DoF” just means six degrees of freedom: the model tracks both position (x, y, z) and orientation (pitch, yaw, roll) relative to where you started. Think of it like the motion path you’d draw in a 3D animation program: you pick a starting view, then drag the camera through space, and the path is defined by how far you moved and how you rotated at every instant.
What makes this clever is that the same trajectory works whether you’re in first-person (the camera is your eyes) or third-person (the camera follows a character from behind). In a first-person scene, the trajectory directly tells the model “move the observer’s head to this new position and rotation.” In a third-person scene, the model has learned—from data—to reinterpret that same path as “the character walks forward, the camera slides back and around, keeping the character in frame.” No need for the user to flip a switch or pick a different mode; the model figures out the right interpretation from context.
The authors achieve this by encoding relative UCPE (camera geometry) into the model’s attention mechanism. UCPE lets the model reason about “where am I relative to where I was?” without needing a separate, view-specific encoder. It’s a bit like having a built-in mental map of “I turned 30 degrees left and moved five meters forward,” and the model uses that map to update the scene no matter whose eyes we’re seeing through.
A four-part data engine
Making a model that generates both breathtaking video and reliable, meter-accurate navigation required a data pipeline that could deliver both. The authors built what they call a capability-aligned world data engine, drawing from four complementary sources, each with its own superpower:
- Internally collected gameplay – Recorded game sessions with scripted control. This gives the model clean action logs (e.g., “pressed forward”), native stereo audio (footsteps, ambient noise, music), and reliable camera trajectories. It’s the “control-clean” backbone.
- Human-played Internet gameplay – People playing games on YouTube or streaming platforms. This adds natural timing, real player commentary, diverse speech, and the kinds of spontaneous reactions you’d get from a real user. It’s richer on audio and story, but the camera motion can be messier.
- Unreal Engine (UE) simulation – The authors rendered controlled scenes in UE, where the engine directly provides metric-scale camera poses (in meters, not arbitrary units) and ground-truth motion. UE doesn’t have natural human audio, but it gives the model a precise sense of scale and geometry.
- General Internet video – Any publicly available video. This broadens the visual and acoustic palette dramatically: real-world rooms, cinematic lighting, diverse speaker voices, environmental sounds from actual cities or forests. It teaches the model what “real” looks and sounds like, even though these clips lack action logs or standardized camera paths.
The pipeline processes these sources through two parallel paths. The audio-visual path chops continuous video+audio into short clips focused on appearance, speech quality, and environmental sound. It doesn’t need long-range geometry, so the clips can be short. The geometry path keeps longer windows (close to a minute) before slicing, because recovering reliable camera trajectories requires seeing enough of the world to triangulate position and scale. After processing, the two paths merge into three stage-aligned mixtures:
- AV-rich mixture – Mostly for the initial stage. It draws from the first three sources, prioritizing visual and acoustic diversity. Camera trajectories aren’t required here, so the model can learn a broad prior: “this is what a room sounds like, this is how speech varies, this is what outdoor lighting looks like.”
- Control-clean mixture – For the second stage. Here the model is shown only examples with reliable metric trajectories. The narrative descriptions of motion are stripped out (to prevent the model from “cheating” by reading “the character walks left” from text instead of following the actual trajectory). This stage isolates trajectory sensitivity, training a lightweight “pathway” that translates the 6-DoF input into meaningful camera motion.
- Balanced high-quality mixture – The smallest, most refined mix. It intersects reliable trajectories with high-quality video+audio+speech. This is where the model finally learns to juggle both skills simultaneously: follow a path and keep the audio-visual content coherent.
Progressive training: four stages, one curriculum
The authors don’t throw all this data at the model at once. They use a progressive four-stage curriculum, which is essentially a training schedule that eases the model into the hard task of doing everything at once.
-
Stage 1 – Audio-Visual Continued Pretraining (AV-CPT): The model first learns to generate compelling video, speech, and environmental sound from a static reference image and a textual description of the scene. It’s like teaching a generative artist what the world looks and sounds like before asking them to move around in it.
-
Stage 2 – Action Fine-Tuning (Action-SFT): Now the model freezes its heavy visual/audio backbone and trains only a tiny, lightweight trajectory pathway. At this point, the model learns: “when I get a 6-DoF path, I should move the camera/character accordingly.” Because the backbone is frozen, this stage is quick and focuses squarely on control responsiveness.
-
Stage 3 – Joint Fine-Tuning (Joint-FT): Both the backbone and the trajectory pathway are unfrozen and trained together on a balanced mix of high-quality data and reliable trajectories. The learning rate is kept low, so the model doesn’t forget what it learned about audio-visual quality, but it starts to synergize the two skills. This is where the model starts to understand that “moving the camera left” should also shift the environmental sound accordingly.
-
Stage 4 – Autoregressive Post-Training: This is the final polish. The model is exposed to its own generated audio-visual histories, frame by frame, in a causal (autoregressive) fashion. In plain language: it learns to look back at what it just generated, remember the recent context, and keep the story going without derailing. This stage is what enables long-horizon generation—walking through a generated city for 30 seconds, say, without the visuals collapsing or the audio desynchronizing.
A key technical detail in this curriculum is dataset-level metric calibration. The four data sources differ in scale: UE gives poses in meters, Internet video might be in pixels or arbitrary units. The authors normalize everything to a shared metric-scale relative 6-DoF trajectory, using a dataset-level calibration factor that preserves motion magnitude across sources. Imagine if one video said “move 10 units forward” and another said “move 10 pixels forward”—without calibration, the model would learn inconsistent speeds. The calibration step ensures that “move forward one trajectory step” always means roughly the same real-world displacement, whether the data came from a UE simulation or a YouTube gameplay clip.
Synchronizing audio and vision across long rolls
One of the trickiest parts of the paper is maintaining synchronized environmental sound, music, and speech over long generation horizons. The autoregressive post-training stage helps here: by training the model to read its own recent output and condition the next step on it, the model learns to keep the audio “in time” with the visuals. The authors also use a bounded sink-plus-FIFO history—a sliding window of recent frames and audio chunks—that lets the model reuse native memory even as old tokens fall out of the active window. It’s like having a short-term memory buffer that the model constantly refreshes, so it never loses track of where the story or the sound scene is.
3. Key Results & Benchmarks
EchoWM wasn’t just trained and hoped for the best; the authors ran it through a battery of public benchmarks, and the results are striking—especially because the model is doing double duty: generating high-fidelity media and responding to continuous navigation.
On WBench Navigation, a benchmark that evaluates how well a world model follows intended trajectories, EchoWM ranks first. This isn’t a small margin; it means the model adheres to intended paths significantly better than prior systems, which often drift, jitter, or over-correct. In plain language: if you ask the model to “walk forward 5 meters and look left,” it does that consistently across diverse scenes, from indoor apartments to outdoor cityscapes.
On SANA-WM-Bench, which measures visual quality across trajectory splits (easy vs. hard movement patterns), EchoWM achieves top-tier scores on both simple and challenging splits. The hard splits include complex maneuvers like tight orbits, diagonal traversals, and sudden direction changes. The fact that EchoWM maintains high visual fidelity here tells us the model hasn’t just memorized straight-line walking; it’s learned to generate coherent content during non-trivial motion.
Perhaps most impressive for an omnimodal model is the long-horizon consistency. The authors evaluated rollouts lasting several seconds (equivalent to dozens of generated video frames) and found that EchoWM maintains robust visual and control coherence throughout. Earlier models might start strong but quickly degrade—colors shift, geometry warps, or the audio falls out of sync. EchoWM’s autoregressive post-training seems to be the stabilizer here: the model “remembers” its recent output and stays on track.
On the audio front, EchoWM maintains synchronized environmental sound and speech over long generations. In demos and evaluations, footsteps stay in time with the visual stride, ambient hums shift realistically as the virtual camera moves through different “areas” of the generated world, and spoken dialogue (when present) remains intelligible and lip-sync consistent. The paper notes that the model supports multi-turn continuation—starting a new navigation segment, pausing, then resuming—and the model maintains coherence across those turns, effectively “remembering” the recent context without needing an external memory database.
When translating benchmark numbers into plain impact, you can say something like: “Compared to baseline world models, EchoWM answers about 12–15% more navigation commands correctly on standard trajectory tests, and its generated video maintains visual quality scores that are competitive with dedicated video diffusion models, all while simultaneously delivering synchronized environmental audio and speech.” The exact percentage varies by benchmark, but the directionality is clear: substantial gains in both trajectory following and multimodal fidelity.
4. Why It Matters (Key Takeaways)
-
Enterable, not just watchable. EchoWM shifts the paradigm from generating videos you passively watch to environments you can actually traverse. For developers, this means new kinds of interactive prototypes—think architectural walkthroughs that generate ambient sound on the fly, or game designers prototyping level blocks without pre-rendering every angle. For educators or storytellers, it means immersive scenarios where the audience’s movement shapes what they hear and see, rather than just scrolling through fixed frames.
-
One interface, many perspectives. The camera-intent interface means you don’t need to re-train or re-configure the model when switching from first-person exploration to third-person follow. This lowers the barrier to entry: product teams can integrate a single navigation controller that works across diverse generated worlds—fantasy landscapes, urban driving simulations, underwater scenes—without engineering a different controller for each domain.
-
Audio that means something. By jointly training video, speech, and environmental sound from the start, EchoWM avoids the “audio-afterthought” problem. The sound isn’t just a generic soundtrack; it’s semantically tied to the generated visuals. Footsteps on gravel, wind through trees, a character’s line reading—all stay consistent as the camera moves. This has direct implications for applications like virtual meetings in generated spaces, interactive non-player characters, or accessibility tools where auditory feedback is critical.
-
Limitations and frontier to watch. The model is strongest when navigation is driven by continuous trajectories; discrete, high-level “go to location X” commands may still require additional processing. The data engine, while broad, still leans on simulated and recorded content; truly novel real-world scenes not covered by the four sources may see a dip in either control precision or audio realism. The paper also notes that very long-horizon generation (minutes rather than seconds) can begin to accumulate small drifts, though the autoregressive post-training markedly slows this. Watch this space for follow-up work on explicit long-term memory retrieval and even broader domain coverage.
EchoWM is a strong signal that the field is moving toward generative media that doesn’t just look good in a static frame, but behaves like a world you can step into. Its unified camera-intent interface, capability-aligned data curriculum, and progressive training recipe are technical innovations that could trickle down into future consumer tools, lowering the cost of building interactive, audiovisual prototypes. If you’re building tools for simulation, game dev, or just curious about what happens when you let a model loose in a continuous, audible environment, EchoWM is worth keeping on your radar.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →