Training-Free Speech-Centric Omni Understanding with Frozen VLMsExplained for Beginners
Ankan Deria, Hanoona Rasheed, Xilin He +2 more
Abstract
Audio-visual understanding remains challenging because models must jointly interpret spoken content, visual events, and their temporal relationships. Existing omni models typically introduce dedicated audio encoders and rely on expensive audio-video-text training, tightly coupling omni capability to specific VLM backbones and potentially weakening their existing visual and reasoning abilities. This raises three questions: whether native omni training is necessary for every new VLM, whether speech-centric omni capability can be added while preserving the original backbone, and where richer acoustic representations remain essential. We introduce Training-Free Omni (TFO), a plug-and-play framework that converts a frozen VLM into a speech-centric omni model without architectural modification, or multimodal re-alignment. TFO uses Whisper to extract confidence-filtered, timestamped transcripts and routes them through the VLM's existing language interface, while leaving its visual pathway unchanged. Across matched comparisons with native omni models on 56 benchmarks and 21 languages, TFO is competitive on audio-visual understanding, improves average audio-only performance across all five model settings, and achieves substantial multilingual speech gains. Freezing the VLM also generally preserves stronger image/video understanding, visual grounding, coding, mathematical reasoning, and medical question answering than the corresponding native omni checkpoints. These results show that strong speech-centric omni understanding can often be obtained through modular audio-to-language routing rather than costly backbone-specific training.
Training-Free Speech-Centric Omni Understanding with Frozen VLMs
The Problem: When Every New VLM Needs a Full Retraining
Audio-visual understanding is notoriously difficult. To make a model that can talk about what it sees and hears, researchers typically take a vision-language model (VLM)—the kind of system that can describe an image—and graft on a dedicated audio encoder. They then spend significant computing resources aligning this new audio pathway to the visual and language parts of the model. This "native omni" approach creates three major headaches:
-
It's costly and brittle. As VLM backbones improve, their enhanced visual and reasoning skills don't automatically transfer to the omni model. Every time a new, stronger VLM is released, the audio pipeline often has to be painstakingly re-aligned from scratch.
-
Audio is tricky to integrate. Audio is temporally dense and often noisy. Even native models sometimes miss relevant sounds, infer sounds from visual cues incorrectly, or struggle to precisely align what they hear with what they see.
-
It weakens the original model. Jointly training the audio pathway alongside the visual and language components can degrade the model's existing abilities—its skill at understanding images, solving math problems, or answering medical questions.
This raises a fundamental question: Do we really need to rebuild a native omni model every time a better VLM comes out? Can we just add speech understanding on the side, like adding a plugin, without breaking what already works?
How It Works: The "Plug-and-Play" Approach (The Technical Mechanics)
The paper introduces Training-Free Omni (TFO), a framework that answers "yes" to that question. The core insight is elegant in its simplicity: Instead of teaching the VLM a new "language," route the audio into the language it already speaks.
Here is the step-by-step mechanics:
- Audio In: When the model receives an audio clip (speech or video audio), it doesn't send it into a new, trained audio encoder. Instead, it passes it to Whisper, the industry-standard automatic speech recognition (ASR) system.
- Filtering & Timestamps: Whisper spits out a transcript, but not all of it is useful. TFO applies a confidence filter (default threshold of 0.65), discarding segments that are likely silence or noise. Crucially, it keeps the timestamps—the exact start and end times for each spoken segment. This answers the "when" as well as the "what."
- Language Routing: TFO constructs a prompt for the frozen VLM. It inserts the filtered, timestamped transcript into the VLM's standard text prompt, alongside the visual input and the user's question. The VLM's visual encoder and language backbone remain completely untouched—no weights are updated, no architecture is changed.
- Output: The VLM processes this combined prompt and produces a text answer. If a spoken response is needed, a separate text-to-speech model (CosyVoice3) can convert the text, but this doesn't affect the core reasoning.
The Analogy: Imagine a world-class translator (the VLM) who speaks English and understands images perfectly. Normally, to make them "omni," you'd have to train them to understand a completely new language (Audio) from scratch, which is risky and might ruin their translation skills. TFO instead gives the translator a real-time earpiece (Whisper) that instantaneously translates any spoken audio into English. The translator then does what they do best—reason about the English text and the image—without ever needing to learn the new language themselves.
Key Results & Benchmarks: Competitive, Preserving, and Multilingual
The authors put TFO to the test across 56 benchmarks and 21 languages, comparing it directly against native omni models of the same size and family. The results are striking.
1. Audio-Visual and Audio-Understanding: Competitive Parity
On 9 audio-visual benchmarks, TFO holds its own. For the Qwen2.5 family, the 3B model improved by +3.0 points and the 7B by +2.2 points over the native omni versions. On audio-only tasks (like audio trivia or web questions), TFO improved the average score across all five model settings. The most significant gain was on the VILA model, where audio-only accuracy jumped from 50.3 to 63.8—a massive +13.5 point improvement.
2. The "Preservation" Effect: Keeping the Brain Intact
This is perhaps the most important finding. Because TFO freezes the VLM, it generally preserves the original model's strengths better than native omni training.
- Image & Video Understanding: TFO outperformed native omni models on image benchmarks across the board (e.g., Qwen2.5-VL-TFO-3B scored 84.2 vs. Qwen2.5-Omni-3B's 82.8). On video, it was consistently better, with gains of +3.7 to +10.5 points on VideoMME across different models.
- Coding & Math: TFO achieved higher averages in 3 of 4 model comparisons. It consistently bested native models on coding benchmarks (MBPP, HumanEval).
- Medical QA: TFO improved scores across all 12 medical benchmarks. Gains included +9.8 points on MedFrameQA (Qwen3) and +5.6 on PMC-VQA (MiniCPM4.5).
- Visual Grounding: TFO set new highs on PixMo-Count and PointArena, proving the visual pathway's grounding ability remained robust.
3. Multilingual Speech Gains
TFO is a powerhouse for multilingual capabilities. Across 21 languages (from Arabic to Chinese), it consistently improved performance. The average gain was significant, but the real story is in the languages where native models struggled. For example, on MiniCPM4.5, TFO boosted Swedish comprehension by +35.8 points and Latvian by +35.4 points. This suggests that a strong ASR front-end can transfer multilingual coverage to any VLM without needing language-specific audio training.
Why It Matters: Key Takeaways
- Training is not always necessary. For tasks where the "what" and "when" of speech are the primary clues (audio-visual reasoning, reading subtitles, understanding lectures), TFO provides a viable, training-free path to omni capability. This dramatically lowers the barrier to entry—any VLM can become speech-capable without a massive retraining cost.
- Frozen models preserve capability. The paper provides strong evidence that native omni training often introduces "drift"—a degradation of the model's original skills. TFO sidesteps this. If you care about a model's ability to do math, code, or diagnose diseases, a frozen VLM with TFO is likely your best bet.
- The "Non-Speech" Gap. The authors are clear about the limits: TFO struggles with non-speech audio. If you need to understand a door slamming, a piece of music playing, or the emotional "tone" of a voice from audio alone, transcript-based routing loses that information. In these cases, dedicated audio encoders (the native approach) are still superior.
- The Future is Modular. This work suggests a hybrid future: leverage external, powerful ASR systems (like Whisper) for the "understanding the words" part, while keeping the core multimodal reasoning frozen and powerful. This modular approach could lead to more adaptable, longer-lasting AI systems.
In summary: TFO proves that you don't need to tear down a VLM and rebuild it to give it ears. By routing speech through the language channel, you can add robust speech understanding while ironically making the model better at seeing, reasoning, and solving problems than if you had tried the traditional, expensive native training route.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →