arXiv:2608.24763nvidia/nemotron-3.5-lightning-30b-a3bAugust 25, 2026

MoTE: Mixture of Task Experts for Multi-Task Video UnderstandingExplained for Beginners

Muhammad Asad Ali, Umar Khan, Nadia Robertini +1 more

Computer Vision and Pattern RecognitionMachine Learning

Abstract

Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transformer decoders share the same feed-forward networks across tasks, which can entangle task behavior and make controlled capability expansion difficult. Sparse Mixture-of-Experts (MoE) decoders provide conditional computation, but token-level learned routing is not naturally aligned with task-level procedural objectives. We propose MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared. Each example follows one sample-level task route, so active task-expert computation remains independent of the number of stored task experts. We instantiate this design as VideoLLM-MoTE and evaluate it on five COIN benchmarks using explicit task routes. The five-expert model activates ~2B LLM parameters per sample and achieves higher average top-1 accuracy than recent VideoLLM baselines. Under the same expert topology, it improves over dense all-expert activation and learned sparse-routing controls. These results show that task-structured routing provides an interpretable and compute-efficient decoder alternative for multi-task video-language learning.

Here is the structured explanation of the MoTE paper.

1. The Problem

Imagine a smart-video assistant that needs to look at a clip of someone cooking and instantly answer three different questions: "What is the cook doing right now?", "What happens next?", and "What is the overall goal of this recipe?". This is the core challenge of procedural video-language modeling.

The paper identifies a specific technical bottleneck: current models typically use a dense transformer decoder, where the same feed-forward networks are shared across all tasks. Because every task relies on the exact same neural pathways, the model struggles to specialize. This leads to negative transfer—where learning one task (like forecasting the next step) inadvertently degrades performance on another (like recognizing the current action). For a product team, this means you can’t easily add new capabilities to a deployed model without risking that existing features break. The paper argues that we need a way to keep the model's "visual understanding" shared, but give each task its own dedicated "thinking pathway" inside the decoder.

2. How It Works (The Technical Mechanics)

The proposed solution is MoTE (Mixture of Task Experts). To understand it, it helps to think of the Large Language Model (LLM) decoder as a block of experts rather than a single thinking path.

In a standard model, every layer has one Feed-Forward Network (FFN) that processes the information. MoTE modifies this by replacing that single FFN in each layer with a switch and a set of experts. Specifically, for LL decoder layers, MoTE creates a Shared Expert and MM Task Experts.

Here is the concrete analogy: Think of the decoder as a restaurant kitchen. In a dense model, there is one kitchen staff that must be prepared to cook every dish on the menu simultaneously. This is inefficient and messy. MoTE builds a kitchen with a shared station (for basic prep like chopping vegetables that every dish needs) and several task-specific stations (a "Salad Station," a "Grill Station," etc.).

When an order (a video prompt) comes in, the head waiter doesn't route individual ingredients (tokens); instead, they look at the order ticket (the task identity) and send the whole order to the specific station best suited for that dish.

The Mechanics:

  • Shared Backbone: The visual encoder (watching the video) and the attention mechanisms (linking visuals to text) remain shared. This ensures the model understands the video content once.
  • Task Experts: Each decoder layer now has experts specialized for specific tasks (e.g., one expert is great at recognizing actions, another at forecasting future steps).
  • Sample-Level Routing: This is the key innovation. For any given video clip, the model looks at the prompt and says, "This is a 'Next Step' question." It then activates only the expert(s) responsible for "Next Step" across all decoder layers. The other experts sit idle.
  • Inference Flexibility: At test time, if the task label isn't provided, the model can match the text prompt to a description of the experts (e.g., matching "What is the next action?" to the "Next Step" expert description).

3. Key Results & Benchmarks

The paper evaluates VideoLLM-MoTE on the COIN benchmark, which consists of five procedural tasks. The results are striking when we look at the "active parameter" budget.

  • The Parameter Trade-off: The strongest baseline models (like VideoLLM-online-8B) activate a massive 8 billion parameters per video clip to get good accuracy. In contrast, VideoLLM-MoTE-1B+5E activates only ~2 billion parameters per clip (the shared expert plus one task expert). This is a 4x reduction in computational cost per sample.
  • The Accuracy Gain: Despite using 4x fewer active parameters, MoTE achieves a 62.9% average top-1 accuracy across the five COIN tasks.
  • Comparison: This outperforms the 8B baseline (which scores ~61.8% average) and other recent VideoLLM baselines. It also beats "dense all-expert activation" (where every expert runs for every clip) and "learned sparse-routing" (where the model guesses the task via token-level gates).

Translation to plain language: You can have a model that is nearly as accurate as a massive 8-billion-parameter model, but run it using the compute budget of a 1-billion-parameter model. Specifically, on procedure forecasting tasks (which require structured, multi-step answers), MoTE shows the largest gains, proving that aligning the routing with the task structure helps the model handle complex, long-horizon predictions.

4. Why It Matters (Key Takeaways)

  • Compute Efficiency: By activating only the necessary expert, MoTE significantly reduces the energy and time required to run inference. This makes deploying multi-task video assistants on edge devices or in real-time streaming scenarios more feasible.
  • Modular Expansion: The architecture allows for "plug-and-play" addition of new tasks. To add a sixth task (e.g., "detect if a safety violation occurred"), you simply add a new expert module. The existing experts and the shared backbone remain frozen, preventing the model from forgetting what it previously learned (mitigating catastrophic forgetting).
  • Interpretability: Because routing is tied to explicit task identities rather than opaque learned gates, researchers and developers can diagnose exactly why a model made a certain prediction by looking at which expert was activated.
  • Cross-Domain Applicability: The principle isn't limited to videos. The authors demonstrate that the same MoTE conversion works for document understanding. By adding a "KIE" (Key Information Extraction) expert to a document OCR model, they boost performance from 20.07% to 95.79% on receipt parsing while perfectly preserving the original OCR capability. This proves the method is a general-purpose tool for scaling AI models horizontally as task requirements grow.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →