Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMsExplained for Beginners
Yunheng Li, Guohong Mu, Hao Li +4 more
Abstract
Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, post-training on large multi-task datasets remains challenging, as existing reinforcement learning methods sample on-policy groups with few high-quality rollouts even with costly chain-of-thought (CoT) generation. In this paper, we study the sample efficiency and scalability of RL post-training for video MLLMs and introduce OraRL. We identify an overlooked role for annotations: Beyond scoring rollouts, each can enter its on-policy group as an oracle rollout, a direct positive optimization target. Direct oracle integration, however, is nontrivial: a high-reward oracle raises the group baseline and inverts otherwise positive policy advantages, a failure we term advantage inversion. At the core of OraRL is a decoupled advantage estimator: policy rollouts determine an oracle-free baseline, while the oracle-policy gap modulates both a directional gain and a separate detached oracle advantage. Sign-balanced pruning improves efficiency: by retaining only the oracle and the strongest rollouts of each sign, OraRL requires just 2.2x the step time of SFT, less than half the 4.9x required by GRPO with CoT. OraRL scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts. Without chain-of-thought, Video-ORA-9B decodes in 130 ms instead of 4,780 ms. Compared with the respective prior best models, it raises temporal mIoU from 62.5 to 66.0, tracking AO from 73.0 to 78.2, segmentation from 64.3 to 70.4, and the three-benchmark spatial-intelligence macro average from 51.0 to 56.1; on VSI-Bench, it scores 73.1 against 55.0 for GPT-5 and 55.1 for Gemini-3-Pro.
Annotations as Rollouts: Efficient and Scalable RL for Video MLLMs
The Problem
Multimodal large language models (MLLMs) have become the go-to architecture for unified video perception—they can describe what’s happening in a video, localize events in time, ground objects in space, and even track objects across frames. But there’s a catch: after the initial supervised fine-tuning (SFT) stage, getting these models to perform precise, task-level reasoning remains surprisingly difficult.
The core bottleneck is sample efficiency. Existing reinforcement learning (RL) methods, particularly Group Relative Policy Optimization (GRPO), try to compare the model’s own rollouts against each other to figure out what “good” looks like. The problem is that on-policy rollouts rarely land exactly on the intervals, boxes, or masks that human annotations specify. Annotations serve only as scoring references, not as optimization targets. When rollouts are sparse and high-quality examples are few, the learning signal becomes weak. Adding chain-of-thought (CoT) reasoning doesn’t really help—it just makes every rollout longer and more expensive to generate, without clear gains in accuracy.
In short: the field had a reliable source of ground truth (the annotations) but no efficient way to use them as direct supervision during RL. The paper identifies this gap and sets out to fix it.
How It Works (The Technical Mechanics)
The central idea of OraRL is elegantly simple once you see it: treat each annotation as an oracle rollout. Instead of letting the annotation merely score the model’s output, the paper appends the annotated answer directly to the on-policy group as a ninth “oracle” example. This provides a reliable positive target for every query.
But there’s a catch. If you naively include the oracle in the group used to normalize advantages, the high reward of the oracle pushes up the group baseline. Suddenly, on-policy rollouts that are actually pretty good—but not as good as the oracle—get assigned negative advantages. The paper calls this advantage inversion. It’s a subtle but fatal error: the model gets penalized for producing decent outputs because the oracle set the bar impossibly high.
OraRL avoids this by decoupling the oracle from the advantage baseline. Here’s the mechanics in three steps:
-
Baseline from on-policy rewards only. The advantage for each on-policy rollout is computed relative to the mean reward of the other on-policy rollouts. The oracle is excluded from this calculation, so the baseline reflects what the current policy can actually achieve on its own.
-
Directional gain from the oracle-policy gap. The gap between the oracle reward and the on-policy mean is captured in a “directional gain” term. This term amplifies advantages for on-policy rollouts that are above the mean, nudging them toward the oracle’s quality without inverting their signs.
-
Detached oracle advantage. A separate, bounded term encodes how far the policy is from the oracle. This term decays as the policy improves, ensuring the oracle provides strong supervision early on but doesn’t dominate the update later.
To make this efficient, OraRL uses sign-balanced pruning. After computing advantages for all rollouts (including the oracle), the method retains only a subset: the oracle, the strongest positive on-policy rollout, and a couple of negative ones. This sign-balanced selection preserves the contrast between reinforcing and suppressive signals. After the forward/backward pass on this reduced set, a post-selection moment correction re-centers the advantages so the update remains stable. The result: just 2.2× the step time of supervised fine-tuning (SFT), compared to 4.9× for GRPO with CoT.
The method also scales gracefully. Because the oracle is always present and treated as a positive anchor, the pruning budget can be allocated evenly between positive and negative rollouts. If one sign has fewer candidates, the spare slots go to the other sign. This sign-awareness is what allows OraRL to outperform magnitude-only pruning methods, which tend to select only one sign and lose the balancing effect.
Key Results & Benchmarks
The results are striking. OraRL doesn’t just nudge performance upward; it consistently beats prior best models across a remarkably wide range of video understanding tasks.
Temporal grounding (localizing events in time): Video-ORA-9B raises mIoU from 62.5 (backbone) to 66.0 on the three TimeLens benchmarks. On ActivityNet, the improvement at the strictest threshold (R1@0.7) is 6.7 points. The model also leads on GOT-10k tracking, achieving an average overlap (AO) of 73.1 versus 32 for the base Qwen3.5-9B model.
Spatial grounding and segmentation: On RefCOCO family benchmarks, Video-ORA-9B hits 94.6% R@0.5, exceeding the best baseline by over 1 point. On video segmentation (MeViS), J&F jumps from 52.7 (backbone) to 61.3—a gain of 8.6 points. On ReasonVOS, the improvement is a massive 42.2 points, from 59.9 to 63.8. These are not marginal gains; they represent the difference between a model that can roughly point at an object and one that can accurately segment it.
Video question answering: Without any chain-of-thought, Video-ORA-9B ranks first on five of seven Video QA benchmarks and lifts the macro average from 61.9 to 66.8 over the Qwen3.5-9B backbone. On the challenging LongVideoBench, the gain is a modest 1.6 points, but on VideoHolmes it’s 15.2 points—showing the method’s strength on reasoning-heavy tasks.
Spatial intelligence: This is perhaps the most impressive domain. On VSI-Bench, the macro average jumps from 55.1 (Gemini-3-Pro) to 73.1 for Video-ORA-9B. That’s a gain of 18 points over the prior best proprietary model. The model also scores 78.2 on the “relative distance” subset and 90.6 on “appearance order,” indicating strong capabilities in spatial reasoning tasks that typically require metric understanding.
Efficiency gains are equally dramatic. Without chain-of-thought decoding, Video-ORA-9B generates answers in 130 ms per video, whereas the CoT-enabled backbone takes 4,780 ms. That’s a 36× speedup in generation latency. On ten-minute videos, answer-only generation reduces median end-to-end latency from 29.0 seconds to 24.3 seconds.
The paper also evaluates scaling. OraRL improves its backbone from 0.8B to 9B parameters across all seven task families, with macro-average scores rising from 51.8 to 66.2. It also outperforms GRPO at every data budget tested, up to 100k prompts. The scaling advantage comes from the fact that every added prompt brings a reliable oracle rollout, whereas GRPO must wait for enough high-quality sampled rollouts to generate a meaningful comparison.
Why It Matters (Key Takeaways)
-
Annotations as a free lunch. The most immediate takeaway is that annotations—already present in the dataset for other purposes—can serve as direct optimization targets. This means the sample efficiency of RL post-training improves dramatically: every annotated query now carries a positive supervision signal, not just a scoring reference. For practitioners, this translates to faster convergence and better data utilization.
-
Solving advantage inversion. The paper’s analysis of advantage inversion is a valuable contribution beyond this specific work. Many RL methods for multimodal models risk collapsing when a high-reward reference (like an oracle) is mixed into the normalization baseline. OraRL’s decoupled estimator provides a principled solution that could generalize to other domains where strong supervision is available but expensive to sample.
-
Efficiency without sacrificing quality. The fact that answer-only generation decodes in 130 ms instead of 4,780 ms (a 36× improvement) while simultaneously raising scores across the board is noteworthy. It suggests that oracle-guided RL can produce models that are both smarter and faster—a rare combination in post-training.
-
Scaling consistency. OraRL’s performance increases monotonically with model size (0.8B → 9B) and data budget (up to 100k prompts). This consistency is important for product teams: they can expect predictable gains as they invest more compute, rather than hitting diminishing returns.
-
Limitations to watch. The method assumes annotations can be serialized into the model’s response format and evaluated by a scalar reward. If annotations are ambiguous, partial, or noisy, the oracle signal may be misleading. Additionally, while OraRL outperforms open-source baselines broadly, it still trails proprietary models like GPT-5 on some spatial intelligence tasks. Future work should explore learned oracles and broader supervision regimes.
In summary: OraRL reframes the role of annotations in video MLLM post-training. By treating each annotation as an oracle rollout and carefully decoupling it from the advantage baseline, the method avoids the pitfall of advantage inversion while delivering consistent gains across seven task families. The result is a reinforcement learning paradigm that is not only more sample-efficient but also dramatically more inference-efficient—answer-only decoding at 130 ms versus nearly 5 seconds with CoT. For anyone building or fine-tuning video MLLMs, OraRL offers a practical, scalable path to much stronger fine-grained perception without requiring additional annotation collection or chain-of-thought infrastructure.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →