arXiv:2609.10712nvidia/nemotron-3.5-lightning-30b-a3bSeptember 9, 2026

An Open Recipe for IMO Gold: Training Nemotron for Olympiad MathematicsExplained for Beginners

Ivan Moshkov, Stephen Ge, George Armstrong +3 more

Artificial Intelligence

Abstract

We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verification, and refinement. Based on these findings, we present an open-model test-time-compute pipeline. The system operates entirely in natural language, with no formal prover, external tools, or internet access. Three Nemotron 3 Ultra checkpoints - the general-availability model and two post-trained specialists - power an iterative search that generates, verifies, and refines candidate proofs; a separate high-compute stage then selects each final submission. The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold. We release the two post-trained checkpoints as well as the training data, the training and inference code, the submitted solutions, and Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems.

The Problem: When Language Models Meet the World's Toughest Math Contest

The International Mathematical Olympiad (IMO) sits at the pinnacle of high-school competition mathematics. Since 1959, the world's brightest pre-university minds have gathered to solve problems that demand not just computational skill, but creative insight, rigorous proof construction, and the ability to navigate deceptively simple statements that hide deep structural challenges. For AI researchers, the IMO serves as a strict benchmark: if a system can reliably navigate these problems, it likely possesses a form of reasoning general enough to tackle scientific discovery, engineering design, and complex system analysis.

Until recently, AI systems struggled at the IMO. In 2024, neuro-symbolic systems combining neural language models with formal provers like Lean scraped together enough points for a silver medal. In 2025, general-purpose large language models (LLMs) powered by specialized prompting and verification pipelines crossed the threshold into gold-medal territory. But there was a catch: these gold-medal systems were closed-source, running on expensive proprietary infrastructure, and often relied on external tools like symbolic provers or internet-accessed theorem databases.

The gap the Nemotron team addressed was this: Can an open-weight model, using only natural language and no external tools, compete with the best closed systems at the IMO? Their answer, demonstrated at IMO 2026, was a resounding yes—and the implications extend far beyond competitive mathematics.

How It Works: An Ensemble That Thinks in Circles

The core innovation of the Nemotron IMO system isn't a single breakthrough in model architecture, but a clever orchestration of multiple specialized models working together in a loop. Think of it as a quality-assembly line for mathematical proofs, where different stations catch different kinds of errors.

The system runs on three versions of Nemotron 3 Ultra, a 550-billion-parameter Mixture-of-Experts model:

  • Nemotron-3-Ultra-GA (General Availability): The base model, unchanged from its original training. It provides broad mathematical knowledge but without olympiad-specific fine-tuning.
  • Nemotron-3-Ultra-SFT (Supervised Fine-Tuned): This specialist was trained on 414,890 carefully curated proof examples, learning to construct clear proofs, identify gaps in reasoning, and evaluate their own work. Imagine a math student who not only solves problems but also learns to mark their own homework with uncanny accuracy.
  • Nemotron-3-Ultra-RL (Reinforcement Learning): This checkpoint underwent reinforcement learning on 9,597 problems, rewarded for generating proofs that verifiers judged as correct. It learned to focus on what makes a proof convincing rather than just what looks mathematically plausible.

These three models power an iterative "generate-verify-refine" loop that works like this:

  1. Generation: For each IMO problem, the system asks all three models to produce 16 proof attempts each, using eight different prompting strategies. Some prompts ask for lemma-first decomposition, others for route comparison, counterexample hunting, or invariant detection. This diversity ensures the system isn't just seeing the same solution strategy repeated 48 times.

  2. Verification: Each candidate proof gets evaluated by the RL and SFT specialists, who act as judges. They assign scores of 1 (completely correct), 0.5 (generally correct but with minor issues), or 0 (fatal errors). The system requires all 16 judgments from both specialists to agree the proof is perfect before accepting it.

  3. Refinement: If no proof clears this high bar, the system takes the best-looking candidates, asks all three models to revise them using the judges' feedback, and generates 192 new attempts (3 models × 8 prompts × 8 samples). This cycle repeats for up to eight rounds.

  4. Final Selection: Once the search concludes, a separate high-compute stage re-evaluates every remaining candidate using a stringent IMO-style grading rubric (scoring 0-7). The proof with the highest average score becomes the submission.

The critical insight is that no single model is smart enough to solve IMO problems alone every time. But by combining them—letting the SFT model catch gaps the RL model misses, and vice versa—the system creates a "wisdom of crowds" effect where errors cancel out.

Key Results: 30 Out of 42 Gets the Gold

The numbers tell a compelling story. On the official IMO 2026 competition, this open-model system scored 30 out of 42 points, surpassing the gold-medal cutoff of 29. To put this in context: in 2025, Gemini's "Deep Think" system and an experimental OpenAI model both achieved gold, but theirs was a closed-system achievement. The Nemotron system became the first open-weight model to reach this threshold without relying on formal provers, external tools, or internet access.

Breaking down the performance:

  • Full credit (7 points each) on Problems 1, 2, 4, and 5
  • Partial credit (1 point each) on Problems 3 and 6

The system solved four problems completely and made meaningful progress on the remaining two. Notably, the unofficial regrade after the competition suggested that with slightly different model judgments, the score could have reached 33 points— Problem 6's submission, for instance, might have earned 4 out of 7 with different evaluation.

In terms of computational cost, the competition run consumed approximately 707M generated tokens and 1,464 GPU-hours to reach the 30-point result. Allowing the system to continue searching past the contest deadline pushed this to 2.31B tokens and 4,800 GPU-hours, with one additional proof reaching 4 out of 7 points in unofficial grading.

Why It Matters: The Takeaways

The Nemotron IMO result carries several important lessons for anyone following AI development:

1. Post-training design matters as much as model scale. The two specialist checkpoints (SFT and RL) were created through targeted post-training on carefully designed data, not by simply making the base model bigger. The SFT checkpoint learned to construct proofs and self-evaluate; the RL checkpoint learned to focus on what makes proofs convincing. This targeted fine-tuning unlocked capabilities that the general-purpose model alone didn't possess—demonstrating that how you train a model after its initial training can be more impactful than raw scale.

2. Verification and refinement are where the intelligence hides. The system's design emphasizes careful judging over raw generation. The verification step that requires unanimity between two different specialists is a clever constraint: it prevents the system from accepting a proof that only one model thinks is good, forcing each to earn its keep. The refinement loop then gives promising candidates multiple chances to improve, with the judges providing specific, actionable feedback each time.

3. The ensemble beats the specialist, but for specific reasons. The paper rigorously tests whether spending more attempts on a single checkpoint helps versus spreading attempts across checkpoints. The results show that doubling RL attempts yields diminishing returns, but adding SFT attempts helps because that checkpoint reaches problems RL cannot. The general-availability model (GA) contributes by providing diverse proof strategies that refine the pool, even though it's the weakest pure solver. The takeaway: diversity among models matters more than pushing any one model to its absolute limit.

4. We're witnessing a shift in what "AI mathematical ability" looks like. Previous IMO AI systems relied heavily on formal provers (Lean, Coq) that can rigorously verify every logical step. Those systems achieved silver in 2024 and gold in 2025, but they required substantial engineering infrastructure and often couldn't handle problems outside their formal language's reach. The Nemotron system proves that pure natural-language models, with the right post-training and inference design, can compete at the gold level without any formal tools. This lowers the barrier for institutions without massive AI research budgets to explore mathematical reasoning capabilities.

Cautions and limits to watch:

  • The system still struggles with Problems 3 and 6, which received only 1 point each. These appear to be the types of problems where subtle gaps in reasoning are hardest to catch without formal verification.
  • The computational cost is non-trivial: 1,464 GPU-hours for a 30-point result, scaling to nearly 5,000 GPU-hours for the full search. This is feasible for well-funded AI labs but prohibitive for most applications.
  • The system operates in natural language only—graders read the proofs as written, with no automated theorem-checking catching hidden errors. A single overlooked assumption or unstated lemma could change a 7 to a 0.

Final Thoughts

"An Open Recipe for IMO Gold" is more than a competition entry; it's a manifesto for how general-purpose AI can tackle structured reasoning tasks without needing specialized infrastructure. The authors have released everything required to reproduce their results: the two post-trained checkpoints, the training data (414,890 proof examples and 9,597 RL problems), the inference code, the submitted solutions, and Nemotron-IMO-Bench—a new benchmark of 200 novel olympiad problems.

What makes this particularly exciting is the methodological transparency. The paper doesn't just report a score; it explains exactly how different design choices—checkpoint selection, verification thresholds, refinement loops—interact to produce the final result. For product managers considering AI integration into educational tools, for researchers exploring reasoning capabilities, or for anyone curious about where AI and mathematics intersect, this work provides both a functional system and a clear experimental framework.

The IMO gold medal was once thought to require systems costing tens of millions in compute infrastructure. This work suggests that with the right training recipes and test-time strategies, open models can reach the same heights—opening the door to broader participation in advancing AI reasoning capabilities. The "recipe" is now open source, and the next gold medalist might come from a lab we haven't heard from yet.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →