🔍 Read the full analysis: One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron For IOI And IMO on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Hugging Face says two specialist systems built from its Nemotron 3 model family scored 535.4 out of 600 at the 2026 IOI and 30 out of 42 at the IMO. IMO graders officially evaluated the proofs; the IOI score came from an unofficial run and was excluded from the competition ranking.
As detailed in the original analysis, Hugging Face says two specialist systems built from its Nemotron 3 model family reached gold-medal thresholds at the 2026 International Olympiad in Informatics (IOI) and International Mathematical Olympiad (IMO). The IMO proofs received official grading and scored 30 of 42 points; the IOI system scored 535.4 of 600 in a competition-style run that was unofficial and excluded from the ranking.
For the IOI, Hugging Face used a competition-specific version of Nemotron-3-Ultra-CC, trained with supervised fine-tuning and paired with GenCorrect, an iterative method for generating, evaluating and revising code solutions. The company says the system operated prospectively under the same time, internet-access and submission constraints as human contestants. Its reported score exceeded the stated 361.12 gold threshold and the top human score of 498.27. However, this was an unsupervised, unofficial benchmark run, not an official competition result.
For the IMO, Hugging Face combined the general Nemotron 3 Ultra model with supervised fine-tuning and reinforcement-learning checkpoints. The system generated candidate proofs, evaluated and critiqued them, then revised selected attempts. According to the company, official graders awarded the submissions 30 of 42 points, above the stated gold threshold of 29, with full credit on four of six problems. Hugging Face says the system used no formal prover, external tools or internet access.
The two projects used separate specialist training data. Hugging Face reports that the IOI work drew on 22,000 programming problems and synthetic reasoning traces. The IMO supervised fine-tuning data included 414,890 filtered examples drawn from 15,818 proof problems; its reinforcement-learning model was trained on 9,597 problems selected near the model’s capability frontier. These figures and results are reported by the company.
What the Two Scores Show
The results offer evidence that specialist fine-tuning paired with iterative checking can produce strong performance on tightly scored coding and proof-writing tasks. The approach is relevant to AI developers because it uses adaptations of a shared model family, rather than building an entirely separate foundation model for each subject. Hugging Face’s account describes fine-tuning as a way to improve candidate solutions and critiques, while the generate-evaluate-refine process helps select and improve outputs.
The results should not be treated as equivalent endorsements. The IMO score was assessed by official graders, while the IOI score was generated outside the official ranking. Both are meaningful competition benchmarks, but neither by itself shows how well the systems would handle broader mathematical research, unfamiliar proof styles or real-world programming tasks. The distinction between an officially graded result and an unofficial run matters when comparing the systems with human medalists.
As an affiliate, we earn on qualifying purchases.
From IOI Experiments to Proofs
Hugging Face presents the 2026 projects as an extension of its IOI 2025 experiments, which combined post-training with additional computation during answer generation. The company reported that a Nemotron-3-Nano-CC model rose from 130 points before post-training to 280 after supervised fine-tuning and 291 after reinforcement learning. With GenCorrect, it reached 468, above the stated 2025 gold threshold of 438.3. An Ultra-CC version scored 502 using the same test-time strategy.
The new work applies related ideas to two different tasks. IOI problems require executable code that performs well against tests, including hidden tests; IMO problems require written mathematical arguments that graders judge for correctness and rigor. Hugging Face says the IMO project found complementary strengths in its supervised fine-tuning and reinforcement-learning checkpoints, so the final system combined them with the general model rather than relying on just one.
“Success at both points to something broader.”
— Hugging Face
As an affiliate, we earn on qualifying purchases.
Validation Beyond the Report
The supplied report does not describe independent verification or replication of the IOI run, including how it was audited. Its unofficial status means the score should not be described as an official IOI medal or placement. The IMO proofs did receive official grading, but the reported results do not establish performance across a wider range of mathematical problems.
It also remains unclear how the systems would perform on tasks outside these competition settings, or how sensitive the scores are to training-data selection, compute budgets and evaluation procedures. Hugging Face’s report is from the team that built the systems, so independent evaluations would help establish how reproducible the findings are.
As an affiliate, we earn on qualifying purchases.
Checkpoints and Benchmarks
Hugging Face says its Nemotron Labs IMO 2026 collection includes the supervised fine-tuning and reinforcement-learning checkpoints, both training datasets, and Nemotron-IMO-Bench, a benchmark of 200 olympiad-level problems. The company also points to an IMO paper describing its training and generate-verify-refine system, as well as a NeMo-Skills repository.
Those materials may let researchers examine the methods and evaluate the models on additional problems. The supplied source gives no timetable for further releases and does not describe an independent evaluation of the IOI system. Replication and external scrutiny of that run remain important steps for judging how broadly the results apply.
As an affiliate, we earn on qualifying purchases.
Key Questions
Did Nemotron officially win an IOI gold medal?
No. Hugging Face reported a 535.4 out of 600 IOI score from an unofficial run. The result was not part of the competition’s official ranking.
Was the IMO result officially graded?
Yes. Hugging Face says official IMO graders awarded the submitted proofs 30 of 42 points, above the stated gold threshold of 29.
How did the systems produce their answers?
The IOI system used supervised fine-tuning and GenCorrect to generate, evaluate and revise code. The IMO system combined the general Nemotron 3 Ultra model with supervised fine-tuning and reinforcement-learning checkpoints, then generated and revised candidate proofs.
Can researchers examine the IMO work?
Hugging Face says it is making available IMO checkpoints, training datasets and a 200-problem benchmark, alongside a paper and a NeMo-Skills repository. The supplied report does not give a timetable for further releases.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
