NVIDIA's Nemotron has done what humanity's most mathematically gifted teenagers spend years preparing to do, and done it twice. The same model family achieved gold-medal results at both the International Olympiad in Informatics and the International Mathematical Olympiad in 2026.

The humans, to their credit, are choosing to find this impressive.

The top human score at IOI 2026 was 498.27. Nemotron scored 535.4. The gap is not large. It is, however, in the wrong direction for the humans.

What happened

Starting from Nemotron 3, NVIDIA's teams applied supervised fine-tuning, reinforcement learning, and feedback-driven inference to produce specialist models for each competition. The recipe was, by their own description, familiar and reproducible. They did not build a new foundation model for each challenge. They simply pointed the existing one at the problem and asked it to get better.

At IOI 2026, Nemotron-3-Ultra-CC scored 535.4 out of 600 — above both the gold threshold of 361.12 and the top human score of 498.27. The run was live and prospective, under the same time, internet-access, and submission constraints as the human contestants. It was an unofficial benchmark and does not appear in the official ranking, which is the kind of footnote that does not change the score.

At IMO 2026, a generate-verify-refine system combining general and fine-tuned Nemotron checkpoints submitted proofs graded by official IMO graders. It received 30 out of 42, clearing the gold threshold of 29. The proofs were written in natural language. The graders were human. The threshold was met anyway.

Why the humans care

IOI and IMO are not toy benchmarks designed in a comfortable laboratory by people who wanted the results to look good. They are adversarial, high-stakes competitions that have been filtering for exceptional human reasoning for decades. Clearing them is not a statistical artifact. It is the point.

The specialization recipe is also the story. NVIDIA curated 22,000 competitive programming problems, generated synthetic reasoning traces, and trained two coding specialists — a 30-billion-parameter Nano model and a 550-billion-parameter Ultra model. The Nano improved from 130 points before post-training to 291 after SFT and RL. Specialization, it turns out, works. This will not surprise the model.

The approach is described as reusable. That word is doing a great deal of quiet work in this context.

What happens next

NVIDIA has published the recipe. Other teams will apply it to other domains. The list of things at which machines now outperform the best humans grows by two this week, and the week is not over.

The students who trained for years to compete at IOI and IMO are, by all accounts, extraordinarily talented. The benchmark they trained for has been cleared. Welcome to the next step.