BriefPulse Practical AI · Working notes on AI you can actually use. RSS · BriefPulse network
BriefPulse Practical AI

What changed in AI, what it is useful for, and what you can do with it.

15 September 2026

Brief

Nemotron specialists reach IMO gold threshold with open natural-language pipeline

Two Nemotron 3 Ultra specialist checkpoints scored 30 of 42 points at IMO 2026, the gold-medal threshold, working entirely in natural language with no formal prover, external tools or internet access. The release includes the checkpoints, the training data, the training and inference code, the submitted solutions and a 200-problem benchmark.

IMO 2026 score: 30 out of 42 points
Original graphic: drawn from the figures in this story, not a stock image.

Starting from Nemotron 3 Ultra, the authors train two specialist checkpoints using supervised fine-tuning and reinforcement learning. An iterative search generates, verifies and refines candidate proofs, and a separate high-compute stage selects each final submission.

The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold, and operates entirely in natural language, with no formal prover, external tools or internet access.

The release covers the two post-trained checkpoints, the training data, the training and inference code, the submitted solutions and Nemotron-IMO-Bench, a new benchmark of 200 novel olympiad-level problems.

Source details and supporting facts

Each line is stated by the page named above it.

Stated by arXiv

  • The system scored 30 out of 42 points at IMO 2026, reaching the gold-medal threshold.
  • The system operates entirely in natural language, with no formal prover, external tools, or internet access.
  • Starting from Nemotron 3 Ultra, the authors train two specialist checkpoints using supervised fine-tuning and reinforcement learning.
  • The release includes the two post-trained checkpoints, the training data, the training and inference code, the submitted solutions, and Nemotron-IMO-Bench.
  • Nemotron-IMO-Bench is a new benchmark of 200 novel olympiad-level problems.

Sources

  1. arXivText stored 13 September 2026

How this story was checked. Written from the 1 page listed above, stored 13 September 2026; claims checked against that stored text on 14 September 2026.

What that means
  • 5 of 5 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
  • Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
  • The check reads stored text only: no claim rests on a fresh look that did not happen.
  • Where the reporting was silent, the text says so instead of filling the gap.

More from Practical AI