Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsYes—AI systems have solved multiple International Mathematical Olympiad (IMO) problems. Google DeepMind reported that AlphaProof and AlphaGeometry 2 solved four of six IMO 2024 problems for 28 of 42 points, a silver-medal-equivalent score. In 2025, DeepMind said an advanced Gemini system reached gold-medal-standard performance using natural-language statements within the 4.5-hour contest limit. Those results demonstrate major progress on defined olympiad benchmarks, not the ability to solve arbitrary unsolved mathematics.
What AI achieved at the IMO
The clearest public milestones come from two different evaluations.
| Evaluation | System | Problems solved | Reported score or result | Important qualification |
|---|---|---|---|---|
| IMO 2024 | AlphaProof plus AlphaGeometry 2 | Four of six | 28/42, equivalent to a silver medal under the IMO scoring rubric | Formalization was required, and the total computational effort exceeded the human contest window. |
| IMO 2025 report | Advanced Gemini with Deep Think | DeepMind reported gold-medal-standard performance | No official human-contest medal was awarded to the system | DeepMind described natural-language input and rigorous proofs produced within the 4.5-hour limit. |
The IMO has six problems worth up to seven points each, so 28 points is a substantial result. It should be read as a benchmark score, not as evidence that a machine can routinely solve every olympiad problem.
How AlphaProof and AlphaGeometry 2 divided the 2024 problems
AlphaProof: algebra, number theory and formal proof search
According to Google DeepMind’s 2024 account and a later Nature description, AlphaProof solved three of the five non-geometry problems: two in algebra and one in number theory. The solved set included the hardest problem in that group.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
- Used Book in Good Condition
AlphaProof trains itself to prove statements in Lean, a formal programming language and proof assistant. Its reinforcement-learning system searches for proof steps that Lean can check, then learns from successful and unsuccessful attempts. Google Research has also described test-time reinforcement learning in which the system generates and learns from many related problem variants while working on a task, allowing problem-specific adaptation.
AlphaGeometry 2: a specialized geometry system
AlphaGeometry 2 solved the geometry problem. DeepMind reported a 19-second solution after the problem had been supplied in formalized form. The system combines language-model guidance with symbolic geometry reasoning and can generate useful auxiliary constructions—extra points, lines or relationships that make a proof possible.
Rank #2
Geometry is handled separately because diagrams, constructions and geometric relations are represented differently from algebra, number theory and combinatorics. DeepMind also reported that AlphaGeometry 2 solved 83% of geometry problems from the preceding 25 years of IMO contests in its historical evaluation; that figure is a system-specific benchmark, not a prediction of performance on every future geometry problem.
Why the 2024 silver-equivalent score is not the same as a human medal
The 28/42 total was mapped to the IMO scoring rubric, which makes it useful for comparison. The workflow, however, differed from that of a student sitting the contest.
Rank #3
- Formalized input: The machine systems worked with formal statements, including a formalized geometry problem, rather than only the ordinary contest text and diagram.
- Different proof format: AlphaProof produced Lean-checkable proofs. Human contestants submit written mathematical arguments judged by IMO coordinators.
- More computation: The Nature account says the main AlphaProof training was halted and its hyperparameters were frozen before the official 2024 problems, but the total effort used to obtain solutions extended beyond the human contest time constraints.
- Specialized components: AlphaProof and AlphaGeometry 2 were separate systems aimed at different mathematical domains, whereas a contestant uses one general problem-solving process across all six questions.
For these reasons, “silver-medal equivalent” accurately describes the score under the rubric, but not an identical human-versus-machine competition.
What changed in DeepMind’s 2025 result
DeepMind’s 2025 announcement described a stronger end-to-end setup. It said an advanced Gemini model with Deep Think received the official problems in natural language and produced rigorous mathematical proofs within the 4.5-hour competition time limit.
This addresses the largest practical objection to the 2024 demonstration: the need to prepare formal statements and use computation outside the contest schedule. It is therefore a more direct test of contest-style performance. Nevertheless, it remains a company-reported evaluation on six fixed problems, not an officially administered IMO entry. “Gold-medal-standard” is DeepMind’s description of the performance; it does not mean the system was ranked alongside student contestants or received an official medal.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to compare AI olympiad results fairly
When a new claim appears, check the following dimensions rather than relying on the medal label alone.
Recommended Free Tools
Best Value
| Comparison axis | Questions to ask |
|---|---|
| Problem domain | Was the system tested on geometry, algebra, number theory, combinatorics, or a mixture? |
| Input format | Did it receive ordinary natural-language statements, a diagram, or a human-prepared formalization? |
| Proof representation | Is the result a Lean-checked formal proof, a generated natural-language proof, or a score based on an external evaluator? |
| Time and compute | Was the system restricted to the 4.5-hour contest limit, or did training, formalization and multi-day search contribute? |
| Scope of success | How many of the six problems were solved, and were partial points counted? |
| Verification | Was the result independently checked or officially administered, or reported only by the system’s developer? |
A claim that scores highly on one axis can still be weak on another. For example, a formal proof checked by Lean provides strong verification of correctness, while natural-language input better reflects the way human contestants encounter problems. Neither fact alone establishes general mathematical intelligence.
Does this mean AI can solve unsolved mathematics?
No. IMO problems are exceptionally difficult, but they are still a fixed, curated benchmark with a known statement and a short expected proof. Success on six contest problems does not demonstrate that a system can choose productive research questions, develop new theories, recognize when a conjecture is false, or sustain an independent mathematical investigation.
The results do establish narrower conclusions:
- AI can now produce verified solutions to several olympiad-level problems.
- Combining learned guidance with symbolic or formal reasoning is effective on problems with precise structure.
- Natural-language, time-limited performance has improved enough for DeepMind to report gold-medal-standard results on the 2025 set.
The appropriate interpretation is rapid progress in mathematical-reasoning benchmarks, not a universal theorem-proving ability or a replacement for human mathematical insight.
Bottom line for students and mathematicians
AI has crossed a meaningful threshold: it can solve selected IMO problems, and the reported performance has moved from a 2024 silver-equivalent result using specialized formal systems to a 2025 gold-medal-standard claim using natural language and the contest time limit. To understand what any future result means, look past the medal analogy and inspect the input, proof checker, time budget, compute, problem coverage and verification method.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

