What IMO gold medal standard means
The International Mathematical Olympiad is an annual competition for pre-college mathematicians, run since 1959. Contestants sit two exam sessions of four and a half hours each, three problems per session, and write full proofs rather than select an answer. The 2025 edition ran on Australia's Sunshine Coast and drew 641 students from 112 countries, including five contestants who scored a perfect 42 out of 42 points across the six problems. [3]
Gold medal is not tied to a fixed score; the cutoff is set after the papers are marked, based on how the field performed that year. In 2025 the cutoff was 35 points out of a possible 42, the same total both Google DeepMind and OpenAI would go on to report for their AI systems days later. [1]
Atlas interpretation: That is the standard both companies are claiming to have cleared: proofs judged rigorous enough, on problems written fresh for that year's contest and never seen in training data, to match the performance of the strongest pre-college mathematicians in the world working the same six-problem, nine-hour paper. [1][3]
Two gold claims, two days apart
OpenAI researcher Alexander Wei posted his company's result on X on July 19, 2025, within roughly an hour of the IMO's closing ceremony: an experimental reasoning model, not released to the public, had solved five of six problems for 35 of 42 points, working from the official problem statements under the same two-session, 4.5-hour-per-session time limit as the students, with no internet access or tools. [2][6]
Google DeepMind published its own result two days later, on July 21. An advanced version of Gemini running in a mode called Deep Think had also solved five of six problems for 35 of 42 points inside the same time limit, and its proofs had been graded after the contest by the IMO's own coordinators. [1]
Both systems posted the identical breakdown: seven points on each of problems one through five and zero on problem six, the same pattern reported for OpenAI's model as for Google's. [1][8]
Google DeepMind said afterward that it had been coordinating with the IMO organization in advance and held its announcement until after the closing ceremony out of respect for the IMO Board's request that AI labs wait until results were independently verified before publishing. [4]
Atlas interpretation: OpenAI's post went out while the closing-ceremony crowd was still on the Sunshine Coast, ahead of that verification for any lab, and reporting on the week described it as widely read as an attempt to claim the moment before Google's more deliberate, IMO-coordinated release could land. [4][5]
What Deep Think actually is
Google DeepMind describes Deep Think as a reasoning mode that explores several possible solution paths in parallel before settling on an answer, rather than committing to one chain of reasoning from the start. The version used for IMO 2025 was trained with reinforcement learning on multi-step reasoning and theorem-proving tasks plus a curated set of high-quality mathematical solutions, and it produced its proofs directly in natural language from the official problem text, without first translating the problems into a formal proof language. [1]
Atlas interpretation: That last detail is a change from Google DeepMind's 2024 systems, AlphaProof and AlphaGeometry, which won silver at that year's IMO but needed problems hand-translated into the Lean proof language and took one to two days of computation per problem. Working end to end in natural language, inside the same 4.5-hour sessions the students sat, is what DeepMind presents as the more significant part of the 2025 result, not simply the higher score. [1]
Why the grading process is the real story
Google DeepMind says its proofs were marked after the competition by the IMO's own coordinators, the same officials who graded the students that year, applying the IMO's grading guideline. OpenAI's proofs were marked differently: by the company's own account, three former IMO medalists it had engaged independently graded each problem, and a score was finalized only once all three agreed. OpenAI did not submit its answers through the IMO's coordination process, and reporting on the event noted that none of the IMO's 91 official coordinators took part in evaluating OpenAI's submission. [6][5]
Google DeepMind researcher Thang Luong put the distinction directly: the IMO's coordinators work from the organization's own grading guideline, and "any evaluation that's not based on that guideline could not make any claim about gold-medal level performance." [4]
Atlas interpretation: Neither grading process is being accused of dishonesty, and both used mathematicians fluent in Olympiad-style proof writing. The gap is independence and standing. Google DeepMind's score was assigned by the people who set that year's grading standard and applied it to the actual field of contestants, the same yardstick every human medal that week was measured against. OpenAI's score was assigned by three mathematicians the company itself selected, working outside that process. Whether OpenAI's model is in fact as capable as Google's is a separate question from what each company's claim can actually support. [6][5][4]
What the IMO itself said
The IMO's own comment on AI participation predates both companies' announcements. On July 19, the day of the closing ceremony and before OpenAI or Google had published anything, the organizers issued a statement noting that a group of AI companies had been invited to a fringe event and had privately tested closed-source models on that year's problems. IMO president Gregor Dolinar was quoted directly: "the IMO cannot validate the methods, including the amount of compute used or whether there was any human involvement, or whether the results can be reproduced. What we can say is that correct mathematical proofs, whether produced by the brightest students or AI models, are valid." [3]
Atlas interpretation: That caution was written for every AI claim made that week, Google's included. The IMO did not run an official AI category or award AI-specific medals, so neither company's result carries an IMO seal the way a student's medal does. What separated the two claims was not IMO sanction, since neither had it, but whether the grading ran through the coordinators and rulebook the IMO uses on its own contestants. Google's did. OpenAI's did not. [3][1][5]
Sources
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad
Google DeepMind · Jul 21, 2025
- OpenAI's IMO 2025 gold medal announcement (X thread)
Alexander Wei, via X (OpenAI) · Jul 19, 2025
- Final day of IMO 2025 (Closing Day Statement)
International Mathematical Olympiad 2025 (Sunshine Coast organizing committee) · Jul 19, 2025
- OpenAI and Google outdo the mathletes, but not each other
TechCrunch · Jul 21, 2025
- IMO Rebukes OpenAI for Self-Claimed Victory: "None of 91 Judges Participated in Scoring"
36Kr · Jul 21, 2025
- OpenAI's gold medal performance on the International Math Olympiad
Simon Willison (simonwillison.net) · Jul 19, 2025
- Mathematicians Question AI Performance at International Math Olympiad
Scientific American · Oct 7, 2025
- AI at IMO 2025: a round-up
Xena Project (Kevin Buzzard) · Aug 3, 2025