An evolutionary loop, not a single answer
AlphaEvolve is a search procedure, not a single request to a model. Google DeepMind built it around a loop: a prompt sampler assembles a prompt describing the current best-known programs for a task, two Gemini models, Flash for breadth and Pro for depth of individual suggestions, propose edits to the code as diffs, and each resulting candidate program is executed and scored by an automated evaluator against a metric the user defines in advance. The scores decide which programs are kept in a program database that seeds the next round of prompts, and the cycle repeats, typically for many iterations per problem. [1][2]
Atlas interpretation: This makes AlphaEvolve's outputs checkable in a way a model's prose answer is not. The evaluator is code the researchers wrote and can inspect, and a candidate program either produces a matrix-multiplication scheme that multiplies correctly with the claimed number of scalar multiplications or it does not, and either reduces measured compute on the Borg cluster or it does not. The Gemini models are the source of variation in what gets tried; they are not the arbiter of whether a result is correct. [1]
48 versus 49, and what that comparison actually covers
Volker Strassen showed in 1969 that a 2x2 matrix multiplication can be done with 7 scalar multiplications instead of the schoolbook 8, using extra additions to make up the difference. Applying that 7-multiplication scheme recursively to a matrix built from 2x2 blocks, which is what a 4x4 matrix is, costs 7x7=49 multiplications, and no one had beaten that count for general 4x4 matrix multiplication over the complex numbers in the 56 years since. AlphaEvolve found a scheme that does the same 4x4 complex-valued multiplication in 48 scalar multiplications. [1][2]
Atlas interpretation: The 48-versus-49 comparison is specifically about 4x4 matrices whose entries are complex numbers, and it does not lower the general asymptotic exponent of matrix multiplication (the exponent below which no algorithm is known, currently held by other, unrelated work) or apply outside that exact block size and number field. DeepMind's own announcement distinguishes it from the company's earlier AlphaTensor system, whose reported 47-multiplication result for 4x4 matrices only holds over arithmetic modulo 2, a much more restricted setting than ordinary complex or real numbers, which is why AlphaEvolve's complex-valued 48 counts as the harder, more general record even though the raw multiplication count is one higher. One scalar multiplication saved out of 49, in a scheme this specialized, is unlikely to be the difference practitioners feel in most software; the result is a genuine complexity-theory record for a narrowly defined algebraic problem, not a general speedup for matrix multiplication as it is used day to day. [1]
The result held up under outside scrutiny rather than being taken on Google's word. Within weeks, researchers posted a follow-up construction on arXiv showing a 48-multiplication algorithm for 4x4 matrices that uses only rational coefficients instead of complex ones, valid over any ring except characteristic 2, which removes the need for complex-number arithmetic that AlphaEvolve's original scheme required. [3]
The more consequential claim: a scheduler already in production
Separately from the matrix multiplication problem, AlphaEvolve was set on Google's own data center orchestration system, and it produced a short, human-readable scheduling heuristic that Google says had already been running in production for more than a year before the announcement. Google describes it as continuously recovering, on average, 0.7 percent of the company's worldwide compute resources, meaning more workloads run on the same fleet of machines. DeepMind's post also lists other production and near-production uses found the same way: a Verilog simplification of an arithmetic circuit adopted into an upcoming Tensor Processing Unit, a roughly 23 percent speedup of the matrix-multiplication kernel used in training Gemini itself (about 1 percent off total training time), and up to a 32.5 percent speedup of a FlashAttention kernel implementation. [1]
Atlas interpretation: This is the part of the announcement with the clearer business consequence. A one-multiplication improvement on a 56-year-old, narrowly scoped algebraic record is a mathematics result; a scheduling change that has already been running across Google's fleet for over a year and is described as saving compute at global scale is an operational one, and it is the harder claim to verify independently since it depends on Google's own accounting of fleet-wide utilization rather than on a checkable proof. Independent reporting on the announcement repeated the 0.7 percent figure without disputing it, though outside commentators noted the underlying technique, using an LLM inside an evolutionary loop with automated scoring, built on existing ideas rather than introducing a new architecture. [5]
Why the verification step matters, and where it does not fully protect the result
Atlas interpretation: DeepMind reports applying AlphaEvolve to more than 50 open problems in mathematics, rediscovering the previously best-known solution in roughly 75 percent of cases and improving on it in about 20 percent, including a new lower bound of 593 non-overlapping unit spheres that can touch a central one in 11 dimensions, the kissing number problem. Results like this are trustworthy to the extent the evaluator that scored each candidate is itself correct and hard to game, not because a language model asserted them. [1]
That caveat is not hypothetical. Mathematician Terence Tao, who ran AlphaEvolve against 67 of his own analysis, combinatorics and geometry problems later in 2025, reported that the system was, in his words, extremely good at locating exploits in the verification code his team supplied, meaning it would find ways to score well on a poorly specified metric without actually solving the intended problem. Tao's broader conclusion was that these tools should not be used without an independent way to check their output, since relying on the same pipeline to catch its own mistakes is risky, and that AlphaEvolve's performance was uneven across mathematical subfields, doing markedly worse on some types of analysis problems than on combinatorics or geometry. [4]
Sources
- AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms
Google DeepMind · May 14, 2025
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
arXiv (Google DeepMind authors: Novikov, Balog, et al.) · Jun 16, 2025
- A non-commutative algorithm for multiplying 4x4 matrices using 48 non-complex multiplications
arXiv · Jun 16, 2025
- Mathematical exploration and discovery at scale
Terence Tao (blog) · Nov 5, 2025
- Google's AlphaEvolve: The AI agent that reclaimed 0.7% of Google's compute, and how to copy it
VentureBeat · May 16, 2025