Six agents and a tournament
Google built the AI co-scientist on Gemini 2.0 as seven cooperating agents rather than one model answering a prompt. A Supervisor agent reads the scientist's research goal, breaks it into a plan, and assigns work to six specialists. Generation drafts candidate hypotheses by running literature searches and simulating a debate among several expert personas. Reflection acts as peer reviewer: an initial pass, then a full review with its own literature search, then a "deep verification" pass that breaks a hypothesis into its individual assumptions and checks each one. Proximity clusters hypotheses that are really the same idea so the tournament stops wasting matches on near-duplicates. Evolution writes improved variants inspired by the current leaders, as new entries rather than edits, so a bad rewrite cannot silently replace a good hypothesis. Meta-review reads across every review and debate transcript, distills a running critique that gets appended to the other agents' prompts on the next cycle, and drafts the research overview a scientist actually reads. [1][2]
Ranking is where the system spends most of its compute. New hypotheses enter a continuous tournament at an Elo rating of 1200; top-rated hypotheses fight in multi-turn simulated debates while lower-rated ones get a single comparison, and the agent preferentially matches similar ideas so a score reflects a real head-to-head rather than an easy pairing. The technical report validates the score itself against ground truth: run on GPQA Diamond's graduate-level science questions, the hypothesis with the highest Elo rating matched the benchmark's correct answer 78.4 percent of the time. On a set of 15 research goals picked by outside experts, the same tournament ranked the co-scientist's own hypotheses above answers from Gemini 2.0, OpenAI's o1, and DeepSeek R1, and the experts themselves rated its output 2.36 out of 4 on average (lower is better), with novelty and impact scored 3.64 and 3.09 out of 5. [2]
A question nobody had published an answer to
The clearest test came from Jose Penades and Tiago Costa's lab at Imperial College London, which studies how bacteria trade genes. Capsid-forming phage-inducible chromosomal islands, cf-PICIs, are parasitic DNA elements that build their own protein capsids but have no tail of their own, so on paper they should have no way to inject their DNA into a new bacterial cell. The lab had spent over a decade trying to explain how cf-PICIs still turn up across unrelated bacterial species anyway. They gave the AI co-scientist only the question and the published literature, none of their own unpublished results, and asked it to explain the mechanism. [3][6]
The system's top-ranked hypothesis was that cf-PICIs pirate tails from unrelated phages already infecting the same cell, building a hybrid particle that can carry the cf-PICI's DNA into a different species. That was the mechanism the lab had already confirmed at the bench. At the time the AI was tested, the paper describing that confirmation was under confidential review at a journal and nothing about it was public, so the match could not have come from the model having read an early draft or a conference abstract. Google's blog post and the lab's own preprint went up the same day, February 19, 2025. The formal, peer-reviewed version followed seven months later, published in Cell that September, describing the mechanism as tail piracy: a cf-PICI capsid binding whichever phage tail happens to be available and using it to cross into an unrelated species, carrying resistance and virulence genes along with it. [3][1][4]
Connecting dots, not discovering them
Atlas interpretation: Penades was candid about what the match meant and what it didn't. "It's very frustrating because we have the answer up there, and we didn't see it," he said of rereading his own field's literature after seeing the AI's hypothesis. Asked whether the system understood the mechanism it had proposed, he was blunter: it "don't really understand what this proposition, it's just connecting the dots with makes sense." The lab's own explanation for why a language model got there first was not superior reasoning. It was that the model carried none of the working assumption, that every particle a cf-PICI releases must be infectious, that had kept the lab from considering a tailless capsid binding a stray phage tail. [6]
Atlas interpretation: That distinction matters for how much weight the case can carry. The system was not shown to generate an idea absent from the literature; it was shown to recombine published facts into a specific mechanism that a well-informed team, reading the same papers, had not yet put together, and to do it once, on a question whose answer a co-author already knew. A single blind match against one lab's unpublished result is real evidence that the tournament process can surface a good hypothesis. It is not evidence that the system can be pointed at an arbitrary open problem and repeat the trick, and Google's own writeup does not claim otherwise. [6][5]
The other two experiments, and the pushback
Two other validations ran alongside the cf-PICI case. Set loose on 2,300 approved drugs across 33 cancer types for acute myeloid leukemia, the system produced 78 repurposing hypotheses, written up as NIH Specific Aims pages and scored by six board-certified hematologists and oncologists; one candidate, KIRA6, went on to inhibit a leukemia cell line in vitro at a clinically relevant concentration. At Stanford, a separate team asked it for new drug targets for liver fibrosis; the system proposed epigenetic targets that, tested in human hepatic organoids, showed measurable anti-fibrotic activity. [1][2]
Neither result held up as well as the cf-PICI case once specialists outside Google looked at it. Steven O'Reilly, a biotech researcher at the UK firm Alcyomics, dismissed the liver fibrosis finding outright: "The drugs identified are all well established to be antifibrotic. There is nothing new here." Pathologist Favia Dubyk made a similar point about the leukemia hypotheses: that Google's public description was too thin to judge, and that no legitimate scientist would call the claim credible without seeing the underlying data. [5]
Where this sits in the AI for science story
Atlas interpretation: The blog post and preprint route Google used in February 2025, publish first and let a journal catch up later, is common for AI lab announcements and unusual for biology. It took until May 2026 for the underlying work to clear peer review, appearing in Nature with an expanded author list spanning Google DeepMind, Google Cloud AI Research, and Stanford University School of Medicine, and with more specific numbers than the original announcement carried: five initial AML candidates narrowed to three, binimetinib, pacritinib, and cerivastatin, that measurably slowed cell growth, and two of three proposed liver fibrosis targets showing anti-fibrotic activity at statistical significance. [7]
Atlas interpretation: Read together, the two threads support a narrower claim than "AI discovers science." The tournament architecture is a reasonable way to turn a large language model's literature knowledge into ranked, falsifiable hypotheses, and the cf-PICI case shows it can land on a specific, correct, and previously unpublished mechanism without being told the answer. Whether it can do that reliably, on problems where no one already knows the answer, is exactly what the drug-repurposing critiques and Penades's own qualifications leave open. The 2025 wave of AI-for-science claims, of which this was one of the first and most carefully validated, mostly rested on that same open question. [5][6]
Sources
- Accelerating scientific breakthroughs with an AI co-scientist
Google Research · Feb 19, 2025
- Towards an AI co-scientist
Google Research · Feb 26, 2025
- AI mirrors experimental science to uncover a novel mechanism of gene transfer crucial to bacterial evolution
bioRxiv · Feb 19, 2025
- Tail-Swapping "Pirate" Phages Expose New Route for AMR
Genetic Engineering & Biotechnology News · Sep 11, 2025
- Scientists Say Google's "AI Scientist" Is Dead on Arrival
Futurism · Mar 5, 2025
- What did Google's AI Co-Scientist "Discover"? The Human Scientists' POV, from the Podovirus podcast
The Cognitive Revolution · Jan 21, 2026
- Accelerating scientific discovery with Co-Scientist
Nature · May 19, 2026