AI News Accuracy Study: Errors Across Four Assistants

Journalists testing ChatGPT, Gemini, Copilot and Perplexity found significant issues in 45% of more than 3,000 news answers, with sourcing the largest problem.

How the study was built

The European Broadcasting Union and the BBC ran the study across 22 public service media organizations in 18 countries, working in 14 languages, testing the assistants on real news questions between late May and early June 2025. Professional journalists at each organization scored more than 3,000 of the resulting answers against four criteria: accuracy, sourcing, whether the answer kept opinion separate from fact, and whether it gave enough context to understand the story. The four assistants tested were ChatGPT, Copilot, Gemini and Perplexity. [2][3][1]

Atlas interpretation: That scale is the point of the exercise. A single newsroom testing one assistant in one language cannot tell you whether a failure is a quirk of that newsroom's archive or a property of the model, and the EBU built this study specifically to answer that question across markets and languages at once. [3]

What counted as a significant issue

Forty-five percent of the answers evaluated had at least one issue the journalists rated significant enough to mislead a reader, and 81 percent had some problem once smaller errors were counted too. Sourcing was the largest single category at 31 percent, ahead of accuracy problems (hallucinated or outdated details) at 20 percent and missing context at 14 percent. Some answers also presented opinion or satirical content as settled fact. The problems were concrete: Perplexity told users surrogacy is illegal in the Czech Republic when it is not, ChatGPT described Pope Francis as the sitting pope months after his death, and Gemini claimed NASA astronauts had never been stranded in space despite two crew members having just spent nine months on the International Space Station. [1][3][4]

What a sourcing failure actually looked like

"Sourcing" problems meant missing citations, statements attributed to a source that never said them, and what one report on the study called "ceremonial citations": references that look credible but do not actually support the claim they are attached to. Broken down by assistant, sourcing failures were far from evenly spread. Google's Gemini had a sourcing problem in 72 percent of its answers, against 24 percent for OpenAI's ChatGPT and 15 percent each for Perplexity and Copilot. Gemini's answers had a significant issue of any kind 76 percent of the time, well above the 45 percent average across all four assistants, and the study attributes most of that gap to its sourcing. [4][5]

Atlas interpretation: The summary's figure of "76 percent" for Gemini's sourcing is close but conflates two separate numbers the study reports: 76 percent is Gemini's overall significant-issue rate across all four criteria, while 72 percent is the sourcing-specific figure. Both point the same direction, Gemini trailed the other three assistants by a wide margin, but the sourcing failure rate itself is 72, not 76. [4]

A second reading, not a first one

This was not the first time the BBC had measured this. In February 2025 the BBC alone tested the same four assistants on 100 questions built around its own reporting, asking each to cite BBC articles as sources, with the testing done in December 2024. Journalists found significant issues in 51 percent of responses and some kind of problem in 91 percent; 19 percent of answers citing BBC content contained factual errors such as wrong dates or statistics, and 13 percent of quotes attributed to BBC articles were altered or fabricated outright. Broken down by assistant, Gemini already had the worst error rate at 34 percent, ahead of Copilot at 27 percent, Perplexity at 17 percent and ChatGPT at 15 percent. [6][7]

Atlas interpretation: The October 2025 study exists to test whether the February findings were an artifact of BBC content and the English language specifically, or a property of how these assistants handle news generally. Testing 14 languages and 22 outlets and landing on broadly the same picture, with Gemini again the weakest on sourcing, answers that question: the problem travels across languages and newsrooms rather than sitting in one archive or one market. [3][6]

What the EBU asked for, and who answered

Alongside the report, the EBU and BBC published a "News Integrity in AI Assistants Toolkit" laying out what a good AI news answer looks like on each of the four criteria and a taxonomy of the ways answers go wrong, aimed at both AI developers and newsrooms. The EBU said it is pressing EU and national regulators to enforce existing information integrity and digital services laws against AI assistants, and said it wants formal dialogue with the technology companies involved. OpenAI, Google, Microsoft and Perplexity did not respond to requests for comment on the study by the time the initial coverage ran. [2][5][1]

Sources

  1. AI models misrepresent news events nearly half the time, study says

    Al Jazeera · Oct 22, 2025

  2. Largest study of its kind shows AI assistants misrepresent news content 45% of the time – regardless of language or territory

    European Broadcasting Union · Aug 20, 2026

  3. News Integrity in AI Assistants

    European Broadcasting Union · Oct 22, 2025

  4. AI chatbots flub news nearly half the time, BBC study finds

    The Register · Oct 24, 2025

  5. AI Assistants Get News Wrong in 45% of Cases, Landmark BBC/EBU Study Finds

    WinBuzzer · Oct 22, 2025

  6. AI assistants error prone when it comes to news

    Digital Content Next · Feb 24, 2025

  7. BBC AI Accuracy Study

    The Media Copilot (Substack) · Sep 8, 2026