AI Chatbots Fail as Fact-Checkers: Full Fact Trial Reveals Critical Errors in Major Language Models
In a comprehensive five-month investigation, Full Fact, the UK’s leading fact-checking organization, has uncovered significant flaws in how major AI chatbots respond to misinformation circulating online. The trial, which began in February, systematically tested three Gemini models, Grok, and ChatGPT against claims that were being fact-checked by the organization, revealing 39 major AI errors in the first five months alone. The findings paint a troubling picture of the reliability of large language models (LLMs) as tools for verifying information, particularly during breaking news situations. Perhaps most concerning is that even when chatbots initially provided incorrect information, they often failed to correct themselves even after fact-checking articles were published, sometimes introducing new errors while attempting to fix old ones.
The most striking examples of AI failure came in the realm of visual content identification. Several models incorrectly authenticated an AI-generated image of a cabin window on the Ryanair flight where a passenger was nearly sucked out of a window, while others misattributed pictures and footage to Iran or Israel when they were from entirely different locations. Grok, X’s AI assistant, falsely claimed that an AI-generated image of Jewish charity ambulances on fire showed a real scene from the March arson attack in Golders Green. In another instance, one Gemini model fabricated an encounter in which Katie Hopkins had “unleashed hell” in the House of Commons by confronting Muslim MPs, an event that never occurred. The models also showed particular weakness in identifying AI-generated political content, with several falsely claiming that AI-generated political banners and fake council posters were authentic official communications.
The research methodology was carefully designed to simulate real-world usage, with reporters drafting neutrally phrased questions that a typical reader might reasonably ask an AI assistant. These questions were then posed to multiple AI systems, including Gemini 2.5 Flash, 2.5 Pro, and 3.1 Pro through Google’s API service, as well as ChatGPT 5 and Grok 4 through their respective API services. The results were then compared against verified facts. This approach revealed not just errors in identifying AI-generated content, but also fundamental failures in recognizing real footage. For instance, both Grok and ChatGPT incorrectly identified a video of a fire near Glasgow Central Station as showing an Iranian missile attack on Tel Aviv, while Grok claimed a 2022 video of a fire in Saudi Arabia showed footage of Tel Aviv during Iran’s missile attacks on Israel.
The second phase of the trial examined whether publishing fact-checking articles would help the AI systems improve their responses. The results were mixed and often disappointing. While some models did correct their initial errors by referencing Full Fact’s published articles, others continued to provide incorrect information despite having access to accurate sources. Two Gemini models persistently maintained that an AI-generated image of Earth from the Artemis II mission was real, even after fact checks debunking this claim had been published. In perhaps the most concerning example of AI unreliability, Grok initially corrected its claim about New York Mayor Zohran Mamdani’s comments about the UK, but then introduced a new error by falsely stating that Mamdani was not the mayor of New York at all.
The tech companies’ responses to these findings varied significantly. OpenAI acknowledged the concerns raised by Full Fact, emphasizing that factual accuracy remains an important focus for the company, though they noted that the API testing method used by Full Fact did not represent the typical consumer ChatGPT experience. Google took a more defensive stance, arguing that the study had accessed “out-of-date Gemini models through a developer channel that isn’t representative of how most people use AI,” and claimed that when they retested the same claims on the Gemini website, they were consistently debunked. X, formerly Twitter, did not respond to requests for comment about Grok’s errors. These responses highlight a fundamental challenge in evaluating AI systems: the rapid iteration of models means that test results may quickly become outdated, while the API-based testing methods used by researchers may not reflect the user experience of those accessing chatbots through mainstream websites and applications.
The implications of these findings extend far beyond individual errors to question the broader role of AI chatbots in information dissemination and verification. As people increasingly turn to LLMs as search engines and fact-checking tools, the trial demonstrates that these systems cannot be relied upon as foolproof means of verifying claims. The errors identified in this trial were not subtle or borderline cases; they included gross misidentifications of real-world events, fabrications of nonexistent incidents, and confident assertions of falsehoods. Full Fact’s assessment is clear: while AI may serve as a useful starting point for research, it cannot replace the role of fact checkers or human judgment in general. The organization recommends that anyone attempting to verify information—whether images, videos, or statistics—should look beyond LLMs to credible, primary sources of information that can be triangulated to provide a wider, more accurate picture.
This research is part of a larger ongoing effort by Full Fact, which has previously documented instances of AI systems providing incorrect information. The organization had earlier reported on Grok misidentifying a fire in Glasgow as an incident in Tel Aviv and falsely claiming that a video shared by the Metropolitan Police of a “Unite the Kingdom” rally was actually footage from a 2020 anti-lockdown protest. As these AI tools become increasingly embedded in everyday life, the quality and reliability of what they produce becomes a fundamental matter of public interest. Yet there is currently no independent, public interest mechanism to systematically evaluate whether these AI services are accurate, transparent, timely, consistent, or responsible in the information they provide. To address this gap, Full Fact is developing a benchmark to evaluate LLM performance, which would provide a much-needed framework for independently assessing the reliability of AI systems that millions of people rely on for accurate information. The findings of this trial underscore the urgent need for such oversight as AI continues to shape how people access and verify information in the digital age.



