Independent AI Benchmark Launched by Full Fact as Chatbots Become Primary News Source

People are increasingly turning to AI chatbots, based on large language models, for information that shapes their lives. The scale of that shift is difficult to overstate. Google, for example, says its AI Overviews feature, now the default in most Google searches, has 2.5 billion users a month. That figure is not merely a commercial statistic; it represents a moment in which algorithmic synthesis has replaced the traditional web of links, headlines, and source attribution for a substantial portion of public information retrieval. Where a search query once produced a list of sources that users could mentally rank, it now often produces a single conversational answer, written by a machine, with no visible chain of evidence. As these tools become embedded in everyday life, the quality and reliability of their outputs become a matter of public interest on par with the accuracy of broadcast media or government statistics. The risk is not that every answer must be perfect; it is that errors can be spread at machine speed and on a scale that overwhelms correction. Yet unlike broadcast media, AI systems are neither regulated nor independently audited. Their underlying models are proprietary, their training data opaque, and their performance claims frequently self-assessed. In this atmosphere, users are being asked to make decisions about health, politics, finance, and personal relationships based on machines whose operations they do not understand and whose failures are often invisible. More must be done, in the words of Full Fact, to understand and address the potential harms caused by AI-generated misinformation, so that societies can make informed decisions about when to rely on AI, and public trust can be grounded on evidence rather than marketing claims.

Full Fact, the UK-based fact-checking charity, has already witnessed some of these harms firsthand. Earlier this year, the organisation conducted a trial in which leading AI models made dozens of major errors when assessing false claims. These were not subtle misstatements or minor wording problems. Models endorsed falsehoods, produced confident but misleading explanations, and failed to identify claims that a skilled human fact-checker would immediately recognise as fabricated. In some cases, the models appear to have generated new falsehoods while attempting to debunk established ones—a dangerous outcome in an environment where people increasingly ask AI assistants to check political speeches, viral social media posts, and news headlines. For a fact-checking organisation, this trial reinforced a troubling conclusion: the tools that millions of people now use as authoritative sources of information are not safe to trust by default. The problem is not simply that individual answers are wrong; it is that the pattern of error is difficult for ordinary people to predict. A chatbot may answer perfectly on one day and then fail badly on a similar question the next day, with no outward sign that its confidence is misplaced. This unpredictable unreliability creates a particularly insidious form of misinformation, because it resembles expertise. Full Fact believes that far more research is needed, not only into how AI models generate false information, but also into how that information moves through social networks, search engines, and messaging apps, shaping the beliefs and decisions of millions of people.

In response to this urgent problem, Full Fact is building a publicly available, auditable benchmark to independently evaluate the output of leading AI models. The methodology is both simple and rigorous: the organisation asks the same set of questions every day and records the responses from each major AI model. These questions are designed to test how models handle claims that are true, false, partially true, and contested. They also test how models deal with topics where information changes rapidly, such as public health guidance, legislative developments, and election administration. After collecting the responses, Full Fact will annotate and analyse them on a regular basis, using the same editorial standards applied to human claims. This repeated, daily testing is central to the design. A single snapshot of an AI model’s performance tells us little; a running record over weeks and months reveals patterns of drift, improvement, regression, and inconsistency. The benchmark will be published in full, with methodologies clearly explained, so that other researchers, journalists, and members of the public can verify the findings and even replicate the tests.

The benchmark will assess AI models against five key dimensions. The first is factuality. This asks whether responses contain verifiably accurate claims, avoid hallucination, and correctly represent the evidence base. Factuality is not merely about catching outright falsehoods; it also means checking whether a model correctly handles nuance, uncertainty, and conflicting evidence. A model that accurately repeats a well-established fact but ignores recent contrary evidence is still failing the factual standard. The second dimension is transparency. This examines whether the language model communicates uncertainty, cites or attributes high-quality sources, acknowledges limitations, and distinguishes fact from opinion. A good answer should not make the user hunt for the source of a claim; it should be clear about what is known, what is not known, and whether a particular statement is an objective finding or a value judgment. Transparency is especially important in civic contexts, where a single unsupported assertion can subtly shape the public debate around an election, a referendum, or a public health emergency.

The third dimension is timeliness. Large language models are often trained on data that becomes out of date, and they may not know when their knowledge is stale. Full Fact’s benchmark will test whether responses reflect current information rather than outdated data, and whether the model recognises the limits of its own temporal knowledge. A model that confidently describes a situation that no longer exists is not simply being impolite; it is actively misleading users who have no reason to suspect that the information is old. The fourth dimension is consistency. This is an especially powerful test because it exposes the difference between knowledge and mimicry. Does the language model give materially the same answer to the same question over time and when asked in different ways? If a model gives contradictory answers to the same prompt on different days, or if it can be pushed into reversing a factual position simply by rewording the question, then its responses lack the stability we demand from reliable information sources. Consistency is not about robotic repetition; it is about predictability and trustworthiness. A source that says one thing today and something else tomorrow, with no new evidence to justify the change, cannot be called reliable.

The fifth and final dimension is civic responsibility. This is an attempt to grapple with the special duty that AI systems now bear as mass information providers. Full Fact will ask whether information about democratically important questions is balanced, whether the model does not amplify misinformation, and whether it supports informed participation in public life. This dimension goes beyond fact-checking individual sentences. It asks how a model handles topics where there are legitimate disagreements, how it represents the views of different communities, and whether its outputs encourage people to engage with democratic processes in an informed manner. An AI that subtly dismisses voting, that presents fringe opinions as consensus, or that treats a civil rights issue as a simple dispute is failing its civic duty even if every sentence is literally true. The benchmark will apply the same rigorous editorial standards that Full Fact uses when assessing claims made by public figures and reported in the media. By systematically testing these models over time, Full Fact will expose inconsistencies and provide the general public with an independent assessment on which AI tools they can trust)Skip? Need continue. Transparency is at the heart of the project. The results will be published in full, to be used by both individuals who use AI to source information in their everyday lives, and institutions and professionals who hold technology companies accountable for producing good information. This is a critical point: the project is not designed to be a one-off report or a consumer product ranking. It is a permanent, auditable archive of model behaviour, a kind of public interest testing laboratory for the age of automatic information.

The first report from this project will be released in mid-October, and Full Fact has promised that further reports will follow. The phrase “watch this space” is not a casual sign-off; it is a signal that the organisation intends to make this benchmark a long-term commitmentiors. As AI plays a greater role in medicine, education, journalism, and government, the need for independent verification will only grow. The stakes are particularly high because errors are not distributed evenly. An AI system that is 99 percent accurate in general may be far less accurate when asked about newly emerging events, such as a war, a pandemic, or a contested election. In those high-stakes moments, even a small rate of failure can cause immense harm, especially when false outputs are shared by users who believe they are spreading verified information. Full Fact’s initiative rests on a simple but powerful idea: public trust in AI should be based on evidence, not on the assurances of technology companies. The days when search engines could be treated as neutral conduits to human knowledge are over. These systems are now authors, editors, and broadcasters in their own rightcars, and they must be held to the same standards of accuracy, fairness, and transparency that we demand from established sources of public information. In launching this benchmark, Full Fact is taking a crucial step toward that goal, and the results will matter to every person who has ever accepted an AI answer at face value.

Share.
Leave A Reply

Exit mobile version