AI Misinformation Is Mostly Self-Inflicted, New Research Finds: Outdated PDFs and Content Debt Are Fueling False AI Outputs
The rapid adoption of artificial intelligence across the global business landscape has created a paradox: companies are investing heavily in AI capabilities while simultaneously becoming more exposed to the technology’s ability to spread false or misleading information. According to Bain research, 74% of companies now rank AI among their top three strategic priorities, but the enthusiasm is tempered by rising concern about AI-driven misinformation. Gallagher’s 2026 AI Adoption and Risk Survey found that “AI errors, misinformation and hallucinations” was the leading perceived threat from AI, cited by 57% of respondents. However, a crucial new layer of understanding has emerged: much of this misinformation is not the fault of the AI systems themselves, but rather the result of organizations feeding those systems their own outdated, contradictory, and poorly structured digital content. New research across 187 AI agents and 360,000 verification checks has revealed that 79% of AI misinformation traces back to an organization’s own outdated content, not to AI hallucination. This finding reframes the conversation around AI risk, shifting the blame from machine error to a decades-long failure of content management and governance. The reality is stark: brands are unintentionally arming AI tools with stale, incorrect, or non-compliant information and then wondering why their AI visibility is poor or why AI-generated answers misrepresent them.
The underlying issue is what experts call “content debt” — the accumulation of outdated, duplicated, and forgotten documents that live on company websites, often for years, without any clear owner or governance. For most of the internet era, this hidden estate sprawl was a manageable nuisance because human users rarely stumbled upon these forgotten files. Search engines might index them, but they seldom ranked highly enough to attract attention. AI systems, however, retrieve information differently. They pull from the entire digital ecosystem, including the long tail of neglected documents, and treat that content as a valid source of truth. “AI is exposing a content management problem that has existed for years,” said Dipo Ajose-Coker, head of marketing for Content Technology at RWS Group. “Organizations could get away with hundreds or thousands of forgotten documents because humans rarely found them.” Now, AI retrieval makes that long tail much more accessible, meaning content debt that was previously hidden is suddenly visible to customers, regulators, and competitors. The scope of the problem is broad. Organizations often have multiple copies, formats, repositories, and owners for essentially the same information, with no reliable mechanism for determining what is current, authoritative, or obsolete. PDFs play an outsized role in this dysfunction. They are often treated as mere attachments to a website, yet they represent a significant information estate in their own right. Annual reports, product guides, contracts, brochures, policies, and financial filings are all PDFs that can remain online for years for compliance, reference, or simply because no one remembers to remove them. Updating a PDF is frequently overlooked because it is too expensive, managed by a different team than the website team, handled by an external partner without internal access to source designs, or simply less visible in a content management system than a web page. As a result, PDFs become time capsules of outdated facts, and AI tools amplify them as though they were current truth.
The danger of outdated PDFs is compounded by the fact that AI systems appear to place disproportionate value on PDFs compared to other forms of content. Research across the top 10 AI tools has found that a PDF is up to 2.5 times more likely than a web page to be valued when AI systems construct answers. This is no accident. PDFs are the default format for exactly the kinds of content that an AI system has good reason to regard as authoritative: annual reports, regulatory filings, technical specifications, corporate policies, formal statements, and instructional documents. The format itself has become a signal of trustworthiness to AI systems. “There is an interesting irony here,” Ajose-Coker noted. “Humans often think of PDFs as static attachments, while an AI retrieval system may encounter them as some of the richest and most authoritative-looking sources on an organization’s digital estate.” For a human, a PDF is a finished document; for an AI, it is evidence. The result is that any error, misstatement, or outdated fact contained within a PDF carries outsized influence in AI-generated answers. A PDF with a wrong financial figure, an obsolete policy, or an incorrect product specification will be cited with confidence, and the AI will not know that newer, correct information exists elsewhere on the same company’s website. This means the stakes for PDF accuracy are far higher than for ordinary web pages, and organizations that fail to recognize this are leaving themselves exposed to significant regulatory, commercial, and reputational risk.
Compounding the problem of outdated content is the fact that many PDFs are structurally unreadable by AI systems, even when the information they contain is correct. The way a PDF is designed for human consumption can create serious obstacles for machine interpretation. Real-world examples from annual reports published by UK FTSE 100 companies in regulated industries illustrate the issue vividly. In one annual report, the PDF had been flattened to make it visually appealing, a process that merges individual objects into a single layer and strips away the structural tags that AI relies on to understand the document. The result was a misalignment between the heading structure of a table and the financial data displayed, causing critical data to be mislabeled from the AI’s perspective and financial information to be misquoted. In another report, the design team introduced a new font without embedding it in the document, so the AI could not read the characters and substituted its own, producing gibberish that was then presented as fact. A third PDF was so large that the AI was cut off from sections of the report, while Google Gemini could not access the file at all, needlessly reducing the visibility of high-profile information on one of the most popular AI tools. These are not edge cases; they are examples from respected companies with substantial design and compliance budgets. “A PDF can look perfectly clear to a human while having poor reading order, missing structure, inaccessible tables, flattened content, or other technical issues that make its meaning much harder for a machine to interpret reliably,” Ajose-Coker said. Organizations that design PDFs solely for the human eye are inadvertently creating documents that AI will misread, ignore, or amplify incorrectly.
The good news is that this problem is remediable. The recommended approach is a phased remediation plan that begins with a comprehensive inventory and risk assessment. “Start with an inventory and risk assessment,” advised Ajose-Coker. “Identify what is public, what is still authoritative, what is duplicated, what is obsolete, and what carries the greatest reputational, regulatory, or commercial risk.” This first step is essential because it establishes a baseline, identifies the PDFs that pose the greatest threat, and allows organizations to measure progress. AI-powered tools can assist this process by scanning digital estates for structural weaknesses, outdated indicators, and duplication. After the assessment, actions must be prioritized. No organization can remediate thousands of PDFs at once, so it is critical to focus on the top percentage of documents that are most likely to cause harm or present the greatest opportunity for AI visibility. The next step is to take a hybrid approach to the actual remediation work. Some PDFs can be fixed with automated tools that add structure, embed fonts, and correct tagging; others will require manual intervention, especially where complex layouts, sensitive content, or regulatory constraints are involved. Manual remediation is time-consuming, resource-intensive, and expensive, so the choice between automation and manual effort must be driven by document type, level of risk, and the nature of the issues found. Finally, the execution phase brings it all together. Timelines naturally vary, but the research suggests that significant progress is achievable quickly. Within 30 days, an organization can have a more accurate catalogue of its PDFs and a clearer understanding of the scale and distribution of its digital estate. Within 90 days, it can have documents available in structured form for AI systems, an inventory with remediation mapped out, and the knowledge needed to ensure new PDFs are designed to be read correctly by both humans and machines.
Once the initial remediation is complete, the harder question becomes how to prevent the problem from recurring. The answer lies in governance and a fundamental shift in mindset. Organizations need to stop treating PDFs as isolated attachments to their websites and instead treat the underlying information as a structured, governed asset. “Put governance in place,” Ajose-Coker said. “Organizations need to get control of the information behind them.” This means managing important content as structured, governed data with clear ownership, metadata, versioning, and lifecycle rules, and then publishing that content into PDF, HTML, and other channels from a single controlled source. By doing so, organizations ensure that the PDF version of a policy, report, or product guide is always derived from the same current, authoritative source as the web page version, and that outdated versions are systematically retired or marked as obsolete. This also addresses the issue of “content debt” reaccumulating after an initial cleanup, because governance prevents the careless duplication and abandonment of documents that created the problem in the first place. Achieving this state requires more than new tools; it requires a cultural change that elevates content management to a strategic priority, with ongoing measurement and monitoring of the digital estate, and potentially more senior ownership of the risks associated with AI misinformation. The issue is urgent. As more users rely on AI services like ChatGPT for information discovery, leading brands are not ready to increase AI visibility and reduce the chance of AI misinformation. The research is clear that AI misinformation is not primarily an AI problem — it is a content problem. Organizations that act now can protect their reputations, comply with regulatory expectations, and actually benefit from AI’s ability to amplify their brand; those that delay will continue to watch as AI systems confidently repeat their own outdated, broken, and forgotten content as though it were fact. The path forward is not complicated, but it requires discipline, investment, and a willingness to admit that the source of the problem has been sitting on the company’s own server all along.


