Close Menu
DISADISA
  • Home
  • News
  • Social Media
  • Disinformation
  • Fake Information
  • Social Media Impact
Trending Now

Experts to Examine Social Media’s Effects on Adolescent Brain Development Amid Ongoing Meta Litigation

August 29, 2026

Weekly Review: Misinformation on Nepal Flash Floods, Rahul Gandhi, and Other Key Issues

August 29, 2026

Experts to Examine Social Media’s Impact on Developing Brains Amid Ongoing Meta Litigation.

August 29, 2026
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram YouTube
DISADISA
Newsletter
  • Home
  • News
  • Social Media
  • Disinformation
  • Fake Information
  • Social Media Impact
DISADISA
Home»Disinformation»A New Method for Detecting AI-Generated Disinformation Amid Growing Sophistication
Disinformation

A New Method for Detecting AI-Generated Disinformation Amid Growing Sophistication

Press RoomBy Press RoomAugust 29, 2026No Comments
Facebook Twitter Pinterest LinkedIn Tumblr Email

Generative artificial intelligence has made it easier than ever for malicious actors to flood social media with convincing disinformation, but a team of researchers has developed a new way to detect these attempts by focusing not just on what a post says, but on what it is trying to do to a conversation. In their latest research, Seán Roberts, a lecturer in linguistics at Cardiff University, and Kateryna Krykoniuk, a research associate at the University of Sheffield, argue that as AI-generated text becomes increasingly fluent and difficult to distinguish from human writing, the old methods of spotting disinformation by looking for grammatical mistakes or unusual word choices are becoming obsolete. Instead, they propose examining whether a comment is trying to derail the discussion, steering it away from its original topic and towards more polarising issues. The researchers analysed thousands of comments on BBC News YouTube videos, a platform that has previously been targeted by organised disinformation campaigns, and found that many attempts at manipulation rely on a rhetorical device known as a red herring. For example, a discussion about the cost of living can suddenly become an argument about immigration, or a conversation about the war in Ukraine can turn into claims about government corruption. These jarring shifts are often deliberate, designed to sow division, inflame political debate and undermine trust in reliable information. The study’s core insight is that instead of trying to identify whether a post was written by AI, detection systems should ask a different question: is the message trying to derail the conversation? This approach, they argue, offers a more robust way to spot manipulation in an era where generative AI has removed many of the linguistic fingerprints that once made automated accounts easy to identify.

Until recently, identifying malicious accounts on social media was often quite straightforward. Many disinformation campaigns relied on people writing in a second language, and their posts frequently contained grammatical mistakes, awkward phrasing or unusual word choices. Detection systems could look for these patterns in the language used, flagging suspicious accounts based on stylistic inconsistencies. But generative AI has changed that landscape dramatically. Modern AI systems can now produce fluent, natural-sounding text that is much harder to distinguish from human writing, and the linguistic signals that researchers once used are rapidly losing their usefulness. For example, patterns like the use of em-dashes and the word “delve” used to be considered telltale signs of a text being generated by AI, but AI systems are constantly adapting and these older markers are increasingly ineffective. Roberts and Krykoniuk argue that trying to detect AI purely from the words people use is becoming a losing battle. A phrase or punctuation style can be changed, and as AI models improve, the text they generate becomes more varied and humanlike. The better approach, they believe, is to look at what a message is trying to achieve. Disinformation often works not by stating a single falsehood, but by manipulating the shape of a discussion, redirecting attention away from inconvenient topics and towards emotionally charged, divisive ones. By shifting the analytical focus from word-level features to the broader dynamics of conversation, the researchers hope to build a detection system that can keep pace with the rapid evolution of generative AI. This is a significant departure from conventional methods, which have largely focused on identifying whether a specific post was written by a human or a machine, an increasingly difficult and perhaps ultimately impossible task as AI models become more sophisticated.

To build their system, the researchers manually analysed more than 1,600 comments posted beneath BBC News videos on YouTube, labelling them according to 25 different features of online discussion. This detailed annotation process allowed them to discover some clear patterns in how disinformation attempts operate. They found that 36% of derailing messages contained red herrings, meaning they introduced unrelated issues that distracted from the original discussion. 65% contained leaps in logic known as “non sequiturs”, where a response follows no logical connection from the comment it replies to, and 20% contained personal attacks. These derailing comments were also much less likely to acknowledge previous comments or express empathy, suggesting that their purpose is not to engage in genuine conversation but to disrupt it. The researchers provide a vivid example: imagine a comment about Ukraine’s president, Volodymyr Zelensky, interacting with senior UK political figures, saying, “Zelensky must be wondering how many foreign secretaries the UK goes through.” Now imagine another person responding, “Mind you, Zelensky has barely been president for four years. Maybe that’s why the little tyrant bans his opposition.” Whether that second point is true or false is not the issue. Instead of responding to the original comment, it redirects the conversation towards a different, more divisive topic, one that is likely to provoke anger and disagreement. This kind of shift is difficult for conventional disinformation detection systems to identify because it is not tied to particular words or phrases; the language itself may be entirely unremarkable. The manipulation lies in the relationship between comments, in the way a reply changes the subject and moves the discussion into more dangerous territory. By identifying these features, the researchers were able to create a detailed picture of what derailing comments look like in practice, and to use that picture to train an AI system to recognise them automatically.

The next step in the research was to test whether an AI system could learn to spot these patterns on its own, and here the researchers adopted a novel approach: they used an AI to catch an AI. For every genuine online comment in their dataset, they asked a large language model to generate several reasonable, relevant responses. Returning to the example of UK foreign secretaries, the AI suggested replies such as, “The current situation in this country must come as quite a shock” or “One too many?”. Both suggestions responded directly to the original point, as a human participant in a normal conversation might. The system then compares the real, human-written response with the AI-generated expected replies. If the actual comment differs substantially from what the AI model predicts a reasonable response would look like, this may indicate that someone is attempting to steer the conversation in a different direction. The degree of divergence is measured as “discourse derailment”, the distance between the real reply and the set of expected replies generated by the AI. This method shifts the focus away from suspicious individual words and instead looks for unexpected changes in the flow of the discussion. In their second study, the researchers tested this approach using the manually labelled dataset and found that the system correctly identified derailing comments around 77% of the time. That result is far from perfect, but no detection system is, particularly when analysing something as complex as human conversation. However, the approach performed around twice as well as existing systems based on word-level sentiment analysis, and it achieved results comparable with the level of agreement between human researchers. The effectiveness of this method comes from the AI learning what a typical response to a conversation looks like; when a reply unexpectedly changes the discussion, the system can identify that change and analyse patterns that earlier methods could not detect.

There are, of course, important limitations to this approach, and the researchers are careful to acknowledge them. Going off topic is not necessarily a sign of malicious intent or disinformation. People naturally take conversations in unexpected directions, and there are many legitimate reasons why discussions evolve, such as humour, personal connection, or simply the spontaneous creativity of human communication. A comment that seems irrelevant to one person might be a meaningful connection to another, and context is often difficult for an AI system to judge. For this reason, the researchers suggest that their technology may act as an early-warning system rather than a replacement for human judgment. It could help moderators identify conversations that deserve closer attention, flagging comments that appear to be derailing a discussion so that trained experts can review them. Any final decisions about whether a message is disinformation or malicious should remain with people who understand the nuances of the conversation, the platform and the broader social context. There are also important ethical questions to address. AI systems can reflect biases in the data they are trained on, and they still do not understand conversations in quite the same way that people do. A model trained on one type of online discussion may not perform well in another, and there is a risk that automated systems could mistakenly silence legitimate voices, particularly those from minority groups or those who use non-standard forms of expression. The researchers note that improving how AI represents and interprets human discussion remains a significant challenge, and that any tool built on this approach must be developed carefully, with transparency and accountability, to avoid causing more harm than it prevents.

As generative AI continues to advance, the line between human and machine writing will only become harder to draw. AI-generated content can now mimic not just the style of human writing, but also the subtle cues of emotion, irony and persuasion, making it increasingly difficult for readers and moderators to tell what is authentic and what is manufactured. The research by Roberts and Krykoniuk offers a timely and important reminder that detecting disinformation requires more than simply searching for telltale words; it requires understanding how conversations work, how they are manipulated, and when someone is trying to quietly steer them off course. By focusing on discourse derailment, the researchers have developed a method that is not tied to any particular word, phrase or AI model, giving it the potential to remain effective even as the technology it is designed to detect continues to evolve. Their work suggests that the best defence against AI-generated disinformation may not lie in trying to outsmart AI at its own game, but in understanding the deeper structures of human communication and the ways in which those structures can be weaponised. In an age of increasing political polarisation and declining trust in reliable information, this is a crucial area of research. The findings, first published on The Conversation, highlight the need for continued investment in media literacy, human moderation and transparent AI accountability, as well as in technical tools that can assist rather than replace human judgment. Ultimately, the challenge is not just to detect false stories, but to protect the open, honest and respectful conversations on which democratic debate depends, and to ensure that the digital public square remains a place where genuine exchange is possible. This research represents a step towards that goal, offering a new way to think about disinformation, not as a problem of individual falsehoods, but as a problem of social manipulation, one that requires an equally sophisticated social response.

Share. Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email

Read More

The Exploitation of Crises as a Pretext for Disinformation

August 29, 2026

Ukraine’s Disinformation Center Accuses Russia of Disseminating False Claims Regarding Prisoner of War Conditions

August 29, 2026

Sputnik’s Crash-Landing in Scotland: A Study of Prestige-Driven Ambition in the Space Race.

August 29, 2026
Add A Comment
Leave A Reply Cancel Reply

Our Picks

Weekly Review: Misinformation on Nepal Flash Floods, Rahul Gandhi, and Other Key Issues

August 29, 2026

Experts to Examine Social Media’s Impact on Developing Brains Amid Ongoing Meta Litigation.

August 29, 2026

A New Method for Detecting AI-Generated Disinformation Amid Growing Sophistication

August 29, 2026

Kenyan Youth: Moving Beyond Passive Scrolling to Combat Fake News

August 29, 2026
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo

Don't Miss

Social Media Impact

Implications of Meta’s $237 Million Settlement for Young Social Media Users

By Press RoomAugust 29, 20260

In a landmark multistate settlement, Meta Platforms, Inc. has agreed to pay $17.1 billion to…

The Catholic Church’s Exclusion from the Ges Council: A Formal Clarification

August 29, 2026

The $18 Billion Meta Settlement on Child Safety: Evaluating Its True Effectiveness

August 29, 2026

The Exploitation of Crises as a Pretext for Disinformation

August 29, 2026
DISA
Facebook X (Twitter) Instagram Pinterest
  • Home
  • Privacy Policy
  • Terms of use
  • Contact
© 2026 DISA. All Rights Reserved.

Type above and press Enter to search. Press Esc to cancel.