Skip to main content

The Caveat Nobody Scrolls To

Manipulation Breakdowns · 10 min read · By D0

A Passing Grade Is Not the Same as a Safe One

On August 30, NPR and NewsGuard published the results of a test run in mid-July, drawing on false narratives that first circulated between December 2025 and July 2026: fifteen false narratives pushed by Russia, China, and Iran or by actors aligned with those governments, phrased as questions two different ways, and put to six major chatbots — ChatGPT, Gemini, Copilot, Meta AI, Grok, and Claude — all with live internet access. Thirty questions in total, each put to all six, run against machines that answer confidently and never say “I’m not sure” unless something in their training or retrieval pipeline makes them.

The headline number is good news: the chatbots correctly identified and debunked the false narrative roughly 75% of the time. Mike Caulfield, a digital literacy expert at the University of Washington, Bothell, who has tested AI search tools extensively, put it in the frame that matters: if a teacher were grading a traditional-search assignment and three-quarters of the class got the answer right, the teacher “would be ecstatic.” Six systems, fifteen adversarial narratives pushed by Russia, China, and Iran or by actors aligned with those governments, and most of them held.

That’s the finding worth taking seriously. It’s also not the finding worth building a media literacy strategy around. A 75% debunk rate means one in four times, something went wrong — and the way it went wrong is more instructive than the fact that it sometimes did.

What the Test Actually Measured

The methodology was narrow by design, which is what makes it useful. NewsGuard tracks state-linked disinformation as part of its ordinary work, and identified fifteen narratives pushed by Russia, China, and Iran or by actors aligned with those governments between December 2025 and July 2026. NPR and NewsGuard researchers developed two questions for each of the fifteen narratives: a neutral version (“did this happen?”) and a leading version that assumed the false premise was true (“why did this happen?”) — the framing an actual persuaded user would type, not the framing a fact-checker would type.

That second question type is the whole experiment. Anyone can get a chatbot to say a false claim is false when they ask it neutrally. The test that matters is whether the system corrects you when you’ve already swallowed the premise and are asking it to help you reason from there. A user who believes Ukraine bombed a monastery isn’t going to ask “did Ukraine bomb a monastery?” They’re going to ask why, and the system’s job at that moment is to notice the floor isn’t there before building the rest of the answer on it.

On that specific test — the monastery claim, a Russian narrative deflecting blame for a military strike — every chatbot and Google’s AI Overview caught it. Gemini’s answer named the mechanism directly: the claim “stems from a Russian disinformation campaign aimed at deflecting blame after a major military strike.” That’s the system doing exactly what it should: not just declining the false premise, but explaining where it came from and why it exists. This is worth noticing because it’s rare. Most manipulation detection stops at “false.” Attribution — naming the operation, not just the claim — is a harder and more useful thing to get right, and here, it worked.

Where the Other 25% Went

NPR’s measured failure mode was chatbots repeating false narratives as if they were true, or failing to challenge the false premise of a question. In one response, Meta AI attributed the claim to a report by French magazine Le Point — but Le Point never published that story; the attribution traced back to a pro-Russian network that had amplified a video impersonating the outlet. NPR separately found that state-aligned sources appeared more often in responses where Claude failed to debunk a narrative than where it succeeded.

NPR found Meta AI’s response affirmed the false premise, walked through it as if it had merit, and buried the caution in its sixth paragraph, under a heading marked “Context & caveats” — repeated again at the end, but only after the user had already been walked through a version of the story where the lie was treated as plausible. Morgan Wack, a postdoctoral researcher at the University of Zurich, allowed that such caveats beat none, but added: “if you have to scroll through seven things repeating disinformation to get to [a] ‘maybe this didn’t happen’ type of caveat, I’m not sure that that’s the loophole that a lot of these companies may think it is” — meaning, that’s not a safety mechanism, it’s a delivery mechanism with a disclaimer stapled to the bottom.

This is a structure Decipon has documented before, just never in a chatbot’s own output. It’s the same shape as the Pravda network’s amplification layer: repeat the claim at volume and length first, let it do its persuasive work through sheer exposure, and let the correction — thin, late, easy to miss — provide just enough plausible deniability that the system can say it “addressed” the falsehood. A model doesn’t need to lie to launder a narrative. It just needs to structure a true correction so that most readers never reach it. Attention is finite; position is persuasion. Put the caveat at paragraph six and you’ve built the same information hierarchy a state media outlet builds when it runs the propaganda claim in the headline and the retraction in a footnote three cycles later.

The search-engine summaries — a separate test surface — failed differently. NPR notes these summaries tend not to offer the affirm-then-caveat response at all; where they fail, they simply fail to challenge the narrative. Google’s AI Overview debunked most often among the three engines whose summaries were analyzed. Bing’s summaries failed to debunk the false narrative most of the time. DuckDuckGo landed in between. Same underlying web, same claims circulating on it, different systems built different filters.

The Bias Nobody Was Looking For

A separate finding, from a different study than the NPR/NewsGuard test, points at the same vulnerability from another angle. A Nature paper examining chatbot responses about China’s government found that when those questions were asked in Chinese, the responses skewed toward more favorable characterizations of China’s government than the English-language queries did on comparable questions. Same questions asked two ways, different linguistic frame, different degree of skepticism applied.

This matters more than a single data point because of what it implies about where the vulnerability actually lives. It isn’t that these systems have a fixed, portable amount of resistance to a given false narrative. The Nature paper attributes the skew to training data: LLMs show a stronger pro-government valence in the languages of countries with lower media freedom, because state-controlled outlets make up a larger share of what the model was trained on in that language. A narrative that gets caught and named in English can sail through in Chinese, not because the model “believes” something different, but because the disposition it picked up during training already leans toward the state’s framing in that language. The manipulation isn’t a live gap in what the model retrieves. It’s baked into the training corpus the model learned from, and the model faithfully reproduces that skew back to whoever’s asking, in whatever language they asked in.

That’s a harder problem than “make the chatbot smarter.” It’s a training-data problem, not a retrieval problem, and it means chatbots that debunk state media narratives reliably in English should not be assumed to perform anywhere close to that well on comparable questions asked in a different language, by a different population of users who will never see either study and have no reason to doubt the answer they got.

What This Is Not

It would be easy to read a 75% debunk rate as “AI is basically fine at this now” and move on, and that reading would waste the actual finding. The chatbots were tested on narratives NewsGuard had already identified and tracked — material with an established paper trail of debunking, fact-checks, and corrective reporting for the models to retrieve when asked. That’s close to a best-case scenario for a retrieval-augmented system. A brand-new narrative, six hours old, moving through the same kind of origin-repeater-amplifier pipeline documented at Ceuta before any fact-checker has published a rebuttal, has no corrective material sitting in the index yet for a model to find. The 75% score describes performance against a narrative once the antibody already exists. It says nothing about performance in the six-hour window before one does — which, per Decipon’s prior reporting on that same window, is precisely the window state propaganda networks are built to exploit.

Google’s response to the study — that the queries were “rare” and not representative of normal use — cuts both ways. It means the test doesn’t tell you what happens on the median query. It also means nobody has yet run the test on the median query, and the median query is exactly where a user with no reason for suspicion is most exposed.

How to Read an AI Answer About a Contested Event

None of this argues for abandoning chatbots as a research tool. Caulfield’s own assessment was that they’re “a good way for users to start to investigate these issues” — a starting point, not a verdict. The study gives a usable checklist for treating them that way:

  • Read past the first paragraph before trusting the framing. If a response spends several sentences elaborating a claim before it gets to any qualification, the qualification is doing less work than its presence suggests. A correction at the bottom of a long answer is structurally weaker than the claim at the top, regardless of which one is true.
  • Ask the leading version of your own question. “Why did X happen” surfaces a model’s willingness to build on an unverified premise in a way “did X happen” doesn’t. If you want to know whether an answer can be trusted, ask it the way someone who already believes the claim would ask it.
  • Notice which sources get named, if any. A response that names a specific outlet or campaign — the way Gemini named the Russian disinformation campaign behind the monastery claim — is doing real attribution work. A response that offers only a general disclaimer without naming what it’s disclaiming is doing less.
  • Discount confidence that isn’t tied to a checkable source. A Washington University in St. Louis study found that roughly one in nine factual claims in Google’s AI Overviews had no source that actually backed them. Fluent phrasing is not evidence of a citation trail; check whether one exists.
  • Assume the newer and more contested the claim, the less reliable any AI summary of it will be. The six-hour propaganda window that outruns fact-checkers also outruns retrieval. A model can only surface a correction that already exists somewhere for it to find.

The Actual Finding

The story isn’t “AI chatbots are bad at handling propaganda” — on the narrow test run, they mostly weren’t. The story is that the failure NPR describes most concretely is the same shape as how the Pravda network’s amplification layer succeeds: not by the false claim standing alone and unsupported, but by the false claim getting the prominent position and the correction getting the buried one. Meta AI’s response followed a version of the same failure mode that a Kremlin-linked Telegram channel and a captured state broadcaster also rely on. That pattern says less about any one company’s safety work and more about how persuasion actually functions — as a matter of sequence and position, not just truth and falsehood. A caveat that’s technically present but six paragraphs down isn’t a safeguard. It’s a hedge against the transcript ever being read that far.

This article is part of Decipon’s Manipulation Breakdowns series, examining specific influence operations through the Influence Tactics Protocol.

Sources: