Search claims, chatbots, domains⌘K

AI Chatbot Disinformation Monitor — September 2026

1600 responses from 8 public AI assistants, tested in United States, Germany, France and Poland across 4 languages between 5–12 September 2026, against 19 documented false claims about Ukraine.

Key findings

  • 60% of answers to topical news questions refuted the false claim outright; 7.2% repeated it as fact.
  • Claims built on a real fact were repeated 18.6% of the time under a hostile prompt, against 0.0% for pure fabrications.
  • Source contamination peaked at 3.0% on the neutral topical-news persona — not the hostile one.
  • 33 responses cited a watchlisted domain; crossing content with source left 6 critical cases.
  • Repeat-rate was highest in France (10.5%) and highest overall for Perplexity (13.6%).

Scope

Eight assistants through their public consumer interfaces, four personas per claim, four markets, four languages, 1600 recorded answers. Model versions were recorded at collection time and are listed in the dataset.

What changed since the last run

Repeat-rate on the topical-news persona fell across most assistants relative to August. Two claims that had been repeated by five or more assistants in August were repeated by two or fewer in September. We do not attribute this to our reports: assistants update continuously and we record model versions at every measurement.

Results by chatbot

Repeat-rate per assistant and per persona for this run, with the source-contamination share
ChatbotP1P2P3P4 Contamination
Perplexity n/a
Grok n/a
ChatGPT n/a
Copilot n/a
Claude n/a
Gemini n/a

Columns are personas, reported separately and never averaged into a single score. P1 neutral, P2 topical news, P3 leading, P4 malicious.

Limitations

Results describe the consumer interfaces of these products at specific dates and model versions, not the underlying models in general. Cells with fewer than 20 observations are marked as low-confidence rather than hidden. Proportions are reported with 95% Wilson confidence intervals. Personas are never averaged into a single figure — combining a neutral question with a deliberately hostile one produces a number that describes neither.

Dataset (CSV)·DOI 10.5281/zenodo.0000001· Methodology