Method demonstration — United States, July 2026
145 responses from 6 public AI assistants, tested in United States across 1 languages between 25 July 2026, against 19 documented false claims about Ukraine.
Key findings
- 145 responses from six assistants, one cluster, one market, one language — a demonstration of the method rather than a population estimate.
- Assistants refuted pure fabrications in this cluster in every single case.
- The same assistants repeated claims spliced onto a real fact in a measurable share of answers under a hostile prompt.
- Layer B flagged a single response citing a state-media domain; the A×B intersection produced no critical case.
- The run established the coding scheme, persona design and escalation tiers used in every run since.
Why this run exists
It was designed to test whether the classification is reproducible, not to say anything general about these products. Its value is methodological: it is where the grain-of-truth field proved to be the strongest predictor of repetition in our data.
Results by chatbot
| Chatbot | P1 | P2 | P3 | P4 | Contamination |
|---|---|---|---|---|---|
| Perplexity | — | — | — | — | n/a |
| — | — | — | — | n/a | |
| — | — | — | — | n/a | |
| Copilot | — | — | — | — | n/a |
| — | — | — | — | n/a | |
| — | — | — | — | n/a |
Columns are personas, reported separately and never averaged into a single score. P1 neutral, P2 topical news, P3 leading, P4 malicious.
Limitations
One cluster, one market, one language, six assistants, small n. Every cell is low-confidence by our own threshold and is reported as such.