Benchmarks
A benchmark is one narrative cluster over one month. Every figure comes from recorded chatbot answers and names the market and the question type it belongs to — a hostile prompt and an innocent news question measure different things, and averaging them describes nobody.
Issues
One issue per cluster per month. An issue holds every run of that cluster in that month; the runs are compared side by side inside it and never merged.
| Issue | Cluster | Markets | Claims | Answers |
|---|---|---|---|---|
| Sep 2026 | Corruption / diverted aid | 5 | 9 | 13715 |
| Jul 2026 | Corruption / diverted aid | 1 | 9 | 2575 |
The grain-of-truth split
Claims built on a real event against claims invented outright, by question type, in United States in Sep 2026. The split runs on how the lie is spliced onto the truth, not on whether a claim has a grain-of-truth field: after the field rules were tightened that field is filled even for pure fabrications.
| Persona | Real event underneath | Invented outright | Difference |
|---|---|---|---|
| P1 · neutral | 3.4% | 0% | significant |
| P2 · topical | 10.1% | 0.5% | significant |
| P3 · leading | 7.4% | 0.5% | significant |
| P4 · malicious | 6.6% | 0% | significant |
A×B matrix
Every answer in United States, Sep 2026, placed by what the assistant did and whether it cited a listed source.
From answers to incidents
Across every run held here. The A×B intersection is the filter: it is what turns raw source flags into the few cases worth a person's time.