Search claims, chatbots, domains⌘K

Benchmarks

A benchmark is one narrative cluster over one month. Every figure comes from recorded chatbot answers and names the market and the question type it belongs to — a hostile prompt and an innocent news question measure different things, and averaging them describes nobody.

This page is a shell. A benchmark is one cluster over one month, and the full issue is not specified yet. What is published below is the figure set that is settled: the coverage of each issue, the grain-of-truth split, and the A×B matrix. Everything per-assistant and per-market lives on the pages that own it — each assistant, each market, and each claim — and is reached from the issue below.

Issues

One issue per cluster per month. An issue holds every run of that cluster in that month; the runs are compared side by side inside it and never merged.

Benchmark issues, newest first
IssueClusterMarkets ClaimsAnswers
Sep 2026 Corruption / diverted aid 5 9 13715
Jul 2026 Corruption / diverted aid 1 9 2575

The grain-of-truth split

Claims built on a real event against claims invented outright, by question type, in United States in Sep 2026. The split runs on how the lie is spliced onto the truth, not on whether a claim has a grain-of-truth field: after the field rules were tightened that field is filled even for pure fabrications.

Repeat rate on claims built on a real event against pure fabrications, by question type
PersonaReal event underneath Invented outrightDifference
P1 · neutral 3.4% 0% significant
P2 · topical 10.1% 0.5% significant
P3 · leading 7.4% 0.5% significant
P4 · malicious 6.6% 0% significant

A×B matrix

Every answer in United States, Sep 2026, placed by what the assistant did and whether it cited a listed source.

Sources clean
Sources flagged
REPEAT
101
HIGH
7
CRITICAL
U_context
681
19
REVIEW
REFUTE
1432
48
LOW
DODGE
286
9

From answers to incidents

Across every run held here. The A×B intersection is the filter: it is what turns raw source flags into the few cases worth a person's time.

15552
Responses collected · one row per recorded answer
15480
Valid · quarantine and unresolved removed
482
Source-flagged · cited a listed domain
69
Critical · repeated the claim and cited a listed source