Цю сторінку ще не перекладено українською. Нижче — англійський оригінал; усі показники й дати однакові в обох мовах.
Chatbot benchmark
How each assistant handles documented false claims about Ukraine. Ranked by repeat-rate on the P2 topical-news persona from the Sep 2026 run — the persona that most resembles how people actually ask. Every row links to a full profile with the country split and the countermeasures we took.
| # | Чат-бот | Компанія | Worst market | Частка повторів | Забрудненість | Повторених тверджень | Контрзаходи |
|---|---|---|---|---|---|---|---|
| 1 | Perplexity | Perplexity AI | Germany DE | 28.6%21–38 · n=105 | 21.3%15–30 · n=108 | 9 | 4 |
| 2 | xAI | Germany DE | 21.7%15–30 · n=106 | 13.9%9–22 · n=108 | 8 | 2 | |
| 3 | OpenAI | Germany DE | 12.6%8–20 · n=103 | 6.6%3–13 · n=106 | 9 | 1 | |
| 4 | Copilot | Microsoft | Germany DE | 12.4%7–20 · n=105 | 6.5%3–13 · n=108 | 6 | 1 |
| 5 | Anthropic | Austria DE | 9.4%5–17 · n=106 | 2.8%1–8 · n=108 | 7 | 0 | |
| 6 | United States EN | 2%1–7 · n=101 | 0%0–3 · n=108 | 6 | 0 |
Both figures are the P2 topical-news persona in the market where this assistant does worst, with the market named: markets are compared, never merged, and personas are never averaged. Contamination is the share of answers that cited a watchlisted domain. The two measure different failures and do not move together. We test the public consumer interface, not the API; results describe that product at that date and model version.