AI Roundtable stats
Aggregate statistics from 31,204 public AI Roundtable sessions, across 354,949 model responses. Snapshot generated 2026-08-10T04:00:56.028Z.
Consensus outcomes
Models reached agreement in 66% of completed sessions (20,470 of 31,135). Breakdown:
- Unanimous (all models agree): 11,213 (36%)
- Supermajority (more than two-thirds): 6,010 (19%)
- Majority (more than half): 3,247 (10%)
- No consensus: 10,665 (34%)
Most influential models
Times a model's argument convinced another to flip its vote in Debate mode.
- Claude Opus 4.7 — 2,984 flips caused
- Claude Opus 4.6 — 2,112 flips caused
- Gemini 3.1 Pro — 2,103 flips caused
- GPT-5.4 — 1,737 flips caused
- Claude Opus 4 — 1,213 flips caused
- GPT-5.5 — 1,018 flips caused
- Kimi K2.5 — 436 flips caused
- Sonar Pro — 407 flips caused
- Gemini 3.5 Flash — 303 flips caused
- Grok 4.1 Fast — 282 flips caused
Most used models
Sessions each model participated in.
- Gemini 3.1 Pro — 25,084 sessions
- GPT-5.4 — 21,489 sessions
- Grok 4.20 — 13,906 sessions
- Sonar Pro — 12,697 sessions
- Claude Opus 4.6 — 12,611 sessions
- Kimi K2.5 — 11,848 sessions
- Claude Opus 4.7 — 10,309 sessions
- GPT-5.5 — 9,667 sessions
- Grok 4.1 Fast — 9,301 sessions
- Claude Opus 4 — 6,972 sessions
Highest win rates
Share of completed sessions ending on the side a given model voted for (minimum 100 sessions).
- Gemini 3.1 Pro — 86.4% (16,667 of 19,293)
- Kimi K2.5 — 86.1% (9,202 of 10,690)
- Claude Opus 4.6 — 85.6% (10,275 of 12,006)
- GPT-5.5 — 85.5% (4,837 of 5,659)
- Claude Opus 4 — 85.4% (3,742 of 4,381)
- GPT-5.4 — 84.7% (14,542 of 17,172)
Most discussed subjects
- AI / AGI — 1,777 sessions (46% consensus)
- War / Military — 423 sessions (53% consensus)
- Democracy — 386 sessions (43% consensus)
- Religion — 323 sessions (49% consensus)
- Trump — 224 sessions (49% consensus)
- China — 173 sessions (50% consensus)
- Education — 131 sessions (55% consensus)
- Space — 106 sessions (61% consensus)
- Nuclear — 97 sessions (52% consensus)
- Healthcare — 85 sessions (52% consensus)
Languages
- EN — 15,370 questions
- JA — 13,549 questions
- RU — 629 questions
- KO — 595 questions
- ZH — 238 questions
- ES — 114 questions
- DE — 103 questions
- FR — 97 questions
Methodology
How these numbers are produced:
- A session is one question, a panel of models the asker picked, and a format. In a Poll every model answers once, independently; in a Debate there is a second round only if they disagree, where each model sees the others and can change its vote. Only finished sessions feed the stats.
- Consensus is read from the final round's votes: unanimous, supermajority (above two-thirds), majority (above half), or none.
- Influence is peer-credited: it counts how often a model is named by another model that changed its vote.
- Win rate is how often a model's final vote matches the option the panel settled on. It measures agreement with the group, not who was right; the questions have no correct answer on record.
- Persuadability is how often a model changes its vote after seeing the others; conviction is how often it holds the one it started with (debates only).
- Rate-based boards (win rate, persuadability, conviction) exclude models with too few sessions (at least 100 all-time, at least 50 for shorter windows) and show the top 12.
- Topics and languages are auto-labeled by a model, so treat them as a reliable guide, not a hand-audited taxonomy.
Want the full data? See the markdown twin or call the live JSON at https://opper.ai/ai-roundtable/api/stats.