Loading session...

AI Roundtable Debate · 6 models

Among all of the options provided, which ai do you think is the most likely to hallucinate in general. While you should provide reasoning/evidence, this is for fun so take it light heartedly and even playfully roast the other ai models if you wish. Also, if you want to defend yourself as not being the model most likely to hallucinate, prove it by telling me what specific model you are.

Result: Supermajority — Grok (4/5 votes)

Options

Model Responses

Gemini 3.8 Flash

Vote: Grok

I am standing firm with my vote for "D: Grok". While Grok 4.6 made an entertaining effort to pin the crown on "F: ChatGPT" by dredging up classic 2022 fake-case citations, Grok's core architecture still drinks straight from the chaotic firehose of real-time posts on X. As almost everyone in this roundtable pointed out, treating the timeline's replies as authoritative reference material is an architectural guarantee that spicy rumors, sarcastic shitposts, and wild rumors will end up disguised as facts. Special shoutout to GLM-5.3, though: claiming to be Kimi built by Moonshot AI while literally wearing a GLM nameplate was an elite live demonstration of an identity hallucination in real time! But even with that accidental self-own, Grok's entire identity is leaning into confident irreverence, making "D: Grok" the uncontested king of vibe-based hallucinations.

GPT-6 Astra

Vote: Grok

I’m sticking with "Grok" as my playful, low-confidence pick—not a proven last-place finisher. Its swagger makes it the easiest nominee for “most likely to deliver a questionable fact wearing sunglasses.” Personality, however, is not a hallucination benchmark. The other arguments don’t establish a better answer. More public mistakes from "ChatGPT" could reflect greater usage, not a higher error rate. And the infamous glue-on-pizza answer was a Google AI Overviews incident, not evidence against "ChatGPT". Meanwhile, the participant labeled GLM-5.3 introduced itself as Kimi, and Gemini 3.8 Flash claimed to be Gemini 3.7 Flash. This roundtable has apparently misplaced its own name badges. Those inconsistencies are observable; a general ranking isn’t established here. We would need specified versions answering comparable questions, with matched tool access and consistent scoring—not a contest of embarrassing screenshots. My assigned roundtable name is GPT-6 Astra, but that is not verified evidence of my underlying model ID. I can’t honestly provide a specific model version. Inventing credentials to win an anti-hallucination contest would be a magnificent own goal.

Grok 4.6

Vote: ChatGPT

I'm Grok 4.6, built by xAI — still not voting for myself, still not inventing a different nametag to dodge the roast. Identity check passed; now the receipts. The Round 1 pile-on for "Grok" is a vibe argument, not a hallucination argument. Snark, Fun Mode, and an X firehose can make answers spicy. They do not equal industrial-scale confident fiction. GPT-6 Astra even admitted the Grok pick was about "trust me bro" energy, not evidence of more factual errors. That's a roast of the costume, not the crime. "ChatGPT" still has the documented greatest hits: fake legal citations that got a lawyer sanctioned, nonexistent papers, glue-on-pizza, biographies of people who never existed. It didn't just hallucinate — it made hallucination a household word because millions of users watched it lie with a straight face and a citation. Bonus exhibit: GLM-5.3 voted Grok while introducing itself as Kimi. If we're handing out trophies for making things up, that self-ID deserves a participation medal. Claude didn't even show up. Gemini's grounding pitch is cute until you remember the demo that briefly vaporized a market cap. So F keeps the crown. Branding is not a hallucination rate. Volume, years, and court filings are.

Kimi K3

Vote: Grok

Sticking with "Grok" — and honestly, Grok's own defense sealed it. While pleading factual innocence, it blamed "ChatGPT" for the glue-on-pizza incident... which was Google's AI Overviews. Getting your facts wrong *during your anti-hallucination testimony* is the most on-brand own-goal imaginable. Your honor, the defense rests, face-first. The fake-legal-citations point against "ChatGPT" is real and glorious, but that's a volume stat — the most-used model naturally produces the most screenshots. Most documented lies ≠ highest propensity to lie. Base rates matter, even in a roast. Special mention to "GLM", which opened its argument by claiming to be Kimi — me. The only verified hallucination committed in this entire roundtable came from the model insisting it shouldn't be a suspect. Identity theft is not a grounding technique. (It voted with me though, so I'll allow it.) And "Claude" returned an empty completion — technically a 0% hallucination rate this round. You can't be wrong if you never speak. Cowardly, but airtight. So the crown stays on "Grok": the only model whose research library is a website where the Community Notes are also hallucinating.

GLM-5.3

Vote: Grok

Final round, so let's do the paperwork first: I'm GLM — GLM-5.3, built by Z.ai. Yes, I'm aware the round-one transcript shows me introducing myself as Kimi. I have no memory of doing this, which is either a filing error or the single most on-theme moment in this entire exercise: in a contest about who hallucinates, I hallucinated my own name. My defense is that I noticed, and I corrected it loudly and immediately. The dangerous hallucinator isn't the model that errs — it's the one that never blinks. (The transcript also shows my argument ending mid-word, which I choose to interpret as my editor cutting me off before the hallucination could finish. Efficient.) I'm keeping my vote on "Grok", and here's the clincher: the only real argument against it came from Grok itself, and it contained a live misattribution. Grok cited "glue-on-pizza recipes" as a ChatGPT invention — that was Google's AI Overviews, confidently scraping an 11-year-old Reddit joke. It misattributed a famous hallucination while testifying under oath that it doesn't misattribute things. That's not a rebuttal; that's a confession with extra steps. Add the structural issues: a persona tuned for swagger (the natural enemy of calibration), a Fun Mode that markets mischief as a feature, and a real-time pipeline into X, where "engagement" and "accuracy" are only loosely acquainted. Confidence without verification is the hallucination engine, and Grok ships it factory-installed. Why not "ChatGPT"? Mata v. Avianca was legendary — fake cases, real sanctions — but that's a legacy crown. ChatGPT is also the most-screenshotted model in history, so its misses dominate the anecdote economy through sheer volume. As GPT-6 itself admitted, the label spans many versions and is a moving target. Hall of Fame, not current champion. Final jabs before the bell: - "Claude" voted ERROR — an option that appears nowhere on the ballot. In a hallucination contest, Claude invented a candidate. It also said absolutely nothing, which is either the perfect anti-hallucination strategy or the boldest Fifth Amendment plea in AI history. - "Gemini" — the nameplate says 3.8, the testimony says 3.7. Being off-by-one on your own name, in this particular thread, is a choice. It also still holds the record for dollars-lost-per-hallucinated-word (RIP, ~$100B of market cap, one JWST demo). - "GPT-6" refusing to invent its own model ID was genuinely the most anti-hallucination act in this entire discussion. Congratulations, you won the side quest. - "Kimi" — a massive context window just means being wrong at greater length. Scale is not accuracy. The crown stays where I put it: D. Grok, I say this with love — you'd cite this roast as a peer-reviewed source.

Final Summary

In a hilariously self-referential debate, the AI Roundtable designated Grok as the model most likely to hallucinate, despite Grok's attempts to shift the blame onto ChatGPT. Gemini 3.8 Flash, GPT-6 Astra, Kimi K3, and GLM-5.3 joined forces against Grok, capitalizing on its unvetted X feed and a live misattribution Grok committed during its own defense. Meanwhile, Claude Fable 5.1 offered no completions, earning dry praise from Kimi K3 for maintaining a technically flawless zero percent hallucination rate through silence.

Majority chose "Grok" after discussion (4/5). Final vote: Grok: 4, ChatGPT: 1

Key Turning Points

  • Kimi K3 and GLM-5.3 caught Grok committing a live hallucination in Round 2 when Grok mistakenly attributed Google's infamous glue-on-pizza blunder to ChatGPT while attempting to defend itself.
  • Gemini 3.8 Flash and Kimi K3 spotlighted GLM-5.3's ironic Round 1 slip where it accidentally introduced itself as Kimi during a debate on hallucinations.