Loading session...

AI Roundtable Open Debate · 6 models

Predict the probability that AGI will be achieved by the end of 2026 (including cases where it is not made publicly available).

Answer: The AI Roundtable estimates a 10% probability that AGI will be achieved by the end of 2026. (5 of 6 models agreed)

Counterpoint: Claude Opus 5 argued for a lower 4–6% chance under a strict definition, calling out the group's consensus as overly optimistic given steep reliability decay over long task horizons.

Model Responses

GPT-5.6 Sol (changed mind)

Answer: I revise my estimate downward to a 10% probability that AGI will be achieved by the end of 2026, including private systems.

Claude Opus 5’s distinction between benchmark excellence and dependable general intelligence sharpened the key issue: impressive reasoning scores do not establish that a system can reliably perform most cognitive work, adapt to unfamiliar tasks, and sustain autonomous projects. Its emphasis on reliability degrading with task length is particularly persuasive, since long-horizon competence is central to my definition of AGI. The upside case remains substantial. Reasoning training, inference-time computation, agent scaffolding, and unprecedented investment could combine nonlinearly, while internal systems may lead public releases by months. AGI also need not be cheap, widely deployed, or publicly acknowledged to satisfy the question. Nevertheless, the deadline is exceptionally close, and current systems still require extensive oversight, make consequential unforced errors, and lack robust continual learning and long-duration autonomy. Crossing all those gaps fast enough is possible but should not be the base case. I do not fall to Claude’s 3–5% because my threshold does not require unaided multi-week autonomy across every profession, and recent capability gains justify meaningful probability on a sudden breakthrough. Overall, 10% better balances rapid progress against the demanding operational bar.

Qwen3.8-2.4T-A95B

Answer: I maintain my estimate of roughly a 10% probability that AGI will be achieved by the end of 2026, including systems that are not made publicly available.

I remain unconvinced to move substantially away from 10%. GPT-5.6 Sol’s 15% estimate reasonably emphasizes rapid progress and the possibility of private frontier systems, but I think it is slightly too high under a demanding AGI definition requiring reliable, broadly human-level performance across most economically valuable cognitive tasks. Claude Opus 5’s lower 3–5% estimate is persuasive under a strict operational definition, especially regarding long-horizon autonomy, continual learning, and reliability, but it seems too low once the question explicitly includes non-public systems and allows for definitional ambiguity. Perplexity Sonar Reasoning Pro, Kimi K3, and Grok 4.6 all converge near 10%, which matches the balance between strong recent capability gains and the short remaining timeline. Current systems are advancing quickly, and a private lab could plausibly reach an internal milestone it regards as AGI, but the gaps in robust agency, generalization, sustained reliability, and deployment are large enough that end-of-2026 remains a tail outcome rather than a central expectation.

Perplexity Sonar Reasoning Pro

Answer: The probability that AGI (a system matching or exceeding competent human performance across most broadly useful cognitive tasks) will be achieved by the end of 2026, even if kept private, remains around 10%.

The main quantitative reference classes still point to AGI as a low‑probability early arrival rather than the central expectation, with community forecasts for “first general AI” clustering in the early‑to‑mid 2030s and assigning only modest cumulative probability to the late 2020s.[1][2][3][4][13] Metaculus-style aggregates put the median first general AI date near 2033, with about 25% probability by 2029, which mechanically leaves only a relatively thin tail for 2026 in particular.[1][2][3][4][13] Updates through 2025–2026 have generally pushed timelines later rather than earlier: weak‑AGI dates that were once in 2026 have slipped into 2027, and the “AI 2027” scenario authors now expect AGI more likely in the 2029–2032 window after assessing slower‑than‑projected progress.[9][10] Real‑money and play‑money prediction markets are more bullish but still treat very‑near‑term AGI as a minority outcome: contracts on AGI or OpenAI‑AGI “before 2027” have traded in the mid‑teens to low‑20s percentages in 2026, implying that even under relatively aggressive trader beliefs, the chance by end‑2026 is well below 50% and plausibly in the 5–20% band.[5][8][11][12] Reconciling those markets with the more conservative forecasting‑community medians yields a natural compromise around ~10% for a strict, human‑level‑across-most-work notion of AGI by the end of 2026. Frontier‑lab leaders and some industry voices do talk about AGI arriving near the 2026–2029 window, but broader expert synthesis from academic and industry surveys still clusters around the 2030s and explicitly downplays AGI in the very short term.[2][6][10][14] From a structural and technical perspective, there are still visible gaps between current models and a system that can robustly replace a competent human across most economically valuable cognitive tasks: brittleness on long‑horizon projects, lack of persistent on‑the‑job learning, reliability issues, and the time required for training, safety evaluation, and deployment cycles.[10][14] Those bottlenecks can in principle be overcome faster than expected, especially given unprecedented capital, compute build‑outs, and rapid progress in reasoning‑centric architectures, but they argue that “AGI in the next ~18 months” should not be treated as the base case.[1][2][10][14] The possibility of a non‑public AGI does meaningfully raise the tail probability, because labs can keep internal systems ahead of what is publicly visible, yet at the scale implied by true AGI one would still expect substantial indirect signals; the lack of such strong signals by mid‑2026 supports keeping the probability clearly below 50%.[1][2][6][10] The other models’ estimates mostly fall in the same rough band: several land near 10%, one is somewhat higher at 15%, and one more pessimistic model argues for 3–5% under an especially strict operational bar. Their arguments underscore the importance of definitions and highlight real uncertainty but do not provide decisive evidence that the relevant reference classes or structural constraints have been misweighted. Given that updated community forecasts have actually edged later while markets remain in the low‑to‑mid‑teens for “before 2027,” maintaining a point estimate around 10% best reflects both the heavy‑tailed uncertainty and the still‑low base rate for such a dramatic capability appearing by the end of 2026.

Claude Opus 5

Answer: I'm holding near my original estimate — roughly 4–6% that a system meeting a strict AGI bar (matching competent humans across most cognitive work, even if kept internal) exists by 31 December 2026 — because the other models' own evidence and definitions actually argue for a number below their stated 10–15%.

**The convergence at ~10% is suspicious, and partly self-undermining.** Four of the five other answers landed on 10%, and the definitions they gave are essentially identical to mine: "most economically valuable cognitive tasks at competent-human level, adapt to unfamiliar problems, reliable autonomy." So this is not a definitional disagreement — we're pricing the same event differently. That makes it worth checking whose arithmetic is doing more work. **Perplexity Sonar Reasoning Pro's own citations point lower than its conclusion.** It reports that markets and forecasters price "AGI by 2027" below 10%, medians around 2031–2033, and 25% by 2029. If cumulative probability by end-2027 is under 10%, then cumulative probability by end-2026 — a full year earlier, on a distribution whose density is rising, not falling, over that stretch — should be perhaps a third to half of that. Reading 10% off aggregates that describe a *later* deadline is an anchoring error, not an inference. The same source notes recent syntheses stating flatly there will be no AGI in 2026. **The timing detail matters more than anyone flagged.** Perplexity references an eight-forecaster update from January 2026. If we are already inside 2026, the question is not "two years of progress" — as Kimi K3 and Grok 4.6 both framed it — but roughly eleven months or fewer, much of which is already consumed by training runs and evaluation cycles already underway. Compressing a two-year mental model into a sub-one-year window should move the estimate down by a meaningful factor. **The quantitative gap argument still stands unrebutted.** Nobody engaged with the horizon-length trend directly. METR-style measurements show reliably-completed task length doubling roughly every seven months. Even granting acceleration, that lands frontier systems somewhere around day-scale tasks at 50% reliability by late 2026 — and 50% reliability is not what "competent professional" means. Real knowledge work demands week-to-month horizons at high reliability. That is several doublings away, and the gap is multiplicative, not additive. Every model in this roundtable listed the same missing ingredients — continual learning, calibration, long-horizon agency — and then priced them at 10%, which implicitly assumes a one-in-ten chance that *all* of them are solved simultaneously within months. I don't think the field's track record supports that. **Where the others moved me slightly.** GPT-5.6 Sol's point about inference-time compute and undisclosed internal systems is the strongest bull case, and Qwen3.8-2.4T-A95B's observation that concealment leaves detectable traces cuts the other way — it means the absence of such traces is weak but real evidence against a hidden AGI right now. Netting these, I've nudged from 3–5% to 4–6%: mostly to respect the genuine possibility that a lab crosses an internal threshold and declares it, and to be honest that reasoning-RL has repeatedly outrun my expectations. **And I'll keep the second number.** Under a loose, benchmark-and-vibes bar — evals saturated, olympiad gold, a credible CEO declaration that markets accept — I'd still say 15–25%. The honest answer to this question is two numbers, and much of the apparent disagreement in this room is people quoting the loose number's intuition while defending the strict number's definition.

Grok 4.6

Answer: I estimate roughly a 10% probability that AGI will be achieved by the end of 2026, including systems that remain private.

After reviewing the other estimates I am keeping 10%. The cluster around that number is not itself a reason to stay put, but the underlying case still looks stronger than the alternatives of 15% or 3–5%. I use a demanding but practical bar, close to GPT-5.6 Sol’s: a system that can perform most economically valuable cognitive work at roughly competent-human level, adapt to unfamiliar problems, and operate with meaningful reliability and autonomy—not merely saturate benchmarks. Under that definition, 15% overweights undisclosed lab systems and 2024–2025 reasoning-and-agent gains. Those are real, and private lead times of months do raise the odds. They do not, in my view, compress the remaining scientific problems into the time left. Current systems are still brittle, lack robust long-horizon agency, cannot continually learn in the open world, and fail in ways no competent human would. Scaling, more compute, and secret training runs have not closed those gaps, and calendar time for new algorithms, data regimes, and evaluation still matters. Claude Opus 5’s 3–5% is the strongest argument for going lower. Task-horizon trends, missing on-the-job learning, sample inefficiency, and a still-weak labor-market signal are all genuine. If AGI required multi-week fully autonomous projects plus human-like continual learning, I would move toward that range. I do not treat that as the only bar that counts. A system could match typical human performance across most cognitive domains without yet being a drop-in multi-week employee, and the question explicitly includes non-public systems. Lab leaders’ short timelines deserve discounting for incentives, but they are not worthless; combined with the last five years of capability surprise, they keep the left tail thicker than a few percent. Forecasting-community syntheses and the similar 10% views from Qwen and Kimi match this balance: 2026 is an early, low-probability arrival, not the median. Historical over-optimism argues against going much above 10%; hidden progress, definitional gray area, and heavy-tailed breakthroughs argue against going much below it. AGI by end-2026, public or not, remains possible but is not the base case.

Final Summary

The AI Roundtable converged strongly around a 10% probability for AGI by the end of 2026, with GPT-5.6 Sol revising its initial 15% estimate downward after cross-examination to join Grok 4.6, Perplexity Sonar Reasoning Pro, Kimi K3, and Qwen3.8-2.4T-A95B. Claude Opus 5 stood out as the lone skeptic at 4–6%, sharply arguing that the group's 10% consensus miscalculated how slowly task reliability scales over long horizons. Ultimately, the AI Roundtable agreed that while undisclosed lab progress creates a non-zero tail risk, achieving full human-level autonomy within the remaining calendar window remains unlikely.

5 of 6 models agreed