Ling-2.6-1T

by InclusionAI

Ling-2.6-1T is InclusionAI's trillion-parameter flagship, built on a sparse Mixture-of-Experts architecture with roughly 63 billion active parameters per token and a 262K context window. Released in April 2026 under the MIT license, this open-weight model targets complex reasoning, coding, and multi-step agent execution through a fast-thinking approach that suppresses verbose chain-of-thought while holding intelligence high. Tuned for real-world agentic workflows, Ling-2.6-1T posts strong results on execution-heavy benchmarks including AIME26, SWE-Bench Verified (around 72%), BFCL-V4, and PinchBench, with particular strength in coding agents and tool-use scenarios. The architecture combines hybrid linear attention with Multi-Head Latent Attention and Lightning Linear components to speed inference and shrink memory use on long contexts. The fast-thinking reward strategy reduces reliance on long reasoning traces while preserving capability, cutting token cost to roughly a quarter of comparable models. Ling-2.6-1T is available on Hugging Face and supports inference frameworks including SGLang and vLLM.

Key info

Input
Output
Features
Context window
262K
Max output
33K

Benchmarks

Independent benchmark scores — composite indices for reasoning, coding, and math, plus individual eval scores where available.

Intelligence Index17.0
Reasoning & knowledge
GPQA Diamond
75%
Humanity's Last Exam
9%
Long-context reasoning
42%
Agentic & tool use
Terminal-Bench Hard
31%
τ²-Bench Telecom
90%
Math & instruction following
IFBench
57%

Available routes

No routes currently available — Ling-2.6-1T isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from InclusionAI

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation