Ling-2.6-1T
Ling-2.6-1T is InclusionAI's trillion-parameter flagship, built on a sparse Mixture-of-Experts architecture with roughly 63 billion active parameters per token and a 262K context window. Released in April 2026 under the MIT license, this open-weight model targets complex reasoning, coding, and multi-step agent execution through a fast-thinking approach that suppresses verbose chain-of-thought while holding intelligence high. Tuned for real-world agentic workflows, Ling-2.6-1T posts strong results on execution-heavy benchmarks including AIME26, SWE-Bench Verified (around 72%), BFCL-V4, and PinchBench, with particular strength in coding agents and tool-use scenarios. The architecture combines hybrid linear attention with Multi-Head Latent Attention and Lightning Linear components to speed inference and shrink memory use on long contexts. The fast-thinking reward strategy reduces reliance on long reasoning traces while preserving capability, cutting token cost to roughly a quarter of comparable models. Ling-2.6-1T is available on Hugging Face and supports inference frameworks including SGLang and vLLM.
Key info
Benchmarks
Independent benchmark scores — composite indices for reasoning, coding, and math, plus individual eval scores where available.
Available routes
No routes currently available — Ling-2.6-1T isn't routed through the Opper gateway right now. It may return.
Contact us about this model →