DeepSeek R1 0528 Qwen3 8B

by DeepSeek

DeepSeek R1 0528 Qwen3 8B is an 8-billion parameter model created by distilling the chain-of-thought of the larger DeepSeek R1 0528 into the Qwen3 8B base architecture, achieving state-of-the-art performance among open-source models on AIME 2024 while staying efficient for inference. The distillation transfers reasoning patterns from the 671B-scale parent, enabling strong reasoning at a fraction of the size and computational cost. Despite having only 8 billion parameters, it scores 86.0% on AIME 2024, surpassing Qwen3 8B by 10 percentage points and matching the performance of Qwen3-235B-thinking, demonstrating the effectiveness of distillation from a reasoning-specialized parent. With a 128K token context window and text-based reasoning focus, DeepSeek R1 0528 Qwen3 8B offers an attractive balance of reasoning capability and efficiency for developers seeking advanced reasoning at a smaller scale. It is released under the MIT license.

Key info

Input
Output
Features
Context window
128K
Max output
32K

Available routes

No routes currently available — DeepSeek R1 0528 Qwen3 8B isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from DeepSeek

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation
DeepSeek R1 0528 Qwen3 8B by DeepSeek — not currently on Opper | Opper AI