DeepSeek R1 0528 Qwen3 8B
DeepSeek R1 0528 Qwen3 8B is an 8-billion parameter model created by distilling the chain-of-thought of the larger DeepSeek R1 0528 into the Qwen3 8B base architecture, achieving state-of-the-art performance among open-source models on AIME 2024 while staying efficient for inference. The distillation transfers reasoning patterns from the 671B-scale parent, enabling strong reasoning at a fraction of the size and computational cost. Despite having only 8 billion parameters, it scores 86.0% on AIME 2024, surpassing Qwen3 8B by 10 percentage points and matching the performance of Qwen3-235B-thinking, demonstrating the effectiveness of distillation from a reasoning-specialized parent. With a 128K token context window and text-based reasoning focus, DeepSeek R1 0528 Qwen3 8B offers an attractive balance of reasoning capability and efficiency for developers seeking advanced reasoning at a smaller scale. It is released under the MIT license.
Key info
Available routes
No routes currently available — DeepSeek R1 0528 Qwen3 8B isn't routed through the Opper gateway right now. It may return.
Contact us about this model →