DeepSeek R1 Distill Qwen 14B
DeepSeek R1 Distill Qwen 14B is a 14-billion parameter model distilled from DeepSeek R1, combining a Qwen base architecture with reasoning capabilities transferred from the larger parent. Released in January 2025 under the MIT license, it reaches 69.7% on AIME 2024 and 93.9% on MATH-500, demonstrating effective knowledge distillation in the 14B range. With a 32K token context window, the model offers a compact option for developers seeking reasoning capability without the requirements of 70B and larger variants. The distillation transfers chain-of-thought patterns from the full R1 model into a more efficient architecture suited to inference on modern accelerators. DeepSeek R1 Distill Qwen 14B bridges lightweight models and reasoning-grade performance, fitting resource-constrained deployments and applications needing moderate reasoning alongside efficient inference.
Key info
Benchmarks
Independent benchmark scores — composite indices for reasoning, coding, and math, plus individual eval scores where available.
Available routes
No routes currently available — DeepSeek R1 Distill Qwen 14B isn't routed through the Opper gateway right now. It may return.
Contact us about this model →