DeepSeek R1 Distill Qwen 14B

by DeepSeek

DeepSeek R1 Distill Qwen 14B is a 14-billion parameter model distilled from DeepSeek R1, combining a Qwen base architecture with reasoning capabilities transferred from the larger parent. Released in January 2025 under the MIT license, it reaches 69.7% on AIME 2024 and 93.9% on MATH-500, demonstrating effective knowledge distillation in the 14B range. With a 32K token context window, the model offers a compact option for developers seeking reasoning capability without the requirements of 70B and larger variants. The distillation transfers chain-of-thought patterns from the full R1 model into a more efficient architecture suited to inference on modern accelerators. DeepSeek R1 Distill Qwen 14B bridges lightweight models and reasoning-grade performance, fitting resource-constrained deployments and applications needing moderate reasoning alongside efficient inference.

Key info

Input
Output
Features
Context window
33K
Max output
16K

Benchmarks

Independent benchmark scores — composite indices for reasoning, coding, and math, plus individual eval scores where available.

Global rank#393 of 596 LLMs
TierEfficient
Output speed0 tok/s
First token0.00s
Intelligence Index9.7
Math Index55.7
Reasoning & knowledge
MMLU-Pro
74%
GPQA Diamond
48%
Humanity's Last Exam
4%
Long-context reasoning
7%
Coding
LiveCodeBench
38%
SciCode
24%
Math & instruction following
AIME 2025
56%
IFBench
22%

Available routes

No routes currently available — DeepSeek R1 Distill Qwen 14B isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from DeepSeek

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation
DeepSeek R1 Distill Qwen 14B by DeepSeek — not currently on Opper | Opper AI