DeepSeek R1 Distill LLama 70B

by DeepSeek

DeepSeek R1 Distill Llama 70B is a 70-billion parameter model based on Llama 3.3 70B Instruct, fine-tuned using outputs from DeepSeek R1 to transfer reasoning patterns and chain-of-thought capabilities. Released in January 2025, it performs strongly on complex mathematical and reasoning tasks, reaching around 70% on AIME 2024 and 94.5% on MATH-500. The distillation approach lets a dense 70B model capture reasoning behaviors from the much larger 671B R1 parent, making it efficient for organizations that prefer dense architectures over Mixture-of-Experts designs. It significantly outperforms the base Llama 3.3 70B model on reasoning benchmarks. Well suited to educational tools, research applications, and mathematical problem-solving, DeepSeek R1 Distill Llama 70B pairs a familiar dense architecture with reasoning-grade performance. As a Llama-derived model it inherits the Llama 3.3 license.

Key info

Input
Output
Features
Context window
131K
Max output
131K

Benchmarks

Independent benchmark scores — composite indices for reasoning, coding, and math, plus individual eval scores where available.

Intelligence Index7.9
Math Index53.7
Reasoning & knowledge
MMLU-Pro
80%
GPQA Diamond
40%
Humanity's Last Exam
5%
Long-context reasoning
10%
Coding
LiveCodeBench
27%
Agentic & tool use
Terminal-Bench Hard
2%
τ²-Bench Telecom
22%
Math & instruction following
AIME 2025
54%
IFBench
28%

Available routes

No routes currently available — DeepSeek R1 Distill LLama 70B isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from DeepSeek

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation