DeepSeek R1 Distill LLama 70B
DeepSeek R1 Distill Llama 70B is a 70-billion parameter model based on Llama 3.3 70B Instruct, fine-tuned using outputs from DeepSeek R1 to transfer reasoning patterns and chain-of-thought capabilities. Released in January 2025, it performs strongly on complex mathematical and reasoning tasks, reaching around 70% on AIME 2024 and 94.5% on MATH-500. The distillation approach lets a dense 70B model capture reasoning behaviors from the much larger 671B R1 parent, making it efficient for organizations that prefer dense architectures over Mixture-of-Experts designs. It significantly outperforms the base Llama 3.3 70B model on reasoning benchmarks. Well suited to educational tools, research applications, and mathematical problem-solving, DeepSeek R1 Distill Llama 70B pairs a familiar dense architecture with reasoning-grade performance. As a Llama-derived model it inherits the Llama 3.3 license.
Key info
Benchmarks
Independent benchmark scores — composite indices for reasoning, coding, and math, plus individual eval scores where available.
Available routes
No routes currently available — DeepSeek R1 Distill LLama 70B isn't routed through the Opper gateway right now. It may return.
Contact us about this model →