Llama 3.1 Nemotron Ultra 253B v1
Llama 3.1 Nemotron Ultra 253B v1 is a reasoning model developed by NVIDIA, derived from Meta's Llama 3.1 405B Instruct using Neural Architecture Search to reduce its memory footprint. With 253 billion parameters and a 131,072-token context window, it targets reasoning-heavy workloads and research applications. The model is post-trained for reasoning, structured output generation, and multi-step problem decomposition, with support for chat preferences, retrieval-augmented generation, and tool calling. The architecture search aims to make a model of this scale more efficient to serve. Its open-weight release supports research, customization, and on-premise deployment for organizations with sufficient compute. Its reasoning focus places it above standard 70B models on tasks involving logic and mathematics.
Key info
Available routes
No routes currently available — Llama 3.1 Nemotron Ultra 253B v1 isn't routed through the Opper gateway right now. It may return.
Contact us about this model →