Llama 3.1 Nemotron Ultra 253B v1

by Meta

Llama 3.1 Nemotron Ultra 253B v1 is a reasoning model developed by NVIDIA, derived from Meta's Llama 3.1 405B Instruct using Neural Architecture Search to reduce its memory footprint. With 253 billion parameters and a 131,072-token context window, it targets reasoning-heavy workloads and research applications. The model is post-trained for reasoning, structured output generation, and multi-step problem decomposition, with support for chat preferences, retrieval-augmented generation, and tool calling. The architecture search aims to make a model of this scale more efficient to serve. Its open-weight release supports research, customization, and on-premise deployment for organizations with sufficient compute. Its reasoning focus places it above standard 70B models on tasks involving logic and mathematics.

Key info

Input
Output
Features
Context window
131K
Max output
8K

Available routes

No routes currently available — Llama 3.1 Nemotron Ultra 253B v1 isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from Meta

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation