Llama 3.1 8B
Llama 3.1 8B is the successor to Llama 3 8B, an 8-billion-parameter instruction-tuned model with a 128,000-token context window, a large jump from Llama 3's 8,192-token limit. Released on July 23, 2024, it supports eight languages including English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai. Trained on over 15 trillion tokens from publicly available sources with a December 2023 knowledge cutoff, Llama 3.1 8B combines supervised fine-tuning with reinforcement learning from human feedback for improved reasoning, coding, and instruction-following. It uses Grouped-Query Attention for inference efficiency and is suited to assistant-style chat, synthetic data generation, and distillation into smaller models. Compared to Llama 3 8B, the 3.1 release offers a roughly 16x longer context window for extended documents and conversations, broader multilingual coverage, and stronger reasoning and coding. Its balance of capability and efficiency makes it a good fit for applications that need long context without a large compute budget.
Key info
Available routes
No routes currently available — Llama 3.1 8B isn't routed through the Opper gateway right now. It may return.
Contact us about this model →