Llama 3.2 1B Instruct

by Meta

Llama 3.2 1B Instruct is Meta's ultra-compact model, released September 2024 and tuned for on-device, mobile, and embedded use where memory and compute are limited. With 1 billion parameters and a 131,000-token context window, it suits local-first inference without cloud dependencies. Despite its small size, the model supports structured output generation, enabling programmatic extraction and JSON formatting for downstream automation. It handles classification, sentiment analysis, intent detection, and simple text generation on modest hardware. Its open-weight design enables on-device fine-tuning and adaptation to specific tasks without connectivity. It is well suited to local applications, embedded systems, and mobile apps, and is often combined with retrieval augmentation to extend capability without scaling parameters.

Key info

Input
Output
Features
Context window
131K
Max output
32K

Benchmarks

Independent benchmark scores — composite indices for reasoning, coding, and math, plus individual eval scores where available.

Global rank#583 of 596 LLMs
TierEfficient
Output speed0 tok/s
First token0.00s
Intelligence Index1.0
Math Index0.0
Reasoning & knowledge
MMLU-Pro
20%
GPQA Diamond
20%
Humanity's Last Exam
6%
Long-context reasoning
6%
Coding
LiveCodeBench
2%
SciCode
2%
Agentic & tool use
Terminal-Bench Hard
0%
τ²-Bench Telecom
0%
Math & instruction following
AIME 2025
0%
IFBench
23%

Available routes

No routes currently available — Llama 3.2 1B Instruct isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from Meta

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation
Llama 3.2 1B Instruct by Meta — not currently on Opper | Opper AI