Qwen3 8B

by Alibaba

Qwen3 8B is a dense 8.2 billion parameter model built on a decoder-only transformer with Grouped Query Attention for inference efficiency. It supports a 128K context window after length extension, with a 32K native window. Open-weight and small enough for local deployment, it offers an integrated thinking mode for complex logical problems, mathematics, and coding, alongside a non-thinking mode optimized for general dialogue and faster responses. FP8 quantization is openly available with minimal performance loss, keeping compatibility with common inference frameworks such as transformers, SGLang, and vLLM. It provides multilingual support across 100+ languages and dialects, with tool-use support for agent integrations.

Key info

Input
Output
Features
Context window
128K
Max output
20K

Available routes

No routes currently available — Qwen3 8B isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from Alibaba

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation
Qwen3 8B by Alibaba — not currently on Opper | Opper AI