Qwen3 8B
Qwen3 8B is a dense 8.2 billion parameter model built on a decoder-only transformer with Grouped Query Attention for inference efficiency. It supports a 128K context window after length extension, with a 32K native window. Open-weight and small enough for local deployment, it offers an integrated thinking mode for complex logical problems, mathematics, and coding, alongside a non-thinking mode optimized for general dialogue and faster responses. FP8 quantization is openly available with minimal performance loss, keeping compatibility with common inference frameworks such as transformers, SGLang, and vLLM. It provides multilingual support across 100+ languages and dialects, with tool-use support for agent integrations.
Key info
Available routes
No routes currently available — Qwen3 8B isn't routed through the Opper gateway right now. It may return.
Contact us about this model →