Qwen3 32B

by Alibaba

Qwen3-32B-FP8 is the FP8-quantized build of the dense 32.8 billion parameter model, using fine-grained quantization with a 128 block size to trim the memory footprint while keeping behavior close to the original BF16 weights. It keeps the full 32.8 billion parameters and the switchable thinking and non-thinking modes, with strong capability in mathematical reasoning, coding, logical inference, and instruction following across 100+ languages. Native context covers around 40K tokens, extending to 131K via YaRN, for long documents and extended reasoning. The dense architecture with FP8 quantization gives predictable inference behavior for memory-constrained GPU deployments, since every parameter activates per token for consistent latency and straightforward scaling. Open-weight under Apache 2.0.

Key info

Input
Output
Features
Context window
41K
Max output
20K

Available routes

No routes currently available — Qwen3 32B isn't routed through the Opper gateway right now. It may return.

Contact us about this model →

Available models from Alibaba

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation
Qwen3 32B by Alibaba — not currently on Opper | Opper AI