Qwen3 32B
Qwen3-32B-FP8 is the FP8-quantized build of the dense 32.8 billion parameter model, using fine-grained quantization with a 128 block size to trim the memory footprint while keeping behavior close to the original BF16 weights. It keeps the full 32.8 billion parameters and the switchable thinking and non-thinking modes, with strong capability in mathematical reasoning, coding, logical inference, and instruction following across 100+ languages. Native context covers around 40K tokens, extending to 131K via YaRN, for long documents and extended reasoning. The dense architecture with FP8 quantization gives predictable inference behavior for memory-constrained GPU deployments, since every parameter activates per token for consistent latency and straightforward scaling. Open-weight under Apache 2.0.
Key info
Available routes
No routes currently available — Qwen3 32B isn't routed through the Opper gateway right now. It may return.
Contact us about this model →