Qwen3 30B A3B
Qwen3-30B-A3B-FP8 is the FP8-quantized build of Alibaba's April 2025 MoE model, using fine-grained quantization with a 128 block size to reduce the memory footprint while keeping behavior close to the original BF16 weights. It retains the architecture of the base model, with 30.5 billion total parameters and 3.3 billion activated per token, the switchable thinking and non-thinking modes for balancing reasoning and speed, and context covering around 40K tokens, extending to 131K via YaRN. It handles mathematics, coding, logical reasoning, and multilingual tasks across 100+ languages, with tool calling for agentic applications. The FP8 format suits memory-constrained deployments without giving up reasoning quality or instruction following. Open-weight under Apache 2.0.
Key info
Available routes
No routes currently available — Qwen3 30B A3B isn't routed through the Opper gateway right now. It may return.
Contact us about this model →