GLM-5.2 Fast

by Z.ai

GLM-5.2 Fast runs Z.ai's GLM-5.2 on a throughput-optimized path that delivers two to three times standard generation speed at the same quality bar. The underlying model, released in June 2026 under an MIT license, is a 744B-parameter mixture-of-experts design that debuted as the top-ranked open-weight model on the Artificial Analysis Intelligence Index. Agentic coding and terminal work are where it shines, with strong showings on SWE-bench Pro and Terminal-Bench, and its context window stretches past one million tokens with output up to 131K. The Fast tier keeps those strengths while cutting time to completion, a good match for interactive coding agents, large-codebase analysis, and reasoning-heavy pipelines where latency matters.

Key info

Input
Output
Features
Context window
1M
Max output
131K
Input price
$2.10 /1M
Output price
$6.60 /1M
  • US residency available
  • Zero data retention on pay-as-you-go
  • No training by default
  • GDPR DPA available

Available routes

GLM-5.2 Fast runs on 2 different routes through the Opper gateway. Compare residency, ZDR, and training posture at a glance β€” full data-handling detail per route below.

ProviderRegionZero data retentionTrainingInputOutput
MultiEnterpriseNo$2.10$6.60
USZero data retentionNo$2.10$6.60

Uptime and availability

One of GLM-5.2 Fast's 2 routes is on the monitoring board, so the figure below is that provider's availability rather than the model's.

100%route uptime, last 30 days
Its other routes aren't measured here yet.

Name a second model in the same request and the gateway tries it on retriable errors, so a busy hour never has to reach your users. Set up a fallback chain.

Measured over the last 30 days from each provider's official status feed via StatusGator. Refreshed hourly. 1 further route is not on the board yet. See uptime for every provider Opper monitors.

Data handling per route

Each route hosting GLM-5.2 Fast has its own privacy posture, residency, and GDPR terms. Postures are maintained by Opper with a last-verification timestamp.

Fireworks β€” United StatesπŸ‡ΊπŸ‡Έ

Zero data retention is available via Opper Enterprise contract. No training on customer data. GLOBAL; SCCs; DPA available.

Zero data retention
Available via Opper Enterprise contract.
Training
No training on customer data.
Logging
None
Third-party access
None disclosed
GDPR DPA
DPA available
Transfer mechanism
SCCs

Wafer β€” United StatesπŸ‡ΊπŸ‡Έ

Zero data retention is on by default on Pay-as-you-go β€” no action required. No training on customer data. US; unknown; DPA available.

Zero data retention
On by default on Pay-as-you-go.
Training
No training on customer data.
Logging
None
Third-party access
Provider may share with subprocessors / partners
GDPR DPA
DPA available
Transfer mechanism
unknown

Get started

Call GLM-5.2 Fast through the Opper gateway with one API key. Let your coding agent set it up, or call it directly β€” Opper is drop-in compatible with the OpenAI, Anthropic, and Google AI SDKs.

Set it up with your agent

Copy this and paste it into a coding agent like Claude Code, Cursor or Codex and it'll wire up Opper for you.

Or call it directly

import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.OPPER_API_KEY,
baseURL: "https://api.opper.ai/v3/compat",
});
const completion = await client.chat.completions.create({
model: "fireworks/glm-5p2-fast",
messages: [{ role: "user", content: "Hello" }],
});
console.log(completion.choices[0].message.content);

Compare GLM-5.2 Fast with…

Side-by-side on privacy, EU hosting, pricing, and benchmarks.

Other models from Z.ai

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation