GLM-5.3-FlashX

by Z.ai

GLM-5.3-FlashX is a faster serving tier of Z.ai's GLM-5.3-Flash. It uses the same model, a mixture-of-experts design with 320B total and 18B active parameters, with infrastructure and inference optimizations that reach generation speeds of up to 200 tokens per second. Like GLM-5.3-Flash, it is the first native multimodal model in the GLM-5 series. It takes text, images, video and files as input and returns text, and its hybrid architecture combines sparse and linear attention to cut attention compute and KV-cache cost. It supports a 1M-token context window and up to 128K output tokens. FlashX suits cases where response time compounds: interactive coding agents, tool-calling loops, visual coding and user-facing assistants. Where raw speed matters less, GLM-5.3-Flash offers the same capabilities, and GLM-5.3 remains the series flagship.

Call GLM-5.3-FlashX on Opper with the OpenAI SDK. Sign up without a credit card. Get started

Key info

Input
Output
Features
Context window
1M
Max output
131K
Input price
$0.37 /1M$0.07 cache read
Output price
$1.25 /1M
  • No training by default
  • GDPR DPA available

Available routes

GLM-5.3-FlashX runs on 1 route through the Opper gateway. Compare residency, zero data retention and training posture at a glance, with full data-handling detail per route below.

ProviderRegionZero data retentionTrainingInputOutputCache read
MultiNo$0.37$1.25$0.07

Data handling per route

Each route hosting GLM-5.3-FlashX has its own privacy posture, residency, and GDPR terms. Postures are maintained by Opper with a last-verification timestamp.

Tencent Cloud — Multi-region

Content is retained: limited debug logs, 30-day retention. No training on customer data. GLOBAL; SCCs; DPA available.

Zero data retention
Retained: limited debug logs, 30-day retention.
Training
No training on customer data.
Logging
Limited debug logs (30-day retention)
Abuse monitoring
On by default, holds flagged content
Caching
Content cached for replay
GDPR DPA
DPA available
Transfer mechanism
SCCs

Get started

Call GLM-5.3-FlashX through the Opper gateway with one API key. Opper is drop-in compatible with the OpenAI, Anthropic and Google AI SDKs.

GLM-5.3-FlashX is a premium model on Opper. Sign up needs no credit card: you get an API key straight away and the free models work in the playground and the API. Add a card to use premium models, pay-as-you-go with no minimum.

Set it up with your agent

Copy this and paste it into a coding agent like Claude Code, Cursor or Codex and it'll wire up Opper for you.

Or call it directly

import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.OPPER_API_KEY,
baseURL: "https://api.opper.ai/v3/compat",
});
const completion = await client.chat.completions.create({
model: "tencent:sg/glm-5.3-flashx",
messages: [{ role: "user", content: "Hello" }],
});
console.log(completion.choices[0].message.content);

Use GLM-5.3-FlashX in your apps

Pick your tool to see how to run it on GLM-5.3-FlashX through Opper.

Compare GLM-5.3-FlashX with…

Side-by-side on privacy, EU hosting, pricing, and benchmarks.

Other models from Z.ai

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation