GLM-5.2 by Z.ai is a 744B open-weight (MIT) mixture-of-experts coding model with a 1M-token context, IndexShare sparse attention, and selectable thinking-effort levels.
Available on
Privacy
- Zero data retention
- Always-on
- Training
- No
- Region
- EU + US
Kilo Code is an open-source AI coding agent for VS Code. It combines autonomous editing, planning and architect modes, and an OpenAI-compatible provider so you can bring the model of your choice. Add Opper as that provider with a base URL and key, and you can run it on any of 700+ models on one gateway, with a choice of region.
Paste this into your coding agent (Claude Code, Cursor, Codex, and more) and it will set up and build with Opper for you.
GLM-5.2 runs on the routes below. Route only to EU providers when your data has to stay in Europe.
| Provider | Region | Zero data retention | Training | Input | Output |
|---|---|---|---|---|---|
| Berget | EU | Zero data retention | No | $1.62 | $5.10 |
| evroc | EU | Zero data retention | No | $1.45 | $5.80 |
| Inceptron | EU | Enterprise | No | $0.75 | $2.90 |
| Nebius | EU | Enterprise | No | $1.40 | $4.40 |
| Nextbit | EU | Enterprise | No | $1.40 | $4.40 |
| Regolo | EU | Zero data retention | No | $2.32 | $6.03 |
| Sference | EU | Zero data retention | No | $1.20 | $4.20 |
| TensorX | EU | Zero data retention | No | $1.50 | $4.50 |
| DeepInfra | US | Enterprise | No | $0.75 | $2.40 |
| Morph | US | Enterprise | No | $1.10 | $4.10 |
| Novita | US | Enterprise | No | $1.40 | $4.40 |
| Wafer | US | Zero data retention | No | $1.26 | $3.96 |
| Alibaba Cloud | Multi | Enterprise | No | $1.19 | $4.15 |
| BytePlus | Multi | Enterprise | No | $1.40 | $4.40 |
| Fireworks | Multi | Enterprise | No | $1.40 | $4.40 |
GLM-5.2 by Z.ai is a 744B open-weight (MIT) mixture-of-experts coding model with a 1M-token context, IndexShare sparse attention, and selectable thinking-effort levels.
Point Cursor's OpenAI base URL at Opper to run it on any of 700+ models.
OpenAI-compatibleCline's OpenAI-Compatible provider points straight at Opper's gateway.
OpenAI-compatibleSwitch Opper models mid-thread and see the tokens, cost, and latency behind each one, in a desktop IDE that runs locally.
Opper built in