Wafer

Tuned serving for open-weight models — GLM, Kimi and DeepSeek.

Wafer serves open-weight models on a stack it tunes per workload, profiling traffic to pick the serving engine, kernels and hardware for each deployment. US-hosted. The route to pick when you want tuned throughput on GLM, Kimi or DeepSeek and US hosting is acceptable.

1 route5 modelsUS🇺🇸 HQ United States
wafer.ai

Models on Wafer

Every model we route through Wafer. Compare residency, ZDR, training posture, and price at a glance — full data-handling detail per route below.

ModelRegionZero data retentionTrainingContextInputOutput
USZero data retentionNo1M$3.00$12.75
USZero data retentionNo1M$1.19$4.40
USZero data retentionNo1M$0.10$0.35
USZero data retentionNo1M$1.26$3.96
USZero data retentionNo1M$0.10$0.25

Data handling per route

Wafer hosts on 1 route. Each route has its own privacy posture, residency, and GDPR terms. Postures are maintained by Opper with a last-verification timestamp.

United States🇺🇸

Zero data retention is on by default on Pay-as-you-go — no action required. No training on customer data. US; SCCs; DPA available.

Zero data retention
On by default on Pay-as-you-go. Derived from the logging and moderation facts.
Training
No training on customer data.
Logging
None
Moderation
Not established
Caching
Not established
Subprocessor access
Not established
GDPR DPA
DPA available
Transfer mechanism
SCCs

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation