Cerebras

Wafer-scale chip inference — among the fastest tokens-per-second in the industry.

Cerebras runs inference on their custom Wafer-Scale Engine, delivering some of the highest tokens-per-second throughput available on Llama and GPT-OSS class models. US-hosted. The default pick when output throughput dominates the workload and US hosting is acceptable.

1 route2 modelsUS🇺🇸 HQ United States
cerebras.ai

Models on Cerebras

Every model we route through Cerebras. Compare residency, zero data retention, training posture and price at a glance, with full data-handling detail per route below.

ModelRegionZero data retentionTrainingContextInputOutput
USNo128K$2.25$2.75
USNo131K$0.35$0.75

Data handling per route

Cerebras hosts on 1 route. Each route has its own privacy posture, residency, and GDPR terms. Postures are maintained by Opper with a last-verification timestamp.

United States🇺🇸

Zero data retention: nothing is logged, held for abuse monitoring or used for training. No training on customer data. US; SCCs; DPA available.

Zero data retention
Yes. Nothing is logged, held for abuse monitoring or used for training.
Training
No training on customer data.
Logging
None
Abuse monitoring
On by default, keeps nothing
Caching
Content cached for replay
GDPR DPA
DPA available
Transfer mechanism
SCCs

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation