Cerebras

Wafer-scale chip inference — among the fastest tokens-per-second in the industry.

Cerebras runs inference on their custom Wafer-Scale Engine, delivering some of the highest tokens-per-second throughput available on Llama and GPT-OSS class models. US-hosted. The default pick when output throughput dominates the workload and US hosting is acceptable.

1 route2 modelsUS🇺🇸 HQ United States
cerebras.ai

Models on Cerebras

Every model we route through Cerebras. Compare residency, ZDR, training posture, and price at a glance — full data-handling detail per route below.

ModelRegionZero data retentionTrainingContextInputOutput
USZero data retentionNo128K$2.25$2.75
USZero data retentionNo131K$0.35$0.75

Data handling per route

Cerebras hosts on 1 route. Each route has its own privacy posture, residency, and GDPR terms. Postures are maintained by Opper with a last-verification timestamp.

United States🇺🇸

Zero data retention is on by default on Pay-as-you-go — no action required. No training on customer data. US; SCCs; DPA available.

Zero data retention
On by default on Pay-as-you-go. Derived from the logging and moderation facts.
Training
No training on customer data.
Logging
None
Moderation
Not established
Caching
Not established
Subprocessor access
Not established
GDPR DPA
DPA available
Transfer mechanism
SCCs

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation