Groq

LPU-based inference — sub-100ms first-token latencies on open-weight models.

Groq's LPU architecture delivers some of the lowest first-token latencies in the industry across Llama, GPT-OSS, Whisper, and Qwen models. US-hosted. The fast-path route for latency-sensitive applications — chat, agents, interactive tool use — where time-to-first-token is the critical metric.

1 route9 modelsUS🇺🇸 HQ United States
groq.com

Models on Groq

Every model we route through Groq. Compare residency, zero data retention, training posture and price at a glance, with full data-handling detail per route below.

ModelRegionZero data retentionTrainingContextInputOutput
USNo131K$0.80$4.00
USNo262K$0.60$3.00
USNo128K$0.15$0.60
USNo128K$0.07$0.30
USNo—$0.0019 / min
USNo—$0.0007 / min
USNo128K$0.07$0.30
USNo128K$0.59$0.79
USNo128K$0.05$0.08

Data handling per route

Groq hosts on 1 route. Each route has its own privacy posture, residency, and GDPR terms. Postures are maintained by Opper with a last-verification timestamp.

United States🇺🇸

Zero data retention: nothing is logged, held for abuse monitoring or used for training. No training on customer data. US; SCCs; DPA available.

Zero data retention
Yes. Nothing is logged, held for abuse monitoring or used for training.
Training
No training on customer data.
Logging
None
Abuse monitoring
No classifier
Caching
Content cached for replay
GDPR DPA
DPA available
Transfer mechanism
SCCs

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation