Groq

LPU-based inference — sub-100ms first-token latencies on open-weight models.

Groq's LPU architecture delivers some of the lowest first-token latencies in the industry across Llama, GPT-OSS, Whisper, and Qwen models. US-hosted. The fast-path route for latency-sensitive applications — chat, agents, interactive tool use — where time-to-first-token is the critical metric.

1 route8 modelsUS🇺🇸 HQ United States
groq.com

Models on Groq

Every model we route through Groq. Compare residency, ZDR, training posture, and price at a glance — full data-handling detail per route below.

ModelRegionZero data retentionTrainingContextInputOutput
USNot by defaultNo131K$0.80$4.00
USNot by defaultNo128K$0.15$0.60
USNot by defaultNo128K$0.07$0.30
USNot by defaultNo$0.0019 / min
USNot by defaultNo$0.0007 / min
USNot by defaultNo128K$0.07$0.30
USNot by defaultNo128K$0.59$0.79
USNot by defaultNo128K$0.05$0.08

Data handling per route

Groq hosts on 1 route. Each route has its own privacy posture, residency, and GDPR terms. Postures are maintained by Opper with a last-verification timestamp.

United States🇺🇸

Zero data retention is not on by default — abuse monitoring, 30-day retention. No training on customer data. US; SCCs; DPA available.

Zero data retention
Not by default — abuse monitoring, 30-day retention. Derived from the logging and moderation facts.
Training
No training on customer data.
Logging
Abuse monitoring (30-day retention)
Moderation
Not established
Caching
Not established
Subprocessor access
Not established
GDPR DPA
DPA available
Transfer mechanism
SCCs

Start building with 700+ models

One API key. Every major provider. Up and running in minutes.

Get startedView Documentation