The 10 best LLM gateways in 2026
By Felix Wunderlich -
Updated August 3, 2026. We revisit this comparison as the market moves.
Picking an LLM gateway in 2026, an AI gateway as many vendors now brand it, is harder than it should be, because the market itself keeps moving: two of the best-known products were acquired this year, one was archived outright, and the survivors compete on very different terms, from free edge proxies to enterprise platforms starting at $499 a month. This comparison ranks the top ten LLM gateways we would actually shortlist, with pricing, data-residency posture and market status verified against each vendor's own pages as of August 2026, linked throughout so you can check every claim.
One disclosure up front: Opper is our product, so read that entry knowing where we stand. Every number in this article, ours included, comes from public pricing pages and primary announcements.
Quick picks
- Best for production teams in Europe: Opper, 500+ models on one OpenAI-compatible API with an EU-hosted control plane
- Biggest ecosystem: OpenRouter, 400+ models from 70+ providers and the largest developer community
- Best self-hosted: LiteLLM, the MIT-licensed default when traffic must stay in your VPC
- Best free tier: Cloudflare AI Gateway, core features cost nothing
- Best for the Vercel stack: Vercel AI Gateway, zero markup and a stated no-retention policy
- Best for air-gapped enterprise: TrueFoundry, VPC, on-prem and air-gapped deployment as first-class options
The best LLM gateways compared
| Gateway | Best for | Pricing | Open source | Data residency |
|---|---|---|---|---|
| Opper | Production teams in Europe | No markup on inference, 3% fee on credit purchases, BYOK free | No (hosted) | EU-hosted control plane, requests pinnable to EU regions |
| OpenRouter | Ecosystem breadth | No markup, 5.5% fee on credit purchases | No (hosted) | Per-provider policies, EU processing on enterprise only |
| LiteLLM | Self-hosting | Free (MIT), enterprise tier unpriced publicly | Yes (MIT) | Your infrastructure |
| Vercel AI Gateway | Vercel-stack teams | Zero markup, paid privacy add-ons | No (hosted) | Stated zero data retention, no published EU region |
| Cloudflare AI Gateway | Free edge caching | Core free, 5% fee on unified-billing credits | No (hosted) | Logging on by default, residency undocumented |
| Portkey | Gateway plus observability | Free tier, then $49/month | Gateway core | VPC option on enterprise |
| Helicone | Budget observability | Free tier, then $79/month | Yes (Apache-2.0) | EU or US region on its cloud |
| Kong AI Gateway | Kong-based enterprises | 5-model cap on Plus tier, enterprise custom | Core plugins | Data planes in your infra, EU control-plane geo |
| Bifrost | Raw performance, self-hosted | Free (Apache-2.0), enterprise unpriced | Yes (Apache-2.0) | Your infrastructure |
| TrueFoundry | Regulated enterprise | Free tier, then $499/month | No | VPC, on-prem, air-gapped |
The market moved hard in 2026
Any roundup written last year is already out of date. Palo Alto Networks completed its acquisition of Portkey on May 29, 2026, making it the core AI gateway of Prisma AIRS. Mintlify acquired Helicone in March and placed the product in what its announcement calls "maintenance mode", with security updates, bug fixes and new models still shipping. TensorZero was archived in June, with its team returning remaining seed capital to investors. OpenRouter raised a $113M Series B in May, and by late July was reportedly in talks to be acquired by Stripe in a deal reported at around $10 billion, though nothing has been announced and the talks could still collapse. Consolidation is real, so who owns your gateway, and whether it is still actively developed, now belongs on the evaluation checklist next to price and latency.
1. Opper
Opper is the gateway on this list built for teams that need European hosting without giving up model choice. It puts 500+ models from OpenAI, Anthropic, Google, Mistral, xAI and others behind one API with drop-in compatible endpoints for the OpenAI, Anthropic and Gemini SDKs, so switching models, or entire providers, is a one-line change. Fallbacks are built in, and the same model can be wired across multiple upstreams, for example Claude Sonnet on Anthropic direct plus AWS Bedrock, with automatic retry typically completing in about 180ms.
The residency story is the differentiator: the control plane runs in the EU, traces and logs stay in EU storage, and requests can be pinned to EU regions on AWS Bedrock, Azure, GCP or Berget AI. Tracing and logging are free and on by default, with per-call cost, latency and token accounting, and structured outputs and evaluations built in rather than sold as a separate product. Pricing is pay-as-you-go with no markup on inference and a 3% fee on credit purchases, and BYOK is free. Uptime is published on a public status page, and we run measured latency benchmarks rather than quoting synthetic overhead numbers.
One honest con: self-hosting is available on enterprise plans only, so smaller teams run on the hosted gateway. If you want the full comparison against the biggest name here, we maintain one at Opper vs OpenRouter.
2. OpenRouter
OpenRouter is the reference point among hosted gateways, with 400+ models from 70+ providers behind one API and the largest developer community in the category, and after its $113M Series B it is also the best-funded independent. As of late July 2026 it is reportedly in acquisition talks with Stripe, though no deal has been announced. Inference is passed through at provider rates with no markup, and the business model sits on credit purchases instead: a 5.5% fee with a $0.80 minimum on card payments, 5% via crypto. BYOK is free for the first million requests each month, then carries a 5% fee.
The trade-off is privacy by patchwork: each upstream provider has its own retention policy, which OpenRouter surfaces in a per-provider table where some entries read "Zero retention" and others "Unknown", and processing pinned to the EU is an enterprise-only feature. OpenRouter states it does not use your inputs or outputs for model training, but that is a training claim rather than a no-retention claim. For catalog breadth and community momentum it is the reference point; for compliance-sensitive workloads you will be reading provider tables carefully.
3. LiteLLM
LiteLLM is the default answer when the requirement is "our traffic never leaves our infrastructure". The MIT-licensed proxy covers 140+ provider integrations with virtual keys, budgets, load balancing and guardrails, all in the free open-source core, and self-hosting means data residency is simply your own infrastructure. It is the most widely deployed open-source gateway and the one most roundups measure everything else against.
The costs are operational rather than financial: you run it, scale it and patch it yourself, and the enterprise tier with SSO, RBAC and audit logs has no public pricing, only a sales conversation. Teams that want LiteLLM's flexibility without operating it can compare the hosted route in our Opper vs LiteLLM breakdown.
4. Vercel AI Gateway
Vercel AI Gateway routes to hundreds of models with zero markup on tokens, including BYOK, and its stated data posture is among the strongest of the hosted options: a zero-data-retention policy under which prompts and outputs are deleted immediately after the request completes. For teams already deploying on Vercel it is the path of least resistance.
Two caveats. The privacy features beyond the baseline are metered paid add-ons, with team-wide ZDR-only routing at $0.10 per thousand requests on Pro and Enterprise plans, and there is no published EU hosting option, so residency-driven buyers get a strong retention story but not a location story.
5. Cloudflare AI Gateway
Cloudflare AI Gateway is the free option that earns its place: analytics, caching and rate limiting cost nothing, it sits on Cloudflare's edge, and unified billing carries a 5% fee on credits purchased with no markup on provider rates. As a caching and control layer in front of providers you already pay, it is hard to argue with the price.
Know what you are getting, though: logs, including request and response data, are enabled by default for each gateway unless you opt out in settings, free-plan storage caps at 100,000 logs, and the docs state no retention period or storage geography, so there is no documented way to keep gateway logs in Europe.
6. Portkey
Portkey built one of the more complete gateway-plus-observability suites, with an open-source gateway core and a cheap on-ramp: a free tier with 10,000 logged requests a month, then a $49/month Production tier with 100,000 logs and 30-day retention, up to enterprise deals with VPC and private-cloud deployment. It has long been the mid-market pick for teams that want guardrails, caching and logging in one product.
The 2026 development to factor in: Portkey is now part of Palo Alto Networks, positioned as the core AI gateway of Prisma AIRS. The standalone product remains on sale today, and what the acquisition means for its independent roadmap is not yet stated either way. Note also that Portkey is observability-first, so logging requests is the core function, with retention set by tier.
7. Helicone
Helicone remains the budget-friendly way to get observability and a gateway together: Apache-2.0 open source, a free tier with 10,000 requests, a $79/month Pro tier, and a cloud AI gateway with 100+ models at 0% markup. Notably for this list, its cloud offers a choice of EU or US regions, which most US-hosted rivals do not.
The caveat is its status: after Mintlify's acquisition in March 2026, the product is officially in "maintenance mode", in the announcement's own words, with security updates, bug fixes and new models continuing but no new feature roadmap. The repository remains genuinely active as of July 2026, so it is a stable tool to run, just not a bet on future capabilities.
8. Kong AI Gateway
Kong AI Gateway extends the Kong API platform to LLM traffic, with multi-provider routing, PII sanitization, semantic caching and MCP governance. Its architecture is a genuine residency advantage: data planes, the part that sees your prompts, run in your own infrastructure, while Konnect control planes are available in six geos including the EU, with only authentication, billing and usage shared across geos.
Pricing is the friction: the affordable Konnect Plus tier caps AI Gateway at five unique LLM models with one million API requests a month included and $200 per additional million, and removing the model cap means a custom-priced enterprise contract. For organizations already running Kong it is a natural extension; as a standalone AI gateway choice it is heavy.
9. Bifrost
Bifrost, from the team at Maxim AI, is the performance-focused newcomer: an Apache-2.0 Go gateway with MCP support and vendor-run benchmarks claiming under 100 microseconds of overhead at 5,000 RPS, routing to 1,000+ models. The repository is active and growing fast.
It is also the youngest project on this list, its headline performance numbers are the vendor's own rather than independently verified, and enterprise pricing is not public. A strong candidate for teams that want a fast self-hosted gateway and are comfortable with early-stage software.
10. TrueFoundry
TrueFoundry aims squarely at regulated enterprises: VPC, on-prem and air-gapped deployment are first-class options rather than special requests, and the platform spans model serving and MLOps beyond the gateway itself, a scope it broadened by acquiring Seldon AI in June 2026. The free Developer tier covers 50,000 requests a month.
The entry price for production is the sticking point: the Pro tier starts at $499/month, an order of magnitude above Portkey's $49 or Helicone's $79, which makes sense for enterprise procurement and less sense for a team of five.
If your prompts have to stay in Europe
This is the dimension most roundups skip, and it is where the field separates. GDPR applies to your AI traffic whenever personal data flows through prompts, wherever the gateway vendor is established, so the practical question is what each product lets you control. Fully self-hosted gateways, LiteLLM and Bifrost here, inherit whatever infrastructure you run them on. Kong keeps prompt-carrying data planes in your infrastructure with an EU control-plane geo. Helicone offers an EU region on its cloud. OpenRouter processes in the EU only on enterprise contracts, Vercel publishes a strong retention policy but no EU region, and Cloudflare documents neither retention period nor storage location for gateway logs.
Opper is built EU-first: the control plane runs in the EU, traces and logs stay in EU storage, and requests can be pinned to EU-hosted model deployments on AWS Bedrock, Azure, GCP or Berget AI, with a DPA and published sub-processor list. For the compliance-driven evaluation in full, see our AI compliance overview.
LLM gateway FAQ
How do I choose between LLM gateways?+
Compare on four axes: the pricing model (no markup with a fee on credit purchases, per-logged-request tiers, or seat pricing), where your data flows (hosting region, logging defaults, retention), open source versus hosted (operational control versus nothing to run), and whether the product is still actively developed, which 2026's acquisitions made a real question. The comparison table above scores all ten gateways on these dimensions, and the fundamentals are covered in our LLM gateway guide.
Which LLM gateway is best in 2026?+
It depends on what you optimize for. Opper is built for production teams that want European hosting with observability included, OpenRouter has the largest catalog, LiteLLM is the self-hosted open-source standard, Cloudflare is free as an edge layer, and TrueFoundry targets air-gapped enterprise deployments. The comparison table above maps each gateway to the buyer it fits best.
Does an LLM gateway add latency?+
Less than most teams assume. In our measured benchmark of time to first token, gateway routes landed within tens of milliseconds of calling OpenAI directly, and some were faster. The full methodology and confidence intervals are in our LLM router latency benchmark.
Which is the best open-source LLM gateway?+
LiteLLM is the established default, MIT-licensed with 140+ provider integrations, if you have the team to operate it. Bifrost is the performance-focused newer alternative in Go, and Helicone pairs an Apache-2.0 gateway with observability but has been in maintenance mode since its acquisition. If you want the flexibility without running infrastructure, hosted gateways with BYOK support such as Opper keep your own provider keys while operating the gateway for you.
Which LLM gateways keep data in Europe?+
Opper runs its control plane in the EU, keeps traces and logs in EU storage, and can pin requests to EU-hosted model deployments. Helicone offers an EU region on its cloud, Kong keeps prompt-carrying data planes in your own infrastructure, and self-hosted gateways inherit your infrastructure. OpenRouter offers EU processing on enterprise contracts only, and Vercel and Cloudflare publish no EU hosting option for their gateways.
What happened to Portkey, Helicone and TensorZero?+
2026 was the year of consolidation. Palo Alto Networks completed its acquisition of Portkey in May, making it the core AI gateway of Prisma AIRS. Mintlify acquired Helicone in March and placed it in maintenance mode, with security updates and new models still shipping. TensorZero was archived on GitHub in June, with the team returning remaining capital to investors.
Try the European pick
If your shortlist ends at "largest catalog" the answer is OpenRouter, and if it ends at "we run everything ourselves" it is LiteLLM. If it ends at production traffic with European hosting, built-in observability and no markup on inference, create a free Opper account and make your first call in minutes, or browse the 500+ models it routes to first.