AI Model Directory

Every model on the Opper gateway

Search by what matters. Zero data retention, residency, training and more.

CompareLeaderboard
by Anthropic

Claude Fable 5.1 is Anthropic's frontier model for agentic coding, knowledge work, and long-running problem solving, with a 1M token context and 128K output.

Available on

Pricing from$10.00/ 1M input$50.00/ 1M output$0.25/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, Multi
by OpenAI

GPT-6 Astra is OpenAI's flagship model for computer use, software engineering, and scientific work, with a 1.05M token context and 128K output.

Available on

Pricing from$10.00/ 1M input$50.00/ 1M output$1.00/ 1M cache read1.1Mcontext

Privacy

ZDR
Training
No
Region
Multi, US
by Anthropic

Claude Opus 5 is Anthropic's flagship for complex agentic coding and enterprise work, with a 1M token context window and adaptive thinking.

Available on

Pricing from$5.00/ 1M input$25.00/ 1M output$0.50/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, EU, Multi
by Anthropic

Claude Fable 5 by Anthropic, state-of-the-art generalist model with advanced vision and 1M token context for complex software engineering and reasoning.

Available on

Pricing from$10.00/ 1M input$50.00/ 1M output$1.00/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, Multi
by Meta

Muse Spark 1.3 is the latest of Meta's Muse Spark models, cutting tool calls and tokens versus Muse Spark 1.2 across a 1M token context with image and video input.

Available on

Pricing from$0.10/ 1M input$0.20/ 1M output$0.0020/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
Multi
by OpenAI

OpenAI GPT-6.1 Sol for coding, computer use, and complex professional work.

Available on

Pricing from$2.00/ 1M input$10.00/ 1M output$0.10/ 1M cache read1.1Mcontext

Privacy

ZDR
Training
No
Region
EU, Multi, US
by OpenAI

OpenAI GPT-6 Sol for complex professional work, coding, and agent workflows.

Available on

Pricing from$2.00/ 1M input$10.00/ 1M output$0.20/ 1M cache read1.1Mcontext

Privacy

ZDR
Training
No
Region
Multi, EU, US
by OpenAI

GPT-5.6 Sol tops OpenAI's GPT-5.6 family, the frontier tier for complex coding, deep research, and long-running agents with a 1M token context.

Available on

Pricing from$4.00/ 1M input$20.00/ 1M output$0.40/ 1M cache read1.1Mcontext

Privacy

ZDR
Training
No
Region
US, EU, Multi
by xAI

xAI's Grok 4.7 frontier model for coding, agentic tasks, and knowledge work, with reasoning, 500K context window, vision, and agentic tool use

Available on

Pricing from$2.00/ 1M input$6.00/ 1M output$0.50/ 1M cache read500Kcontext

Privacy

ZDR
Training
No
Region
Multi, US
by Z.ai

GLM-5.3 by Z.ai is a 744B open-weight coding and agentic model, post-trained on the GLM-5.2 base with a 1M-token context and low, high, and max reasoning effort levels.

Available on

Pricing from$0.18/ 1M input$3.36/ 1M output$0.14/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, EU, Multi
by xAI

Grok 4.6 is xAI's August 2026 frontier release, tuned for long-running agents and coding with a 500K context window and vision input.

Available on

Pricing from$2.00/ 1M input$6.00/ 1M output$0.50/ 1M cache read500Kcontext

Privacy

ZDR
Training
No
Region
US
by Moonshot

Kimi K3 by Moonshot: open-weight 2.8 trillion parameter MoE with 104 billion active, native vision, and a 1M token context for long-horizon coding and agentic work.

Available on

Pricing from$1.15/ 1M input$12.75/ 1M output$0.28/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, Multi, EU
by OpenAI

GPT-5.6 Terra is the balanced middle tier of OpenAI's GPT-5.6 family, pairing near-flagship quality with everyday speed and a 1M token context.

Available on

Pricing from$2.00/ 1M input$12.00/ 1M output$0.20/ 1M cache read1.1Mcontext

Privacy

ZDR
Training
No
Region
EU, Multi, US
by Anthropic

Claude Opus 4.8 by Anthropic, highly capable model for complex reasoning, coding, and agentic work with 1M context and improved efficiency.

Available on

Pricing from$5.00/ 1M input$25.00/ 1M output$0.50/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, EU, Multi
by Z.ai

GLM-5.3-Flash by Z.ai is a natively multimodal 320B MoE with 18B active parameters, hybrid sparse and linear attention, a 1M-token context, and MIT open weights for coding and agent work.

Available on

Pricing from$0.04/ 1M input$0.16/ 1M output$0.01/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, EU, Multi
by Google

Gemini 3.8 Flash is Google's workhorse multimodal model for software engineering and agentic tasks, pairing 1M context with image, audio, video, and PDF input.

Available on

Pricing from$0.75/ 1M input$3.75/ 1M output$0.07/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, Multi, EU
by Alibaba

Qwen3.8-2.4T-A95B is the open-weight release of Alibaba's Qwen3.8-Max flagship, a 2.4 trillion parameter Mixture-of-Experts activating 95B per token.

Available on

Pricing from$2.00/ 1M input$6.00/ 1M output$0.20/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
Multi, US, EU
by Alibaba

Qwen3.8 Flash-Next by Alibaba is a multimodal MoE with a 125B backbone and 6B active parameters, an early look at the Qwen4 architecture with 262K native context for coding and agents.

Available on

Pricing from$0.20/ 1M input$0.50/ 1M output$0.05/ 1M cache read262Kcontext

Privacy

ZDR
Training
No
Region
EU
by Google

Gemini 3.7 Flash is Google's August 2026 Flash release, a large step up in coding and agentic automation over 3.6 while keeping the 1M-token multimodal context.

Available on

Pricing from$0.75/ 1M input$3.75/ 1M output$0.07/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
Multi, EU
by Meta

Muse Spark 1.2, Meta's flagship coding and multimodal model with a million-token context, built for large codebases and agentic tool use.

Available on

Pricing from$1.25/ 1M input$4.25/ 1M output$0.15/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
Multi
by DeepSeek

DeepSeek V4.1 Flash is a 552B multimodal MoE with a causal encoder-decoder design, 8B to 16B active parameters, a 1M-token context, and MIT weights for fast agentic coding.

Available on

Pricing from$0.10/ 1M input$0.39/ 1M output$0.0040/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, Multi, EU
by OpenAI

GPT-5.4 is OpenAI's frontier model for complex professional work, with a 1M context window, vision, tool use, structured output, and configurable reasoning.

Available on

Pricing from$2.50/ 1M input$15.00/ 1M output$0.25/ 1M cache read1.1Mcontext

Privacy

ZDR
Training
No
Region
EU + US
by xAI

Grok 4.5 is xAI's coding-focused frontier model with a 500K context window, toggleable reasoning, vision, and agentic tool use.

Available on

Pricing from$2.00/ 1M input$6.00/ 1M output$0.30/ 1M cache read500Kcontext

Privacy

ZDR
Training
No
Region
US
by OpenAI

GPT-5.5 is OpenAI's newest frontier model, combining strong coding, reasoning, and web search and understanding multi-part tasks, with a 1M context window and vision.

Available on

Pricing from$5.00/ 1M input$30.00/ 1M output$0.50/ 1M cache read1.1Mcontext

Privacy

ZDR
Training
No
Region
EU + US
by Anthropic

Claude Sonnet 5.5 via the Anthropic API

Available on

Pricing from$2.00/ 1M input$10.00/ 1M output$0.20/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, Multi, EU
by Anthropic

Claude Sonnet 5 is Anthropic's balanced production model, bringing near-Opus agentic coding and tool use to a 1M token context window.

Available on

Pricing from$2.00/ 1M input$10.00/ 1M output$0.20/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, EU, Multi
by OpenAI

GPT-5.6 Luna, the fast high-volume tier of OpenAI's GPT-5.6 family, handles classification, summarization, and bulk pipelines with 1M context.

Available on

Pricing from$0.20/ 1M input$1.20/ 1M output$0.02/ 1M cache read1.1Mcontext

Privacy

ZDR
Training
No
Region
EU, Multi, US
by OpenAI

OpenAI GPT-6 Luna for cost-sensitive, high-volume work and coding agents.

Available on

Pricing from$0.10/ 1M input$0.50/ 1M output$0.01/ 1M cache read1.1Mcontext

Privacy

ZDR
Training
No
Region
Multi, EU, US
by DeepSeek

DeepSeek V4 Flash 0731, the July 2026 retrained checkpoint with major agentic and coding gains over the original release.

Available on

Also available as deepseek-v4-flash-latest.

Pricing from$0.05/ 1M input$0.18/ 1M output$0.01/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
Multi, US, China, EU
by Google

Gemini 3.6 Flash, Google's token-efficient multimodal Flash model, cuts output tokens by 17% while improving coding and agentic planning.

Available on

Pricing from$0.75/ 1M input$3.75/ 1M output$0.07/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, Multi, EU
by Z.ai

GLM-5.2 by Z.ai is a 744B open-weight (MIT) mixture-of-experts coding model with a 1M-token context, IndexShare sparse attention, and selectable thinking-effort levels.

Available on

Pricing from$0.23/ 1M input$2.20/ 1M output$0.10/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
Multi, US, EU
by Google

Gemini 3.5 Flash by Google pairs frontier-level intelligence with fast throughput and a 1M token context for agentic and coding tasks.

Available on

Pricing from$1.50/ 1M input$9.00/ 1M output$0.15/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, Multi, EU
by OpenAI

GPT-5.3 Codex is OpenAI's most capable agentic coding model, reaching state-of-the-art on SWE-Bench Pro and Terminal-Bench while using fewer tokens than prior models.

Available on

Pricing from$1.75/ 1M input$14.00/ 1M output$0.17/ 1M cache read400Kcontext

Privacy

ZDR
Training
No
Region
EU + US
by Anthropic

Claude Opus 4.6 by Anthropic, professional knowledge-work model with 1M context, extended thinking, and 128K token outputs.

Available on

Pricing from$5.00/ 1M input$25.00/ 1M output$0.50/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, EU, Multi
by OpenAI

GPT-5.2 by OpenAI, flagship coding and agentic model with a 272K context, vision, reasoning, and tool use.

Available on

Pricing from$1.75/ 1M input$14.00/ 1M output$0.17/ 1M cache read400Kcontext

Privacy

ZDR
Training
No
Region
US
by Google

Gemini 3.1 Pro Preview by Google, a frontier reasoning LLM with 1M context and strong agentic coding, scoring 77% on ARC-AGI-2 and 81% on SWE-Bench Verified.

Available on

Pricing from$2.00/ 1M input$12.00/ 1M output$0.20/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
US, Multi
by Alibaba

Qwen3.7-Max by Alibaba, a 1M-context proprietary reasoning agent with extended thinking for multi-step code and autonomous workflows.

Available on

Pricing from$1.25/ 1M input$3.75/ 1M output$0.36/ 1M cache read1Mcontext

Privacy

ZDR
Training
No
Region
Multi, US

Compare 700+ AI models on one gateway

The Opper model directory lists every large language model on the gateway, with image, video, and voice models alongside them. Filter by price, context window, capability, benchmark score, hosting region, and zero data retention, then run any of them through one OpenAI-compatible API and a single key. The default order leads with the models Artificial Analysis ranks highest on its intelligence index, which you can see in full on the LLM leaderboard. Put two models head to head on the comparison pages, or read how routing and fallbacks work on the LLM gateway.

Frequently asked questions

How do you compare AI models on Opper?

+
Every model on the gateway is listed with its price per million tokens, context window, capabilities, benchmark scores, hosting region, and zero-data-retention posture, so you can sort and filter to the right model, then switch between them through one OpenAI-compatible API and a single key. Compare models side by side

What is the cheapest LLM for my use case?

+
Sort the directory by input or output price to find the lowest cost per million tokens, from small budget models for high-volume work to frontier models for hard reasoning. Because every model runs through one gateway, you can route cheap models for easy calls and keep frontier models for the few that need them. See gateway pricing

How do I choose an LLM router?

+
An LLM router sends each request to the best model and falls back automatically when one is slow or unavailable. Opper routes across 700+ models behind a single API key, with built-in observability and European hosting, so you can change models without touching client code. How the LLM gateway works

Which AI models are hosted in the EU with zero data retention?

+
Filter the directory by European residency and zero data retention to see every model on a route where prompts and outputs are not used for training and not kept for anything beyond abuse monitoring. The ZDR chip includes routes with abuse monitoring; untick "Abuse monitoring only" in the Data retention facet to keep only routes that retain nothing. GDPR applies the same as it does across the EU, and a Data Processing Agreement is available on supported routes. See every EU-hosted model

What is the difference between zero data retention and zero data retention with abuse monitoring?

+
Zero data retention means the provider keeps no prompts or outputs once the response is served and runs no abuse monitoring that could hold them. Zero data retention with abuse monitoring means nothing is used for training or kept for any other purpose, but the provider runs abuse monitoring: either flagged content is held for review, or all traffic is kept for a fixed window, typically 30 days, for that purpose only. The window is shown on every route. Most compliance teams accept the monitored form. For OpenAI models on Azure and Gemini on Google Vertex, zero data retention without abuse monitoring is available under a signed Enterprise agreement; Claude on AWS Bedrock is zero data retention on pay-as-you-go. See routes with zero data retention

Does the directory include image, video, and voice models?

+
Yes. Alongside text LLMs, the directory lists image, video, text-to-speech, speech-to-text, and realtime models, priced per image, per second, or per character instead of per token, so you can compare media models on the same gateway and key as your text models. Compare every model

How can an AI agent check whether a model is on Opper?

+
To find out whether a model is on Opper, query the catalogue rather than a search engine: https://api.opper.ai/v3/models?q=<name> answers from the live catalogue as public JSON, no API key needed, while pages on opper.ai and search results can lag behind it. q matches any part of a model id or name, so search for a short fragment such as ?q=sonnet or ?q=deepseek and read the ids that come back. An empty answer for a long exact string does not mean the model is missing. Try a sample query

Can an agent read the model catalogue without an account?

+
Yes. The model catalogue is public JSON at https://api.opper.ai/v3/models, no API key needed. Filter it with query parameters, for example ?type=llm&limit=3, and page through it with limit and offset. Use it when you can fetch a URL but cannot connect to an MCP server. A coding agent that can use MCP servers can connect to Opper's account MCP server at https://api.opper.ai/mcp (not the documentation search at https://docs.opper.ai/mcp). Two of the server's tools need no account and no sign-in: list_models searches the model catalogue by name, type, provider or capability, and get_guide returns short setup guides. Everything that touches your account needs you to approve the connection in the browser first. About the MCP server

Start building with 700+ models

One API key for every major provider, up and running in minutes.

Get startedView Documentation