Claude Opus 5 is Anthropic's flagship for complex agentic coding and enterprise work, with a 1M token context window and adaptive thinking.
Available on
Privacy
- Zero data retention
- Always-on
- Training
- No
- Region
- EU + US
AI Model Directory
Search by what matters. Zero data retention, residency, training and more.
Claude Opus 5 is Anthropic's flagship for complex agentic coding and enterprise work, with a 1M token context window and adaptive thinking.
Claude Fable 5 by Anthropic, state-of-the-art generalist model with advanced vision and 1M token context for complex software engineering and reasoning.
GPT-5.6 Sol tops OpenAI's GPT-5.6 family, the frontier tier for complex coding, deep research, and long-running agents with a 1M token context.
Kimi K3 by Moonshot: open-weight 2.8 trillion parameter MoE with 104 billion active, native vision, and a 1M token context for long-horizon coding and agentic work.
Qwen3.8-Max, Alibaba's 2.4T-parameter MoE flagship with hybrid thinking, vision input, and a context window approaching one million tokens.
Claude Opus 4.8 by Anthropic, highly capable model for complex reasoning, coding, and agentic work with 1M context and improved efficiency.
Muse Spark 1.2, Meta's flagship coding and multimodal model with a million-token context, built for large codebases and agentic tool use.
GPT-5.6 Terra is the balanced middle tier of OpenAI's GPT-5.6 family, pairing near-flagship quality with everyday speed and a 1M token context.
Claude Sonnet 5 is Anthropic's balanced production model, bringing near-Opus agentic coding and tool use to a 1M token context window.
Claude Opus 4.7 by Anthropic, advanced vision and coding model with 1M context for complex long-running agentic workflows.
GLM-5.2 by Z.ai is a 744B open-weight (MIT) mixture-of-experts coding model with a 1M-token context, IndexShare sparse attention, and selectable thinking-effort levels.
GLM-5.2 Fast, a speed-optimized serving of Z.ai's 744B MoE flagship with a million-token context and frontier open-weight reasoning.
GPT-5.6 Luna, the fast high-volume tier of OpenAI's GPT-5.6 family, handles classification, summarization, and bulk pipelines with 1M context.
Gemini 3.5 Flash by Google pairs frontier-level intelligence with fast throughput and a 1M token context for agentic and coding tasks.
DeepSeek V4 Flash 0731, the July 2026 retrained checkpoint with major agentic and coding gains over the original release.
Also available as deepseek-v4-flash-latest.
DeepSeek V4 Flash, lightweight 284B MoE with 1M context and hybrid thinking for cost-efficient agents.
Gemini 3.6 Flash, Google's token-efficient multimodal Flash model, cuts output tokens by 17% while improving coding and agentic planning.
Claude Sonnet 4.6 by Anthropic, full upgrade across coding, computer use, and long-context reasoning.
Gemini 3.1 Pro Preview by Google, a frontier reasoning LLM with 1M context and strong agentic coding, scoring 77% on ARC-AGI-2 and 81% on SWE-Bench Verified.
Qwen3.7-Max by Alibaba, a 1M-context proprietary reasoning agent with extended thinking for multi-step code and autonomous workflows.
GPT-5.3 Codex is OpenAI's most capable agentic coding model, reaching state-of-the-art on SWE-Bench Pro and Terminal-Bench while using fewer tokens than prior models.
MiniMax M3: natively multimodal 428B Mixture-of-Experts model with ~23B active parameters, MiniMax Sparse Attention for long context, and 59.0% on SWE-Bench Pro
DeepSeek V4 Pro, flagship 1.6T MoE with 1M context and hybrid thinking for frontier reasoning and agents.
DeepSeek V4 Pro 0813, the GA build of the 1M-context flagship MoE, tuned for agentic tool use — on Fireworks
Kimi K2.6 by Moonshot: open-weight agentic model with long-horizon coding, 300-agent swarms, video understanding, and 256K context.
Claude Opus 4.6 by Anthropic, professional knowledge-work model with 1M context, extended thinking, and 128K token outputs.
Kimi K2.7 Code by Moonshot: open-weight, code-specialized agentic model on a 1 trillion parameter MoE, tuned for long-horizon software engineering and tool use.
Kimi K2.7 Code Fast serves Moonshot's 1T-parameter agentic coding model at 100+ tokens per second for long-horizon software tasks.
Claude Opus 4.5 by Anthropic, highly capable model for coding and agents with extended thinking and a tunable effort parameter.
GLM-5.1 by Z.ai is a 754B open-weight (MIT) coding and agentic reasoning model with 202K context that tops SWE-Bench Pro at 58.4.
GPT-5.4 Mini is OpenAI's strongest small model, improving over GPT-5 Mini in coding and reasoning while running more than 2x faster, with a 400k context window.
GPT-5.4 Nano is the fastest, lightest GPT-5.4 variant, built for classification, data extraction, and lightweight coding subagents that need low latency.
Qwen3.7-Plus by Alibaba, a 1M-context multimodal agent with vision, tool calling, and agentic iteration for GUI and CLI automation.
GLM-5-Turbo by Z.ai is a speed-optimized agentic model built for OpenClaw workflows, with thinking modes, reliable tool calling, and 200K context.
MiniMax M2.7: self-evolving 229B model with 56.22% SWE-Pro, native multi-agent teams, and 1495 ELO on GDPval-AA (highest open-weight)
MiniMax M2.7-highspeed: throughput-optimized M2.7 variant delivering the self-evolving model's capabilities at higher output speed with Agent Teams
Qwen3.6-27B by Alibaba: 27B open-weight multimodal model with strong agentic coding (77.2 SWE-bench Verified) and reasoning across 262K context.
Claude Sonnet 4.5 by Anthropic, strong coding and agentic model scoring 61.4% on the OSWorld computer-use benchmark, 30+ hour focus.
Gemini 3.5 Flash Lite is Google's low-latency workhorse for high-volume automation, pairing 1M context with audio, video, and PDF understanding.
Kimi K2.5 by Moonshot: native multimodal 1T-parameter MoE with vision, 256K context, hybrid reasoning, and an agent swarm for parallel task execution.
The Opper model directory lists every large language model on the gateway, with image, video, and voice models alongside them. Filter by price, context window, capability, benchmark score, hosting region, and zero data retention, then run any of them through one OpenAI-compatible API and a single key. The default order leads with the models Artificial Analysis ranks highest on its intelligence index, which you can see in full on the LLM leaderboard. Put two models head to head on the comparison pages, or read how routing and fallbacks work on the LLM gateway.