Claude Opus 5.5 via AWS Bedrock (EU region)
Available on
Privacy
- ZDR
- Training
- No
- Region
- EU
EU data residency
Every model you can run in the EU through Opper, with where it runs, what the provider keeps and what it costs.
Each card shows only the model's routes in Europe, so the hosts, zero data retention and prices are the ones you get there. Norway applies the GDPR in full through the EEA, so its routes are included. EU or EEA, and why it matters
Claude Opus 5.5 via AWS Bedrock (EU region)
Claude Opus 5 is Anthropic's flagship for complex agentic coding and enterprise work, with a 1M token context window and adaptive thinking.
GPT-5.6 Sol tops OpenAI's GPT-5.6 family, the frontier tier for complex coding, deep research, and long-running agents with a 1M token context.
GLM-5.3 by Z.ai is a 744B open-weight coding and agentic model, post-trained on the GLM-5.2 base with a 1M-token context and low, high, and max reasoning effort levels.
Kimi K3 by Moonshot: open-weight 2.8 trillion parameter MoE with 104 billion active, native vision, and a 1M token context for long-horizon coding and agentic work.
GPT-5.6 Terra is the balanced middle tier of OpenAI's GPT-5.6 family, pairing near-flagship quality with everyday speed and a 1M token context.
Claude Opus 4.8 by Anthropic, highly capable model for complex reasoning, coding, and agentic work with 1M context and improved efficiency.
GLM-5.3-Flash by Z.ai is a natively multimodal 320B MoE with 18B active parameters, hybrid sparse and linear attention, a 1M-token context, and MIT open weights for coding and agent work.
Gemini 3.8 Flash is Google's workhorse multimodal model for software engineering and agentic tasks, pairing 1M context with image, audio, video, and PDF input.
Claude Opus 4.7 by Anthropic, advanced vision and coding model with 1M context for complex long-running agentic workflows.
Qwen3.8-2.4T-A95B is the open-weight release of Alibaba's Qwen3.8-Max flagship, a 2.4 trillion parameter Mixture-of-Experts activating 95B per token.
Qwen3.8 Flash-Next by Alibaba is a multimodal MoE with a 125B backbone and 6B active parameters, an early look at the Qwen4 architecture with 262K native context for coding and agents.
Gemini 3.7 Flash is Google's August 2026 Flash release, a large step up in coding and agentic automation over 3.6 while keeping the 1M-token multimodal context.
DeepSeek V4.1 Flash is a 552B multimodal MoE with a causal encoder-decoder design, 8B to 16B active parameters, a 1M-token context, and MIT weights for fast agentic coding.
Claude Sonnet 5 is Anthropic's balanced production model, bringing near-Opus agentic coding and tool use to a 1M token context window.
GPT-5.6 Luna, the fast high-volume tier of OpenAI's GPT-5.6 family, handles classification, summarization, and bulk pipelines with 1M context.
OpenAI GPT-6 Luna for cost-sensitive, high-volume work and coding agents.
DeepSeek V4 Pro, flagship 1.6T MoE with 1M context and hybrid thinking for frontier reasoning and agents.
DeepSeek V4 Pro 0813 is the general availability build of DeepSeek's 1M-context flagship, tuned for agentic tool use and terminal-driven work.
DeepSeek V4 Flash 0731, the July 2026 retrained checkpoint with major agentic and coding gains over the original release.
DeepSeek V4 Flash, lightweight 284B MoE with 1M context and hybrid thinking for cost-efficient agents.
Gemini 3.6 Flash, Google's token-efficient multimodal Flash model, cuts output tokens by 17% while improving coding and agentic planning.
Qwen3.8 27B by Alibaba is a dense 27B open-weight vision-language model with image and video input, a 262K native context, thinking on by default, and Apache 2.0 weights.
GLM-5.2 by Z.ai is a 744B open-weight (MIT) mixture-of-experts coding model with a 1M-token context, IndexShare sparse attention, and selectable thinking-effort levels.
Gemini 3.5 Flash by Google pairs frontier-level intelligence with fast throughput and a 1M token context for agentic and coding tasks.
GPT-5.3 Codex is OpenAI's most capable agentic coding model, reaching state-of-the-art on SWE-Bench Pro and Terminal-Bench while using fewer tokens than prior models.
Claude Opus 4.6 by Anthropic, professional knowledge-work model with 1M context, extended thinking, and 128K token outputs.
Claude Sonnet 4.6 by Anthropic, full upgrade across coding, computer use, and long-context reasoning.
MiniMax M3: natively multimodal 428B Mixture-of-Experts model with ~23B active parameters, MiniMax Sparse Attention for long context, and 59.0% on SWE-Bench Pro
Claude Opus 4.5 by Anthropic, highly capable model for coding and agents with extended thinking and a tunable effort parameter.
Kimi K2.6 by Moonshot: open-weight agentic model with long-horizon coding, 300-agent swarms, video understanding, and 256K context.
GLM-5-Turbo by Z.ai is a speed-optimized agentic model built for OpenClaw workflows, with thinking modes, reliable tool calling, and 200K context.
Kimi K2.7 Code by Moonshot: open-weight, code-specialized agentic model on a 1 trillion parameter MoE, tuned for long-horizon software engineering and tool use.
GPT-5.4 Mini is OpenAI's strongest small model, improving over GPT-5 Mini in coding and reasoning while running more than 2x faster, with a 400k context window.
Kimi K2.5 by Moonshot: native multimodal 1T-parameter MoE with vision, 256K context, hybrid reasoning, and an agent swarm for parallel task execution.
GLM-5V-Turbo by Z.ai is a native multimodal model with vision, text, and video input that turns design mockups into front-end code.
MiniMax M2.5: 229B MoE model trained on hundreds of thousands of real environments, scoring 80.2% on SWE-Bench Verified and 76.3% on BrowseComp
MiniMax M2.7: self-evolving 229B model with 56.22% SWE-Pro, native multi-agent teams, and 1495 ELO on GDPval-AA (highest open-weight)
Gemini 3.5 Flash Lite is Google's low-latency workhorse for high-volume automation, pairing 1M context with audio, video, and PDF understanding.
DeepSeek V3.2, GPT-5-class open MoE that integrates reasoning directly into tool-use.
Qwen3.5 397B A17B by Alibaba: large sparse MoE model with 397B total (17B active) parameters for advanced reasoning and coding.
Qwen3.6-27B by Alibaba: 27B open-weight multimodal model with strong agentic coding (77.2 SWE-bench Verified) and reasoning across 262K context.
Claude Sonnet 4.5 by Anthropic, strong coding and agentic model scoring 61.4% on the OSWorld computer-use benchmark, 30+ hour focus.
GPT-5.4 Nano is the fastest, lightest GPT-5.4 variant, built for classification, data extraction, and lightweight coding subagents that need low latency.
GPT-5 Mini is a fast, cost-efficient GPT-5 model for well-defined tasks and precise prompts, delivering near-frontier intelligence for high-volume workloads.
GPT-5.1 Codex Mini by OpenAI, fast, cost-efficient code model with a 272K context for everyday development tasks.
Gemma 4 31B is Google DeepMind's 30.7B parameter dense open model, the largest of the Gemma 4 family, with 256K context, thinking mode, and native function calling.
Qwen3.6 35B-A3B by Alibaba: sparse MoE (35B total, 3B active) with vision, reasoning, and agentic coding for repository-level tasks.
Qwen3.6 35B A3B FP8 by Alibaba: fine-grained FP8 quantized variant of the sparse MoE model with near-identical performance.
Qwen3.5-122B-A10B by Alibaba: sparse MoE multimodal model with 122B total (10B active) parameters and 262K context for reasoning, vision, and agentic tasks.
Muse Glimmer is Meta's 30B open-weight agentic model, distilled from Muse Spark and built to plan, call tools, and recover from failures on a single consumer GPU.
Claude Haiku 4.5 by Anthropic, fast model matching Claude Sonnet 4 coding performance at a fraction of the cost, more than twice the speed.
Gemini 2.5 Pro by Google, an advanced reasoning LLM with 1M context and thinking mode, strong at complex coding, math, STEM, and deep document analysis.
GLM 4.7 Flash by Zhipu: 30B MoE (3B active) with 200K context, tuned for local deployment and fast agentic coding.
Grok 4.20 Non-Reasoning by xAI delivers fast, direct responses across a 2M context window with real-time X data.
Qwen3.5 9B by Alibaba: 9B multimodal model with vision, tool calling, and reasoning for balanced performance and efficiency.
DeepSeek R1 0528: updated reasoning model with deeper chain-of-thought, system prompts, and structured output
Gemini 2.5 Flash (Google). Fast multimodal model for text, images, video, and audio, with reasoning, tools, and structured output.
GPT-5 Nano is the fastest, lightest GPT-5 variant, built for summarization, classification, and simple tasks that need minimal latency.
Qwen3 235B A22B Instruct 2507 by Alibaba, an instruction-tuned MoE model with a 262K context window.
GPT-OSS 120B by OpenAI: open-weight Apache 2.0 reasoning model, 117B params (5.1B active), near o4-mini on reasoning, tools and structured output.
Mistral Small 4 (March 2026) unifies reasoning, coding, and multimodal in one MoE model with 256K context and configurable reasoning effort.
GPT-OSS 20B by OpenAI: open-weight Apache 2.0 model, 21B params (3.6B active), near o3-mini, for efficient and on-device reasoning.
Qwen3 Coder Next by Alibaba, an 80B ultra-sparse MoE model with 3B active parameters and 256K context for efficient local coding agents.
Gemini 3.1 Flash Lite by Google, a cost-effective, low-latency multimodal LLM with 1M context, thinking mode, and tool use for efficient task automation.
Gemini 2.5 Flash Lite by Google, an efficient multimodal LLM with 1M context, optimized for cost and speed across text, vision, and structured tasks.
Qwen3 VL 30B A3B Instruct by Alibaba, an efficient open-weight MoE vision-language model with 3B active parameters and 131K context.
Llama 3.3 70B Instruct: refined 70B model with improved multilingual performance, structured output, and tool calling for reasoning workloads.
Llama 3.3 70B Instruct in an FP8 build, Meta's multilingual open-weight workhorse with 128K context, quantized for faster serving at near-identical quality.
Amazon Nova Pro, a balanced multimodal model combining accuracy, speed and cost for text, image and video understanding.
Llama 3.1 8B Instruct: efficient open-weight small model with a 131K context, tool calling, and structured output for production deployments.
Amazon Nova Lite, a cost-effective multimodal model with 300K context for document analysis, visual Q&A, and video understanding.
Amazon Nova Micro, a fast text-only model for real-time summarization, translation, classification and simple reasoning.
Claude 3 Haiku by Anthropic, fast and affordable vision model for customer support and batch processing of large documents.
GPT-5.3 Chat by OpenAI, chat-tuned model with a 272K context, vision, and tools for conversational use.
GPT-3.5 Turbo by OpenAI, a legacy text and code model with fast inference, now superseded by more capable alternatives.
OpenAI's most capable embedding model, producing 3072-dimensional vectors with optional dimension reduction for high-accuracy multilingual retrieval and semantic search.
OpenAI's 1.55B-parameter speech-to-text model supporting 99 languages with 10-20% error reduction over Whisper Large v2, released November 2023.
OpenAI's pruned 809M-parameter speech-to-text model with 4 decoder layers, delivering much faster multilingual transcription at minimal quality loss.
OpenAI's legacy embedding model from December 2022, producing 1536-dimensional vectors, now superseded by the text-embedding-3 family.
Gemini 2.5 Flash Image (Google). Fast multimodal model with native image generation and conversational editing.
Gemini Embedding 2 by Google creates multimodal embeddings up to 3,072 dimensions for semantic search and retrieval.
Google Gemma 4 31B IT, instruction-tuned 31B dense model with multimodal reasoning, function calling, and 256K context.
Google Gemma 4 26B-A4B IT, mixture-of-experts open model with only 3.8B active parameters, multimodal reasoning, and function calling.
Mistral OCR 4.1, the current model behind Mistral's Document AI stack, adding paragraph-level bounding boxes and block-level confidence scores to extracted text.
Mistral OCR 4, the June 2026 document model that returns extracted text with bounding boxes, block classification, and confidence scores in 170 languages.
Mistral Medium 3.5, dense 128B open-weight model with 256k context optimized for agentic and coding workflows.
Mistral Small 3.2 24B Instruct is a 128K-context open-weight text and tool-calling model tuned for reliable instruction following and lean deployments.
Mistral Small 3.1 24B is an open-weight instruction-tuned model with 128K context, vision, and function calling at around 150 tokens per second.
Mistral Devstral 2: code agents model for software engineering, 256k context, 123B params, 72.2% on SWE-bench Verified.
Mistral Ministral 3 14B: largest Ministral 3 model, 256k context, Apache 2.0, tuned for local and edge deployment.
Mistral Ministral 3 3B: smallest Ministral 3 model, 256k context, Apache 2.0, ultra-efficient edge deployment.
Mistral Ministral 3 8B, edge-optimized multimodal model with 256k context, vision, and structured outputs.
Mistral Large 3, open-weight mixture-of-experts model with 256k context and native multimodal vision.
Mistral Codestral 25.08: code-optimized model for fill-in-the-middle completion and fast code generation, 128k context.
Mistral Medium 3.1, frontier-class multimodal model balancing reasoning, vision, and agentic control.
Pixtral Large is a 124B multimodal vision model with 128K context, built for document analysis, visual reasoning, and frontier image understanding.
Voxtral Mini Latest routes to the current Voxtral Mini, a 32K-context speech model for voice agents with native audio understanding.
Voxtral Mini Transcribe Realtime 2602 delivers streaming live transcription with 13-language support, speaker diarization, and a 4B edge-deployable model.
Voxtral Mini Transcribe Realtime Latest routes to the current streaming transcription model, a 4B 13-language model for voice agents and live apps.
Voxtral Mini Transcribe Realtime by Mistral: streaming speech-to-text with configurable delay for live transcription, voice agents, and real-time subtitling.
Voxtral Small 2507 by Mistral: 24B audio-native LLM with function calling and structured outputs for speech understanding and agentic chat.
Voxtral Small (latest) by Mistral: open-weight 24B audio-input instruct model that tracks the newest snapshot, with function calling and structured outputs.
Mistral OCR 3, document understanding model with a 74% win rate over OCR 2 on forms, tables, and handwritten text.
Voxtral Mini 2602 is a 4B streaming speech-to-text model for live applications, with speaker diarization and 13-language support.
Mistral Codestral Embed: embedding model specialized for semantic code search and retrieval, 8k context.
Mistral Codestral Embed 25.05: code retrieval embedding model Mistral reports beating Voyage Code 3, Cohere, and OpenAI.
Mistral Embed, general-purpose text embedding model producing 1024-dimensional vectors for semantic search and retrieval.
Mistral Embed 2312, efficient text embedding model with an 8k context for RAG and semantic search applications.
Voxtral Mini TTS by Mistral: 4B text-to-speech with zero-shot voice cloning from about 3 seconds of reference audio, supporting 9 languages.
Voxtral Mini TTS (latest) by Mistral: text-to-speech with zero-shot voice cloning and multilingual support, tracking the newest model snapshot.
Grok 4.20 Reasoning by xAI generates visible chain-of-thought across a 2M context window with real-time data and tool use.
Qwen3-VL-Flash is the speed-focused tier of Alibaba's Qwen3-VL vision line, reading documents, charts, and video at low latency with 262K context.
Qwen3-VL-Plus, Alibaba's flagship vision-language model for document understanding, video analysis, and visual agents, with 262K context and hybrid thinking.
Qwen3-Embedding 8B by Alibaba, a 4096-dimensional embedding model with 36 layers and 32K context for dense retrieval and retrieval-augmented generation.
Qwen3 30B A3B by Alibaba, an efficient MoE model with 3B active parameters and switchable thinking and non-thinking modes.
Qwen3 8B dense model on Fireworks
Qwen3 14B, Alibaba's April 2025 dense model with switchable thinking and direct modes for reasoning and tool use at 40K context.
Qwen3 32B by Alibaba, a dense 32B model with switchable thinking and non-thinking modes for reasoning and speed.
Multilingual E5 Large by Microsoft: semantic search embedding model supporting 100 languages with 1024-dimensional representations.
Multilingual E5 Large Instruct by Microsoft: instruction-tuned embedding model with 1024 dimensions for multilingual semantic retrieval across 100 languages.
Azure AI Speech HD voices (DragonHD), served from Sweden Central. Higher quality and a higher rate than the standard neural voices, which is why they are a separate row.
Azure AI Speech standard neural voices, served from Sweden Central. 579 GA voices across 150+ locales, including German regional voices (de-AT Ingrid and Jonas, de-CH Leni and Jan) that no other provider on Opper offers. The voice name selects both speaker and locale.
Amazon Nova 2 Lite, a fast multimodal model with extended thinking, 1M context, web grounding and code execution for reasoning tasks.
Cohere Rerank v3.5 is a multilingual semantic ranking model that reorders search and RAG results with state-of-the-art precision across 100+ languages and complex enterprise data.
BGE-reranker-v2-m3 by BAAI, lightweight 600M cross-encoder reranker built on BGE-m3 for fast multilingual relevance scoring.
NVIDIA Nemotron 3 Nano 30B is a 30B Mixture-of-Experts model with about 3.5B active parameters, tuned for reasoning, tool-calling, and code.
Talkie 1930 is a 13B open-weight model trained only on pre-1931 English text, giving a 1930 knowledge cutoff for historical reasoning and contamination-free research.
IBM Docling: open-source OCR and document parsing with layout analysis, handling many formats including PDF, images, Office documents, and audio.
Apertus 70B, the Swiss AI Initiative's fully open model from EPFL and ETH Zurich, trained on 15T tokens with over 1,800 natively supported languages.
Brick Complexity Pro, Regolo's hosted prompt-complexity classifier that grades queries easy, medium, or hard so routing picks the right model tier.
green-l-raw GreenPT Backed by Mistral Small 3.2 24B Direct access to the same GreenPT-backed model as green-l without the built-in system prompt. Input €0.25 Output €0.80 Context 128k Max output 32k Released Jun 2025 Text Images Documents Multilingual No system prompt
green-r-raw GreenPT Backed by GPT-OSS Direct access to the same reasoning stack as green-r without the GreenPT system prompt. Input €0.35 Output €0.95 Context 128k Max output 32k Released Aug 2025 Text Images Documents Multilingual No system prompt
green-l GreenPT Backed by Mistral Small 3.2 24B GreenPT-branded chat model tuned for multilingual writing, image understanding, and Dutch grammar guardrails. Input €0.25 Output €0.80 Context 128k Max output 32k Released Jun 2025 Text Images Documents Multilingual Writing assistant Dutch grammar guardrails
green-r GreenPT Backed by GPT-OSS GreenPT-branded reasoning model for advanced analysis, writing, and content generation. Input €0.35 Output €0.95 Context 128k Max output 32k Released Aug 2025 Text Images Documents Multilingual Advanced reasoning Writing & content generation
Expressive multilingual speech synthesis across 50+ languages, with voice selection by reference id, multi-speaker dialogue and prosody control.
Fish Speech S2.1 Pro Realtime from Fish Audio — text model on the Opper gateway.
Faster Whisper Large V3 is a multilingual speech-to-text model supporting 99 languages, a CTranslate2 reimplementation of OpenAI's Whisper that runs up to 4x faster at the same accuracy.
Sao10K's L3.3 70B Euryale v2.3, a 70B Llama 3.3 creative roleplay model and the direct successor to v2.2.
ALIA 40B Instruct 2601, Spain's publicly funded 40B multilingual model from the Barcelona Supercomputing Center, Apache 2.0 and pretrained across 35 European languages.
TheDrummer's UnslopNemo 12B v4.1, a Mistral Nemo fine-tune for adventure writing and roleplay with a 32K context.
Kev-4B System One decision model, a fine-tune by Jared Palmer of Alibaba's Qwen3.5-4B-Base, for typed yes/no decisions (noul), classification (choice) and rubric scoring (score), returning a probability for every option instead of text. Text or structured text input only, up to 8,192 tokens for the state plus one question; trained on states of up to 384 tokens, so accuracy on long documents is lower. Use POST /v3/compat/v1/systemone with model opper/kev-4b. Weights by Jared Palmer: https://huggingface.co/jaredpalmer/kev-4b. Base model: Alibaba's Qwen3.5-4B-Base, https://huggingface.co/Qwen/Qwen3.5-4B-Base.
Undi95's ReMM SLERP L2 13B, a Llama 2 SLERP merge recreating the MythoMax recipe with updated component models.
Mythomax L2 13B by Gryphe is an open-weight Llama 2 model tuned for character-driven roleplay and creative storytelling with a consistent voice across longer exchanges.
Laya System One decision model by ConvAI Innovations (Nandha Kishor M), a non-autoregressive ModernBERT-large encoder (421M) for typed yes/no decisions (noul), classification (choice) and rubric scoring (score), returning a probability for every option instead of text. Text or structured text input only; the English checkpoint reads 512 tokens and truncates longer input. Use POST /v3/compat/v1/systemone with model berget/convaiinnovations/laya. Weights: https://huggingface.co/convaiinnovations/laya (Apache-2.0).
KBLab's KB Whisper Large is a 2-billion-parameter Swedish speech-to-text model trained on 50,000 hours of audio, cutting word error rate by an average 47% versus Whisper-large-v3.
Norwegian and English speech-to-text from the National Library of Norway, a 1.55B-parameter Whisper model trained on 66,000 hours including parliamentary speeches and broadcasts.
Pianissimo from Klang — speech to text model on the Opper gateway.
Every route has its own model id. Call that id on any plan and the request is served on that route, in that location. The examples below call Claude Opus 5.5 on AWS Bedrock in Sweden.
Copy this into a coding agent like Claude Code, Cursor or Codex and it will wire up Opper and call this route.
import OpenAI from "openai";const client = new OpenAI({apiKey: process.env.OPPER_API_KEY,baseURL: "https://api.opper.ai/v3/compat",});const completion = await client.chat.completions.create({model: "aws/claude-opus-5-5",messages: [{ role: "user", content: "Hello" }],});console.log(completion.choices[0].message.content);
A model access rule with Inference location: EU, set for your organization or a project, means no LLM or embedding call can reach a route outside the EU, even by mistake. Rules are part of the Control Plane. How model access rules work
For most teams EU hosting is the right setting. Inference runs in European data centres operated by AWS, Azure, Google and European hosts, Opper itself runs in AWS Stockholm, and many routes carry zero data retention as well.
Opper can run as a dedicated instance on evroc, the Swedish cloud, limited to European-owned providers.
This list as data. The same routes as JSON at opper.ai/models/eu.json (model id, provider, country, zero data retention, training and price per route) and as markdown at opper.ai/models/eu.md.
The whole catalogue. The model catalogue is public JSON at https://api.opper.ai/v3/models, no API key needed. Filter it with query parameters, for example ?type=llm&limit=3, and page through it with limit and offset. Use it when you can fetch a URL but cannot connect to an MCP server. Coding agents that use MCP can connect to https://api.opper.ai/mcp. Two of the server's tools need no account and no sign-in: list_models searches the model catalogue by name, type, provider or capability, and get_guide returns short setup guides. Everything that touches your account needs you to approve the connection in the browser first. About the MCP server