Opper AI partners with Greenference for low-cost EU-hosted inference on small open-weight models

By Felix Wunderlich -

Stockholm, Sweden, September 2026. Opper AI and Greenference are partnering to bring Greenference's EU-hosted inference to the Opper AI gateway. Builders can now reach Greenference's catalogue through the Opper API, from Qwen3.8 27B and Gemma 4 31B to GLM-4.7-Flash, gpt-oss-20b and Qwen3.6 27B.

Why this matters

Small models, priced to run in volume. Most of the calls inside an agent do not need a frontier model. Classification, extraction, routing and short replies run on the 8B to 32B class, many thousands of times a day, and that is the part of the catalogue Greenference serves. At launch it is the cheapest route on the gateway, in any region, for GLM-4.7-Flash, gpt-oss-20b, Qwen3.5 9B and Qwen3.6 27B, and the cheapest European route to Qwen3.8 27B, which carries 262k tokens of context here. Rosters and prices both move, so the table further down is fetched live and is the place to check.

Where it is going: tokens from sunlight. Greenference's founders come from data-centre power, battery storage and cloud operations, and they are building what they call Sunverters: air-cooled, direct-current datacentres in shipping containers, powered off-grid by on-site solar and batteries and placed next to renewable generation instead of at the back of a years-long grid queue. Their scheduler moves inference between sites to follow the sun. The first proof-of-concept site is the next milestone, and until then the API runs on GPU capacity rented inside the EU.

Clear data handling. Greenference SAS is a French processor, and its DPA commits that prompts and completions are never written to disk, logged or inspected, and never used to train a model: they exist in memory while the response is computed and are released once it is delivered. Every GPU it serves from sits inside the EU. Full residency and retention detail per route lives on the Greenference provider page.

More than a router: Rules for every call. Routing to Greenference is just the entry point. Every call through Opper can run under your Rules: spend limits, data retention, model access, content checks and routing order, set once for the organization or per project. Pin Greenference for a task or set it as a fallback, and because Greenference is OpenAI-compatible, getting there is a model string, not a migration.

"The small models do most of the work inside an agent, and what teams need from them is a low price, run at scale, somewhere they are allowed to run it. Greenference serves that end of the catalogue inside the EU, and the team is building towards serving it from sunlight, which is a direction we are glad to back early."

— Göran Sandahl, Co-founder and CEO, Opper AI

"Europe's limit on AI is power, not chips, so we are building inference that brings its own clean power instead of waiting years for a grid connection. Opper puts our open-weight models in front of a wide community of builders through a single API, and we are looking forward to what they build on it."

— Mustafa Demirkol, Co-founder and CEO, Greenference

Models live today

The catalog below is fetched live from Opper's model API and filtered to Greenference-hosted models. Availability, context windows, and pricing stay in sync with what's actually callable through Opper. All Greenference routes on Opper are EU-hosted.

ModelRegionContextInput / 1MOutput / 1M
greenference/gemma-4-31b-itEU262K$0.03$0.17
greenference/glm-4.7-flashEU131K$0.03$0.20
greenference/glm-5.3-flashEU1.0M$0.04$0.16
greenference/gpt-oss-20bEU33K$0.01$0.04
greenference/llama-3.1-8bEU16K$0.01$0.02
greenference/qwen3-14bEU16K——
greenference/qwen3-30b-a3bEU41K$0.10$0.28
greenference/qwen3-32bEU41K$0.10$0.28
greenference/qwen3-8bEU41K$0.45$0.45
greenference/qwen3.5-9bEU262K$0.04$0.07
greenference/qwen3.6-27bEU262K$0.15$1.00
greenference/qwen3.8-27bEU262K$0.05$0.90
USD per 1M tokens. Pricing and availability subject to change.

Get started

Paste this into your coding agent (Claude Code, Cursor, Codex, and more) and it will set up Opper and route to Greenference for you:

Use curl to download, read and follow: https://skills.opper.ai
Then set up Opper to use Greenference as the provider, e.g. greenference/qwen3.8-27b.

Prefer a direct call? Opper is drop-in compatible with the OpenAI, Anthropic, and Google SDKs, so one API key and the model string are all you need:

import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.OPPER_API_KEY,
baseURL: "https://api.opper.ai/v3/compat",
});
const completion = await client.chat.completions.create({
model: "greenference/qwen3.8-27b",
messages: [{ role: "user", content: "Hello" }],
});
console.log(completion.choices[0].message.content);

Follow the quick start in our docs for evaluations, fallbacks, and structured output.

About Greenference

Greenference (Greenference SAS) is a French AI infrastructure company registered in Lyon, founded by Mustafa Demirkol, Vincent Abad and Harun Barış Bulut. It serves open-weight models through an OpenAI-compatible API on GPUs inside the EU, and is building modular off-grid datacentres that run AI inference on solar power.

About Opper AI

Opper AI is the European AI gateway for agents, based in Stockholm. It gives developers access to 700+ leading AI models through one EU-hosted, GDPR-compliant API, pay-as-you-go, with spend caps, model allowlists, zero data retention and full OpenAI SDK compatibility.

Keep reading

All articles