Run Models at Half the Price with Flex and Priority Tiers

By Jose Sabater -

Run the same model at half the price, or get faster tokens when it matters. OpenAI, Google, AWS and xAI sell each model at more than one service tier. Flex processing runs your request for about half the price, in exchange for slower responses. Priority processing gets you faster tokens at about twice the price. Opper now supports both, on every provider that sells them, with the same service_tier field OpenAI uses:

  • service_tier: "flex": about half the price. The request can queue, for minutes at worst.
  • service_tier: "priority": faster tokens, at about 2x the price.

Same model, same answers, different speed and price.

How to use flex and priority with service_tier

from openai import OpenAI

client = OpenAI(base_url="https://api.opper.ai/v3/compat", api_key="YOUR_OPPER_API_KEY")

resp = client.chat.completions.create(
    model="gpt-5.4-mini",
    service_tier="flex",
    messages=[{"role": "user", "content": "Classify this support ticket."}],
    timeout=600,  # flex can queue
)

It works on Chat Completions, Responses and Messages. The response says which tier served it (service_tier in the body, X-Opper-Served-Service-Tier in the headers), and you pay that tier's price.

What happens when a tier has no capacity

  • Flex fails rather than running at full price. You never pay standard by surprise.
  • Priority falls back to standard on the same model, billed as standard.
  • Models without a tier run at standard. The field is safe to leave on.

Flex with a fallback to standard

Each tier is its own model id, so it fits the fallbacks you already use. In the models array, flex first and the same model at standard if flex has no capacity:

{
  "model": "openai:flex/gpt-5.4-mini",
  "models": ["openai/gpt-5.4-mini"]
}

Or put both in a dynamic route Pool ordered Cheapest first: flex wins on price, standard takes over when it can't.

Which providers support flex and priority

ProviderFlexPriority
OpenAI✓✓ (and ultrafast on GPT-6 Astra)
Google Gemini API✓✓
Google Vertex AI✓✓, also in the EU
xAI✓, also in the EU
Azure OpenAI✓
AWS Bedrock✓, also in the EU✓, also in the EU
BytePlus✓

Every tier has its own endpoint and price, listed with a tier mark on the Models page in the Opper platform and in GET /v3/models.


Get started