Documentation · Reliability & limits

Guides

Reliability & limits

Fallback models when a provider fails, parallel requests and key rules: models, IP addresses, spending alerts.

Fallback models#

If a model's provider fails — a 5xx, a 429 or a dropped connection — flua retries on its own first, and if that does not help, sends the request to a similar model: from the same lab or the same price tier, never noticeably pricier. You get an answer instead of an error, billed at the price of the model that answered.

The response headers tell you which model answered:

response headers
x-flua-model: claude-sonnet-5-5
x-flua-fallback-from: claude-opus-5-5

The model field of the response body also names the model that answered. Fallbacks apply to text models in Chat Completions, Responses and Messages — never to images or embeddings. When the request carries pictures, only a model that can see them is chosen.

Turning it off#

For one request — with a body field or a header:

no_fallback.py
response = client.chat.completions.create(
    model="claude-opus-5-5",
    messages=[{"role": "user", "content": "Hi!"}],
    extra_body={"fallback": False},           # or
    # extra_headers={"x-flua-fallback": "off"},
)

For every request of a key — with the “Fallback model on outages” switch in the key settings.

Which model backs up which#

ModelFallback
claude-opus-5-5claude-sonnet-5-5
claude-sonnet-5-5gpt-6.1-sol
claude-fable-5-1claude-opus-5-5
gpt-6.1-solgemini-3.1-pro-preview
gpt-6-astragpt-6.1-sol
gpt-6-lunaqwen3.8-flash
gemini-3.8-flashdeepseek-v4.1-flash
gemini-3.1-pro-previewgemini-3.8-flash
grok-4.7glm-5.3
deepseek-v4.1-flashqwen3.8-flash
qwen3.8-flashqwen3.8-omni-flash
qwen3.8-maxqwen3.8-flash
qwen3.8-omni-flashqwen3.8-flash
kimi-k3qwen3.8-max
glm-5.3glm-5.3-flash
glm-5.3-flashdeepseek-v4.1-flash
minimax-m3deepseek-v4.1-flash
mimo-v2.6-flashdeepseek-v4.1-flash
mimo-v2.6-promimo-v2.6-flash
hy4-previewdeepseek-v4.1-flash

Each model's error rate over the last day and week is on the status page.

Parallel requests#

An account runs a limited number of requests at once — by plan:

  • Free — 5 at once, 60 per minute per key
  • Plus — 10 at once, 180 per minute per key
  • Pro — 20 at once, 600 per minute per key
  • Max — 50 at once, 2000 per minute per key

On top of that, every running request holds its estimated cost against the balance (by max_tokens, or 8,192 output tokens). A new request starts only if the balance covers the ones already running — so a burst of long requests to an expensive model cannot drive the balance deep into the red. The first request still needs only a positive balance.

429 Too Many Requests
{
  "error": {
    "message": "Too many requests in progress: up to 5 at once on the free plan. Wait for one to finish.",
    "type": "rate_limit_error",
    "code": "concurrency_limit"
  }
}

Key rules#

Each key's settings can limit which models it calls and from which IP addresses (single addresses or CIDR ranges, IPv4 and IPv6), set a spending limit and a spending alert per day, month or in total — by email and Telegram.

A key limited to some models lists only those in GET /v1/models. A request for another model gets 403 model_not_allowed, one from another address gets 403 ip_not_allowed.