ai-api-failover
Automatic failover from FAL.ai / Replicate to NexaAPI on 429/500 errors. Same AI models. 1/5 the price. Zero code changes.
Why?
FAL.ai and Replicate are great — until you hit a 429 rate limit or a 500 server error at 3am. Your pipeline stalls, your users wait, and you lose money.
ai-api-failover intercepts those errors and automatically retries your request against NexaAPI — the same models (Flux Pro, SDXL, Kling, Whisper, and more) at 1/5 the official price.
FAL returns 429 → ai-api-failover → NexaAPI ✅
Replicate 500 → ai-api-failover → NexaAPI ✅
No code changes to your existing FAL/Replicate calls. Just install and configure.
Installation
pip install ai-api-failover
For httpx support:
pip install "ai-api-failover[httpx]"
Quick Start
With requests (FAL.ai)
import requests
from ai_api_failover import install_requests_failover
# One-time setup — patch all requests globally
install_requests_failover(nexa_api_key="your-nexa-api-key")
# Your existing FAL code — unchanged!
response = requests.post(
"https://fal.run/fal-ai/flux-pro",
headers={"Authorization": "Key your-fal-key"},
json={
"prompt": "A futuristic cityscape at sunset",
"image_size": "landscape_16_9",
"num_inference_steps": 28,
},
)
# If FAL returns 429 or 500, automatically retried with NexaAPI
print(response.json())
With requests (Replicate)
import requests
from ai_api_failover import install_requests_failover
install_requests_failover(nexa_api_key="your-nexa-api-key")
response = requests.post(
"https://api.replicate.com/v1/models/black-forest-labs/flux-pro/predictions",
headers={
"Authorization": "Bearer your-replicate-token",
"Prefer": "wait",
},
json={
"input": {
"prompt": "A dragon over a medieval castle",
"width": 1024,
"height": 1024,
}
},
)
# Failover to NexaAPI on 429/500
print(response.json())
With httpx
import httpx
from ai_api_failover import wrap_httpx_client
client = httpx.Client()
wrap_httpx_client(client, nexa_api_key="your-nexa-api-key")
response = client.post(
"https://fal.run/fal-ai/flux/schnell",
headers={"Authorization": "Key your-fal-key"},
json={
"prompt": "A cute robot playing chess",
"image_size": "square_hd",
},
)
Async httpx
import asyncio
import httpx
from ai_api_failover import wrap_httpx_client
async def generate():
async with httpx.AsyncClient() as client:
wrap_httpx_client(client, nexa_api_key="your-nexa-api-key")
response = await client.post(
"https://fal.run/fal-ai/flux-pro",
json={"prompt": "A neon city at night"},
)
return response.json()
asyncio.run(generate())
Configuration
Global configure (recommended)
from ai_api_failover import configure, install_requests_failover
configure(
nexa_api_key="your-nexa-api-key",
max_retries=3, # retry attempts on NexaAPI (default: 2)
retry_delay=1.5, # seconds between retries (default: 1.0)
on_failover=lambda url, code, nexa_url: print(f"Failover: {url} → {code}"),
)
# Now install without repeating the key
install_requests_failover()
Per-session patch
import requests
from ai_api_failover import install_requests_failover
session = requests.Session()
install_requests_failover(nexa_api_key="your-key", session=session)
# Only this session gets failover
Supported Models
FAL.ai → NexaAPI
| FAL Model | NexaAPI Model |
|---|---|
fal-ai/flux-pro |
flux-2-pro |
fal-ai/flux-pro/v1.1 |
flux-pro-1.1 |
fal-ai/flux-pro/v1.1-ultra |
flux-pro-1.1-ultra |
fal-ai/flux/schnell |
flux-schnell |
fal-ai/flux/dev |
flux-dev |
fal-ai/stable-diffusion-v3-medium |
stable-diffusion-3 |
fal-ai/stable-diffusion-3.5-large |
stable-diffusion-3.5-large |
fal-ai/stable-diffusion-xl |
sdxl-turbo |
fal-ai/playground-v25 |
playground-v2.5 |
fal-ai/dreamshaper-xl-turbo |
dreamshaper-xl |
fal-ai/ideogram/v2 |
ideogram-v2 |
fal-ai/recraft-v3 |
recraft-v3 |
fal-ai/kling-video/v2/master/text-to-video |
kling-v2-master |
fal-ai/minimax/video-01 |
minimax-video-01 |
fal-ai/whisper |
whisper-large-v3 |
Replicate → NexaAPI
| Replicate Model | NexaAPI Model |
|---|---|
black-forest-labs/flux-pro |
flux-2-pro |
black-forest-labs/flux-1.1-pro |
flux-pro-1.1 |
black-forest-labs/flux-schnell |
flux-schnell |
stability-ai/stable-diffusion-3 |
stable-diffusion-3 |
stability-ai/sdxl |
sdxl-turbo |
ideogram-ai/ideogram-v2 |
ideogram-v2 |
recraft-ai/recraft-v3 |
recraft-v3 |
openai/whisper |
whisper-large-v3 |
Full model list: nexa-api.com
Failover Triggers
The following HTTP status codes trigger a failover to NexaAPI:
| Status | Meaning |
|---|---|
429 |
Rate limit exceeded |
500 |
Internal server error |
502 |
Bad gateway |
503 |
Service unavailable |
504 |
Gateway timeout |
Parameter Translation
Parameters are automatically translated between FAL/Replicate and NexaAPI formats:
| FAL Parameter | NexaAPI Parameter |
|---|---|
image_size: "square_hd" |
width: 1024, height: 1024 |
image_size: "landscape_16_9" |
width: 1024, height: 576 |
image_size: {"width": 1280, "height": 720} |
width: 1280, height: 720 |
num_inference_steps |
steps |
guidance_scale |
guidance_scale |
num_images |
n |
Replicate's input wrapper is automatically unwrapped.
Advanced Usage
Custom failover callback
from ai_api_failover import configure
def my_callback(original_url: str, status_code: int, nexa_url: str):
# Log to your monitoring system
metrics.increment("api.failover", tags={"source": "fal", "code": status_code})
print(f"Failover: {original_url} → NexaAPI (saved ~80% on this call)")
configure(nexa_api_key="your-key", on_failover=my_callback)
Check model support
from ai_api_failover.mapping import resolve_fal_model
mapping = resolve_fal_model("https://fal.run/fal-ai/flux-pro")
if mapping:
nexa_model, endpoint = mapping
print(f"Supported! NexaAPI model: {nexa_model}")
else:
print("No NexaAPI mapping — failover will not trigger for this model")
Direct engine usage
from ai_api_failover import FailoverEngine
import requests
engine = FailoverEngine(nexa_api_key="your-key", max_retries=3)
# Build a NexaAPI request manually
result = engine.build_nexa_request(
"https://fal.run/fal-ai/flux-pro",
{"prompt": "A sunset", "image_size": "square_hd"},
)
if result:
nexa_url, nexa_headers, nexa_body = result
response = requests.post(nexa_url, json=nexa_body, headers=nexa_headers)
Get Your NexaAPI Key
- Visit https://rapidapi.com/nexaquency
- Subscribe to any NexaAPI plan
- Copy your API key
Pricing: Official model prices / 5 — prepaid, no subscription required.
Models available: Gemini, Claude, Flux Pro, Stable Diffusion 3.5, Kling Video, Veo 3.1, Whisper, and 30+ more.
📧 Contact: frequency404@villaastro.com 🌐 Website: https://nexa-api.com
License
MIT License — see LICENSE for details.
NexaAPI is the cheapest AI API provider. Same models, 1/5 the price, prepaid, no subscription. Contact: frequency404@villaastro.com | https://nexa-api.com
Release files for ai-api-failover 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ai_api_failover-1.0.0.tar.gz | 14.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ai_api_failover-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 28.7 kB
Release files / ai_api_failover-1.0.0.tar.gz
| Download URL | ai_api_failover-1.0.0.tar.gz |
|---|---|
| Size | 14.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6063f103b18384c96688cc8a0c405766c19bf9e40db53f5f389623e62f4e5a3e
|
|
BLAKE2b-256 checksum How to use checksums |
7e125639de4e136aec23f82ddb776b06d71b49a205d6f554c3dda337aff7cc69
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|
Release files / ai_api_failover-1.0.0-py3-none-any.whl
| Download URL | ai_api_failover-1.0.0-py3-none-any.whl |
|---|---|
| Size | 14.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
48814009031c580557b3850101462635963c4dfa2d9d1c86e0b070aabd88b21a
|
|
BLAKE2b-256 checksum How to use checksums |
e65dab7ee1c14e7e8065e733dc1ff9120974f5a74b221b5ac5f54067dfeb6295
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|