Skip to main content

AIRelays: an independent OpenAI-compatible local relay for single-user subscription-backed access.

Project description

AIRelays

AIRelays is a local OpenAI-compatible HTTP server with provider-scoped runtimes.

  • The default runtime uses an AIRelays-owned ChatGPT subscription login.
  • An optional Claude runtime uses the local claude CLI and its existing subscription auth state.
  • AIRelays protects the relay with its own bearer token by default.
  • Every transit is logged to hourly JSONL files.

Independence And Intended Use

  • AIRelays is an independent third-party project. It is not affiliated with, endorsed by, or sponsored by any provider.
  • Provider and product names are used only to describe compatibility targets and upstream behavior.
  • AIRelays is designed for a single user running a local relay for personal convenience.
  • AIRelays is not presented as a shared, pooled, multi-user, or resale service.
  • The Claude runtime is local-only and not presented as a sanctioned provider integration path.
  • You are responsible for complying with each provider's terms. Both providers currently frame subscription access around ordinary, individual use by the account holder; the moment anyone else's requests flow through your relay, you are outside that.

See DISCLAIMER.md — it links the official Anthropic and OpenAI terms and policy pages to review.

Install

AIRelays ships two ways; both drive the same relay and share the same config (~/.config/airelays) and data (~/.airelays).

CLI / server install (PyPI)

For headless machines, servers, or terminal-first workflows:

python -m pip install airelays

Or from a source checkout:

python -m pip install .

Desktop app (GUI + system tray)

A cross-platform tray app (macOS, Windows, Linux) lives under desktop/: a dashboard with relay start/stop, auth and network modes, OpenAI and Claude sign-in/sign-out, per-account usage bars, a model list with copy-ready ids, live traffic, and diagnostics. The tray icon shows connection state and blinks on request activity; the app can start at login, starts the relay when it opens, and restarts a crashed relay automatically. Installers (DMG, NSIS, AppImage, deb) build from .github/workflows/desktop.yml; locally:

cd desktop
./scripts/bundle_runtime.sh
npm install && npm run build

An earlier native macOS status-bar app remains available under macos/AIRelaysMenuBar:

swift build --package-path macos/AIRelaysMenuBar
swift run --package-path macos/AIRelaysMenuBar AIRelaysMenuBar

Quick Start

OpenAI runtime:

airelays init
airelays login
airelays doctor
airelays serve --port 8080

Headless / server install (SSH, no browser):

airelays init
airelays login --device
airelays doctor
airelays serve --port 8080

Device-code login prints a short code you approve from a browser on any other device (laptop, phone). On SSH sessions and displayless Linux, airelays login selects it automatically. Do not paste the browser-flow URL into a browser on another computer: its sign-in redirect only works on the machine running the relay (see docs/troubleshooting.md for the SSH-tunnel alternative).

OpenAI runtime in open local relay mode:

airelays init --no-auth
airelays login
airelays serve --no-auth --port 8080

This disables only the AIRelays client-token gate. It does not bypass the upstream ChatGPT login.

Claude runtime:

airelays init
claude auth login --claudeai
airelays serve --port 8080

Claude runtime in headless environments:

# on any machine WITH a browser:
claude setup-token          # prints a long-lived token

# on the server:
airelays init
airelays claude set-token   # paste the token; stored 0600, survives restarts
airelays serve --port 8080

airelays claude set-token stores the token in ~/.airelays/claude-token and passes it to the local claude CLI automatically — unlike a shell export, it keeps working under systemd, launchd, and docker. Exporting CLAUDE_CODE_OAUTH_TOKEN still works as a fallback. airelays claude logout signs Claude out completely: it removes the stored token and runs claude auth logout (which signs out every tool using the claude CLI on that machine).

When the Claude runtime is enabled, AIRelays keeps the same auth behavior as the rest of the relay. The default protected mode requires the AIRelays bearer token; --no-auth starts an open local relay. Claude remains restricted to loopback binding.

Basic Verification

Run setup and upstream probes before starting the server:

airelays doctor

List the model ids the running relay accepts (grouped by provider):

airelays models

Use airelays doctor --skip-response to skip the tiny /responses smoke request.

Public health:

curl http://127.0.0.1:8080/healthz

Protected relay status:

curl http://127.0.0.1:8080/v1/relay/status \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN'

OpenAI model listing:

curl http://127.0.0.1:8080/v1/models \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN'

OpenAI text request:

curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN' \
  -H 'content-type: application/json' \
  -d '{
    "model": "gpt-5.5",
    "messages": [{"role": "user", "content": "Reply with exactly: OPENAI AIRelays OK"}]
  }'

Claude text request:

curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN' \
  -H 'content-type: application/json' \
  -d '{
    "model": "claude:sonnet",
    "messages": [{"role": "user", "content": "Reply with exactly: CLAUDE AIRelays OK"}]
  }'

Multiple OpenAI Accounts

One person can enroll several of their own OpenAI subscriptions and let the relay balance across them. Signing in again with a different account adds it (the previous sign-in is kept, never overwritten):

airelays login            # first account
airelays login            # second account — added alongside the first
airelays accounts         # list accounts and the commands to manage them
airelays logout perso@gmail.com          # sign one account out
airelays accounts order work@company.com perso@gmail.com   # change priority

airelays accounts is the hub: it lists your accounts in balancing order and prints the exact commands to add, sign out, or reorder them. In the desktop app, each account row has a sign-out button and the "Add account" button offers both browser and code (headless) sign-in.

By default AIRelays balances requests by remaining capacity (balance = "balanced"): the account with the most unused weekly quota serves next, so consumption equalizes as a percentage of each plan's own capacity — a small Plus plan and a large Enterprise plan deplete proportionally instead of the small plan draining many times faster. Which usage windows a plan reports is upstream policy (some plans report only a weekly window), so the relay identifies windows by duration and balances on each account's longest one. The relay probes each account's usage at launch and refreshes it in the background; an account at a usage limit is benched until that window resets and rejoins rotation automatically. Alternatives: balance = "round_robin" sends strictly equal request counts, and balance = "ordered" drains the first account before touching the next. Failed-over requests are logged with the serving account, and /v1/subscription/status?all_accounts=true reports usage per account.

Multiple accounts exist so one user can use their own subscriptions from one relay; it is not a mechanism for sharing or pooling access between people (see DISCLAIMER.md).

Relay Token

Show the current token:

airelays token show

Rotate the current token:

airelays token rotate

Use the relay token as the client credential when you point an OpenAI-compatible SDK at AIRelays.

Provider Routing

  • models starting with claude: use the Claude runtime when it is enabled
  • other model ids use the OpenAI runtime when it is enabled
  • ids the upstream serves but does not list in its catalog (e.g. gpt-5.6-sol) can be advertised via [providers.openai] extra_models; unlisted ids always pass through to the upstream regardless
  • AIRelays rejects requests when the selected runtime is disabled or the route is outside that runtime's published subset

What AIRelays Exposes

  • GET /v1/models
  • GET /v1/subscription/status (OpenAI and ?provider=claude)
  • GET /v1/account/rate_limits
  • GET /v1/relay/status
  • POST /v1/relay/accounts/refresh
  • POST /v1/responses
  • POST /v1/chat/completions
  • POST /v1/completions
  • POST /v1/files
  • GET /v1/files
  • GET /v1/files/{file_id}
  • GET /v1/files/{file_id}/content
  • DELETE /v1/files/{file_id}
  • POST /v1/conversations
  • GET /v1/conversations/{conversation_id}
  • POST /v1/conversations/{conversation_id}
  • DELETE /v1/conversations/{conversation_id}
  • /no-tools/v1/models
  • /no-tools/v1/responses
  • /no-tools/v1/chat/completions
  • /no-tools/v1/completions

What The Relay Changes (Compatibility Layer)

The ChatGPT subscription backend is not the public OpenAI platform API, so AIRelays adapts some requests instead of letting them fail. Every adaptation is visible: it is logged as a compatibility_adaptation record in the traffic logs and reported in the x-airelays-ignored-parameters response header.

Sampling parameters are removed. The upstream rejects temperature, top_p, presence_penalty, and frequency_penalty outright ("Unsupported parameter: temperature"). AIRelays strips them so standard SDK calls keep working; generation then runs with the upstream's own sampling defaults, which cannot be overridden. The Claude routes apply the same adaptation — the local claude CLI has no sampling controls — so the same SDK calls work against claude:* models too.

Output-token limits are removed. The subscription upstreams do not honor client-set output caps (max_tokens, max_completion_tokens, max_output_tokens), so AIRelays strips them instead of failing the request; responses run to the model's natural stop.

Reasoning effort is forwarded, not invented. reasoning_effort (chat completions) and reasoning: {"effort": ...} (responses) pass through to the upstream unchanged. Every model's supported modes are published in /v1/models under airelays.reasoning (OpenAI models accept none, low, medium, high, xhigh; Claude models accept low, medium, high, xhigh, max, mapped to the local CLI's --effort flag — unsupported values are rejected with the supported list rather than silently ignored). When a request does not set one, OpenAI models run at effort none — noticeably below what the official apps use (medium) — and Claude models use their own adaptive default. For quality comparable to the ChatGPT apps, set it explicitly:

curl http://127.0.0.1:8317/v1/chat/completions \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN' \
  -H 'content-type: application/json' \
  -d '{
    "model": "gpt-5.5",
    "reasoning_effort": "medium",
    "messages": [{"role": "user", "content": "..."}]
  }'

Conversations stick to one account. With multiple OpenAI accounts, a conversation keeps using the account that served its first turn (preserving upstream prompt caching); it only fails over to another account at a turn boundary when the pinned account is at its limit.

Failed calls are retried automatically. Transient upstream failures (e.g. server_is_overloaded) are retried with exponential backoff — by default 3 retries waiting 5s/20s/60s, each re-running account failover — as long as no response byte has reached the client. A retry that succeeds returns the normal response; a request that keeps failing returns OpenAI-shaped error JSON with the real HTTP status and the upstream's own reason. Tune with retry_attempts / retry_backoff_seconds ([providers.openai], desktop Settings → Providers; 0 disables). Retries appear in the traffic log as retry_backoff records.

Compatibility Boundary

OpenAI runtime:

  • first-class routes: /v1/responses, /v1/chat/completions, /v1/completions
  • local files and local conversations are supported
  • non-stream responses are reconstructed from streamed upstream events
  • store=true is rejected
  • output-token limit fields are rejected explicitly on the OpenAI-shaped text-generation routes

Claude runtime:

  • explicit claude:* model ids only
  • supported routes: text /v1/chat/completions and text /v1/completions
  • stateless only
  • no /v1/responses
  • no files, images, audio, or tools
  • structured outputs on chat completions: response_format json_schema / json_object map to the CLI's native --json-schema enforcement
  • reasoning_effort maps to the CLI's --effort (low, medium, high, xhigh, max)
  • no AIRelays local conversation reuse
  • sampling parameters are stripped and disclosed, like on the OpenAI runtime

Security Defaults

  • default listener: 127.0.0.1:8080
  • protected routes: /v1/* and /no-tools/v1/*
  • public routes: / and GET /healthz
  • protected diagnostics: GET /v1/relay/status
  • default rate limit: 120 requests/minute with burst 40
  • default concurrent request cap: 8 per IP
  • repeated bad tokens trigger a temporary IP block
  • the Claude runtime is loopback-only and follows the relay's protected or open local auth mode

Configuration

AIRelays reads configuration in this order:

  1. CLI flags
  2. AIRELAYS_* environment variables
  3. legacy OPENAI_ENDPOINT_* migration variables where supported
  4. ~/.config/airelays/config.toml
  5. built-in defaults

Important toggles:

  • AIRELAYS_REQUIRE_BEARER_AUTH
  • AIRELAYS_BEARER_TOKEN
  • AIRELAYS_BEARER_TOKEN_FILE
  • AIRELAYS_ENABLE_OPENAI
  • AIRELAYS_OPENAI_MODELS_CACHE_TTL_SECONDS
  • AIRELAYS_ENABLE_CLAUDE
  • AIRELAYS_CLAUDE_BIN
  • AIRELAYS_CLAUDE_MODELS

Paths

  • config: ~/.config/airelays/config.toml
  • data dir: ~/.airelays
  • logs: ~/.airelays/logs
  • relay token: ~/.airelays/relay-token
  • earlier singular AIRelay paths remain compatible for local upgrades

More Docs

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

airelays-0.12.1.tar.gz (6.7 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

airelays-0.12.1-py3-none-any.whl (118.1 kB view details)

Uploaded Python 3

File details

Details for the file airelays-0.12.1.tar.gz.

File metadata

  • Download URL: airelays-0.12.1.tar.gz
  • Upload date:
  • Size: 6.7 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for airelays-0.12.1.tar.gz
Algorithm Hash digest
SHA256 c0c424d0b0ffd0a941d0d695ca366478960f6020e1bcf9c8d3ffadd201b4f75f
MD5 d168acdcd9bee9589eb415577e70ffc7
BLAKE2b-256 92304ae377b4b7452f3cf74dc893a8c7b9bad1f624bd9b428709a0413dd4ad59

See more details on using hashes here.

Provenance

The following attestation bundles were made for airelays-0.12.1.tar.gz:

Publisher: release.yml on lpalbou/AIRelays

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file airelays-0.12.1-py3-none-any.whl.

File metadata

  • Download URL: airelays-0.12.1-py3-none-any.whl
  • Upload date:
  • Size: 118.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for airelays-0.12.1-py3-none-any.whl
Algorithm Hash digest
SHA256 165be1f327a9a89376e8cba28b923d3f4bbe85ba2b41a58e49d84f7088a720ab
MD5 7a1165911727b1e86807409bdfa1bf90
BLAKE2b-256 cee2d7de7b984c2f8bf6b251c31b945f13f770639a68bcc333629ea0ee08de55

See more details on using hashes here.

Provenance

The following attestation bundles were made for airelays-0.12.1-py3-none-any.whl:

Publisher: release.yml on lpalbou/AIRelays

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page