Skip to main content

LLM Proxy CLI

CI Status PyPI Version License

A lightweight, fault-tolerant CLI tool for delegating LLM tasks to expert models across multiple providers (Nvidia NIM, Groq, OpenAI, Anthropic Claude, Gemini).

⚡ Features

  • Multi-Provider Support: Seamlessly route requests to nvidia, groq, openai, anthropic, or gemini.
  • Dynamic Model Discovery: Never hardcode a model name again. Use auto-smart or auto-fast and the router will auto-select the best model.
  • Active Liveness Verification: Pings candidates with a minimal chat request to drop fake/gated models before they crash your task.
  • Automatic Fallbacks: Provide a comma-separated list of models. If one fails, it instantly falls back to the next.
  • Circuit Breaker: Built-in health tracking and cooldowns to prevent spamming dead endpoints.
  • Reasoning Extraction: Automatically extracts reasoning_content from natively supported models (e.g., DeepSeek-R1 or Nemotron) and outputs them to stderr.
  • Streaming Native: Built on the official OpenAI SDK for fast and reliable streaming chunks.

🏗️ Architecture & Under the Hood

  • Language: Python 3
  • Libraries: openai, filelock, anthropic
  • Design Pattern: Circuit Breaker, Chain of Responsibility (Fallback Routing), and Dynamic Caching.

The router uses a FileLock-backed JSON state (circuit_breaker.json) to track failures across concurrent runs. If an endpoint times out or returns a 5xx error more than MAX_FAILURES times, the circuit trips and forces the router to skip that endpoint for the next 120 seconds, immediately trying the next fallback model.

📦 Installation

# Install via pip
pip install llm-proxy-cli

# Or for local development:
# git clone https://github.com/cadakerem/llm-proxy-cli.git
# cd llm-proxy-cli
# pip install -e .

🔑 Configuration & API Keys

The router looks for API keys in your environment variables or in ~/.config/llm-proxy-cli/keys.json.

Supported environment variables:

  • NVIDIA_API_KEY
  • GROQ_API_KEY
  • OPENAI_API_KEY
  • ANTHROPIC_API_KEY
  • GEMINI_API_KEY

💻 Usage

The tool is designed to automatically discover and use the best model without you having to memorize model names (using auto-smart and auto-fast).

Instead of guessing which model is currently the best or active on the API, simply use auto-smart (for complex coding/reasoning tasks) or auto-fast (for quick tasks).

# Auto-select the smartest model on Nvidia (e.g., Nemotron or Llama 3.1 405B)
llm-proxy-cli -m "nvidia:auto-smart" -p "Write a React button."

# Auto-select the fastest model on Groq
llm-proxy-cli -m "groq:auto-fast" -p "Summarize this text."

2. Chained Automatic Fallback

If Nvidia goes down or hits a rate limit, you can instantly fall back to Groq's best model by separating them with a comma:

llm-proxy-cli -m "nvidia:auto-smart,groq:auto-smart" -p "Refactor this python script."

3. Specific / Manual Model Selection

If you have a specific model you want to use, you can still hardcode it directly:

# Use Laguna, and fallback to a specific Groq model if it fails
llm-proxy-cli -m "nvidia:poolside/laguna-xs-2.1,groq:groq/compound" -p "Explain quantum entanglement."

⚠️ Troubleshooting & Known Quirks: Nvidia EULA (404 Not Found)

Nvidia NIM requires users to manually accept the End User License Agreement (EULA) for certain models on their website before using them via API. If you haven't accepted the EULA for a dynamically discovered model, Nvidia returns a cryptic 404 Not Found error. The Smart Router intercepts this behavior automatically and will print a clear warning.

How to Fix:

  1. Log into the Nvidia Build Portal.
  2. Search for the exact model name shown in the warning and click to run a quick test prompt to accept the terms.
  3. Or bypass auto-discovery completely by explicitly hardcoding a model:
llm-proxy-cli -m "nvidia:meta/llama-3.2-11b-vision-instruct" -p "Hello"

🧑‍💻 Developer & Contributions

Developed by Kerem Barbaros Karnabat (@cadakerem).

Note on Repository Structure: The core routing logic, dynamic model discovery, and circuit breaker patterns are entirely contained within llm_proxy_cli.py to ensure maximum portability. Unit tests are located in the tests/ directory, and SKILL.md provides instructions for integrating this tool as a native AI agent skill.

Contributions, issues, and feature requests are welcome! Feel free to check the Issues page.

📜 License

This project is licensed under the MIT License.

Metadata

Release files for llm-proxy-cli 0.5.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-proxy-cli 0.5.5
File Size Uploaded
llm_proxy_cli-0.5.5.tar.gz 15.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-proxy-cli 0.5.5
File Interpreter ABI Platform
llm_proxy_cli-0.5.5-py3-none-any.whl Python 3 none any Details

Total release size: 26.9 kB

Release files / llm_proxy_cli-0.5.5.tar.gz

Download URL llm_proxy_cli-0.5.5.tar.gz
Size 15.2 kB
Tags Source
SHA-256 checksum
How to use checksums
7d59789218f1e2c6d32125dc246ebb4ece40555b771af4cf13dff4ea819807c0
BLAKE2b-256 checksum
How to use checksums
c0d2b1182951e84dbeadc7824bf37d3f958f46e12cae8b0d63627cae464f95c0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release files / llm_proxy_cli-0.5.5-py3-none-any.whl

Download URL llm_proxy_cli-0.5.5-py3-none-any.whl
Size 11.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d67776aa68a9755e0bb3717b0b83781648f9e98c0049742064850f54661b371c
BLAKE2b-256 checksum
How to use checksums
94e5f6711e0de6e3861f98c46d18090e068589bf37fc62587541a687ba5f236f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.5.5 This release

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page