Skip to main content

Centralized routing, retry, and API cycling utilities for LLM integrations

Project description

Central Tools for LLM Orchestration

Utilities for routing workloads to the most appropriate LLM, retrying transient failures, and cycling API keys to stay within rate limits. The first provider family supported is Google (Gemini and Gemma), with room to add more providers.

Features

  • Depth-aware routing to the right model family.
  • Pluggable retry strategies with exponential backoff for transient errors.
  • API key cycling with per-minute and per-day rate limit tracking (supports up to five keys per provider, fewer if not available).
  • Reusable abstractions for providers, routing rules, and orchestrator workflows.

Quickstart

pip install -e .
from central_tools import (
    LLMOrchestrator,
    LLMRequest,
    TaskDepth,
    APIKeyCycler,
    LLMRouter,
    ExponentialBackoffRetry,
    load_env_file,
    load_api_keys_from_placeholders,
    build_cycler,
    GOOGLE_API_KEY_PLACEHOLDERS,
)
from central_tools.providers.google import google_provider_factory
from central_tools.config import RateLimitConfig, build_router

router = LLMRouter()
router.register_model(
    depth=TaskDepth.DEEP,
    model_name="gemini-2.5-pro",
    provider_id="google",
    provider_factory=google_provider_factory("gemini-2.5-pro"),
    metadata={"family": "gemini"}
)

cycler = APIKeyCycler(max_keys=5)

rate_limits = [
    RateLimitConfig(name="per_minute", period_seconds=60, max_calls=15, cooldown_seconds=10),
    RateLimitConfig(name="per_day", period_seconds=24 * 3600, max_calls=1500),
]

load_env_file("config/.env")
api_keys = load_api_keys_from_placeholders(
    GOOGLE_API_KEY_PLACEHOLDERS,
    rate_limits=rate_limits,
)
build_cycler(cycler, api_keys)

retry_strategy = ExponentialBackoffRetry(max_attempts=4)

orchestrator = LLMOrchestrator(router=router, retry_strategy=retry_strategy, api_cycler=cycler)

request = LLMRequest(
    prompt="Summarize the latest research on retrieval-augmented generation.",
    depth=TaskDepth.DEEP,
    metadata={"family": "gemini"},
)
response = orchestrator.generate(request)
print(response.text)

See examples/quickstart.py for a more complete illustration including custom error handling and .env loading.

Environment File

Copy config/.env.template to config/.env and fill in your secrets (the repository copy is ignored by load_env_file if it is missing). The example script automatically loads this file and falls back to any GOOGLE_API_KEY_* variables that are already defined in your shell.

Cycle test (validate keys & rotation)

We provide a small test script that cycles through configured API keys and asks each registered model for a short confirmation (constrained to ~10 words) to validate both access and the APIKeyCycler behaviour.

Run the test from the repository root:

python examples/cycle_test.py

The script will:

  • Load config/.env if present (or use GOOGLE_API_KEY_1..5 environment variables).
  • Register keys in the APIKeyCycler and then iterate each key once, calling every registered model.
  • Report successes (short model-tagged replies) and failures (rate limits or other provider errors). Results are recorded in examples/cycle_test_results.md when you run the script locally.

Sample outcome (already executed in this workspace):

  • Key 1: all models returned short confirmation replies (SUCCESS).
  • Key 2 & 3: rate-limited (HTTP 429) — cycler started cooldowns for these keys.
  • Key 4: all models returned short confirmation replies (SUCCESS).

If you see repeated 429s for a key, check billing/quota and either increase quota or remove the key from rotation.

Discovering Available Models

Model availability changes over time. You can list the current Gemini endpoints exposed to your project:

import google.generativeai as genai

genai.configure(api_key="YOUR_KEY")
for model in genai.list_models():
    print(model.name)

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

central_tools_llm-2025.11.0.tar.gz (25.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

central_tools_llm-2025.11.0-py3-none-any.whl (28.4 kB view details)

Uploaded Python 3

File details

Details for the file central_tools_llm-2025.11.0.tar.gz.

File metadata

  • Download URL: central_tools_llm-2025.11.0.tar.gz
  • Upload date:
  • Size: 25.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.2

File hashes

Hashes for central_tools_llm-2025.11.0.tar.gz
Algorithm Hash digest
SHA256 e32ce6f61e826cef675233a9f495988388ae7f8789474bc1bec30bf111cbf50b
MD5 12551f03e4ee0fec3cf9f01559a99a64
BLAKE2b-256 5773a3168ca7b86d665fcb0e2461c7f1c415092d9032476601480cc3c43554bd

See more details on using hashes here.

File details

Details for the file central_tools_llm-2025.11.0-py3-none-any.whl.

File metadata

File hashes

Hashes for central_tools_llm-2025.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 68acfe752e9ef9bb840b1ffc16c452aadf0f2428b1ad5dbf5d11b8b7738e6641
MD5 c66adc0045a5c910a7c82fbe3d233389
BLAKE2b-256 3caf988f8b9f8e97f370a0eaa5ccc9b306078677f1c50c75adf3b0de423daabf

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page