Skip to main content

Smart LLM model router - auto-starts proviz-server binary, no Docker required

Project description

ProvizElekto

Smart LLM model router. Picks the best model for each call based on context size, rate limits, and capabilities — and retries automatically on failure.

Your app → pz.call(step, fn)               → CallResult
           pz.call_litellm(step, messages) → CallResult
                    ↕  (automatic)
           select → LLM call → report → retry on failure
                    ↕
              proviz-server (Rust)
          rate-limit state · catalog

Key difference from LiteLLM fallback: LiteLLM retries after failure. ProvizElekto picks the right model before the call — skipping models that are rate-limited or near their quota, can't fit the context, or lack required capabilities — then retries with the next eligible model automatically.

Two roles depending on the path

In the regular flow, the server is a pure router — it picks the model and returns credentials; your code makes the actual LLM call.

In the batch flow, the server becomes the caller:

# Regular: YOUR code calls the LLM
Your app → POST /select → ModelCandidate → your code → Mistral/OpenAI/...
                                                ↓
                                        POST /report

# Batch: the SERVER calls Mistral on your behalf
Worker A ──┐
Worker B ──┤ POST /batch/submit → server accumulates over window_secs
Worker C ──┘
                    ↓ server → POST Mistral /v1/batch/jobs (50% discount)
                    ↓ server polls until complete
Worker A ──┐
Worker B ──┤ GET /batch/result/{id} → response
Worker C ──┘

The batch path pools requests from all workers into a single Mistral job — the only way to qualify for Mistral's 50% batch discount. No individual worker can do this on its own, so the server acts as the aggregation point and makes the Mistral call itself.

Deployment note: when using batch, the server process (including Docker) must have the Mistral API key env vars set. In the regular flow, API keys only need to be present in the caller's environment.

Features

  • Context-aware selection - don't waste a 128k model on a 1k prompt
  • Proactive quota tracking - sliding-window counters (RPM/TPM/RPD/TPD) plus atomic in-flight reservations; avoids over-booking before any 429 fires
  • Provider-anchored windows - every successful call forwards x-ratelimit-remaining-* headers back to the server; the window floor is clamped to provider reality so internal estimates can't drift below what the provider actually sees
  • Scored selection - multi-component scoring: fast headroom (RPS/RPM/TPM, 25%), daily budget (RPD/TPD, 20%), quality (20%), cost (15%), latency (10%), traffic balance (10%). Over-quota models stay eligible with lower scores — AllModelsExhausted only fires when every model is in reactive 429 cooldown.
  • Traffic shaping - per-brand traffic_weight steers load proportionally across providers in a 5-minute rolling window; under-served brands get a higher score on the traffic component
  • Capability filtering - hard requirements for function calling, JSON mode
  • Quality floor - reject models below a quality threshold per step
  • Model groups - define named pools of models (e.g. "fast-chat", "coding-tier1") and restrict selection to that pool
  • Your keys, your models - curated catalog, no vendor proxy
  • Zero-infra - pip install proviz-elekto auto-starts the Rust server as a subprocess
  • Any language - HTTP API, not a library binding
  • Pluggable storage - SQLite (default) or PostgreSQL

Installation

ProvizElekto consists of a Rust server and various clients.

pip install proviz-elekto          # core only
pip install proviz-elekto[litellm] # + built-in LiteLLM integration

The proviz-server binary is bundled in the wheel.

CLI tool (proviz) is also included:

proviz --help

Documentation

Quickstart

With LiteLLM (recommended)

from proviz_elekto import ProvizElekto

pz = ProvizElekto(db_path="./proviz.db")
# or PostgreSQL: pz = ProvizElekto(database_url=os.environ["DATABASE_URL"])

result = pz.call_litellm(
    step="verdict",
    messages=[{"role": "user", "content": "Summarize this document..."}],
    estimated_tokens=2500,
    requires_json_mode=True,
)
print(result.provider, result.candidate.model_slug, result.total_tokens)
# → mistral mistral-small-latest 312

call_litellm() selects the best available model, calls it, reports the outcome, and retries with the next eligible model on any failure — automatically.

With a custom LLM caller

import anthropic

client = anthropic.Anthropic()

def my_llm(candidate):
    return client.messages.create(
        model=candidate.model_slug,
        max_tokens=1024,
        messages=[{"role": "user", "content": "Hello"}],
    )

result = pz.call("verdict", my_llm, estimated_tokens=100)
print(result.candidate.brand_slug, result.prompt_tokens)

Pass any callable that accepts a ModelCandidate and returns a response. ProvizElekto wraps it with the same select → report → retry loop.

Low-level API

If you need direct control over selection and reporting:

candidate = pz.select(step="verdict", estimated_tokens=2500)
try:
    response = my_llm_call(candidate)

    # Read provider rate-limit headers (Mistral/OpenAI style; Anthropic style also supported)
    hdrs = getattr(response, "_hidden_params", {}).get("additional_headers") or {}
    rem_req = hdrs.get("x-ratelimit-remaining-requests")
    rem_tok = hdrs.get("x-ratelimit-remaining-tokens")

    pz.report_success(
        candidate.model_id,
        estimated_tokens=candidate.estimated_tokens,  # releases in-flight reservation
        actual_tokens=response.usage.total_tokens,    # improves TPM window accuracy
        remaining_requests=int(rem_req) if rem_req is not None else None,
        remaining_tokens=int(rem_tok)   if rem_tok is not None else None,
    )
    # report_success is fire-and-forget — returns immediately, HTTP call runs in background
except RateLimitError as exc:
    msg = str(exc).lower()
    if "day" in msg or "daily" in msg:
        error_type = "tpd"
    elif "token" in msg:
        error_type = "tpm"
    else:
        error_type = "rpm"
    pz.report_rate_limit(candidate.model_id, error_type)  # synchronous — must complete before retry
except Exception:
    pz.report_error(candidate.model_id, "other")

estimated_tokens in each report call releases the in-flight reservation made at selection time. Omitting it is safe (legacy clients work unchanged) but leaves the in-flight counter inflated until the next selection clears it.

report_success is non-blocking: the HTTP call to proviz runs in a background daemon thread so the caller receives the LLM result without waiting for the round-trip. report_rate_limit and report_error remain synchronous because the model must be blocked in proviz before the retry select() call.

License

Apache-2.0

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

proviz_elekto-0.9.0-py3-none-win_amd64.whl (4.2 MB view details)

Uploaded Python 3Windows x86-64

proviz_elekto-0.9.0-py3-none-musllinux_1_2_x86_64.whl (5.0 MB view details)

Uploaded Python 3musllinux: musl 1.2+ x86-64

proviz_elekto-0.9.0-py3-none-manylinux_2_36_x86_64.whl (4.9 MB view details)

Uploaded Python 3manylinux: glibc 2.36+ x86-64

proviz_elekto-0.9.0-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (4.7 MB view details)

Uploaded Python 3manylinux: glibc 2.17+ ARM64

proviz_elekto-0.9.0-py3-none-macosx_10_12_x86_64.macosx_11_0_arm64.macosx_10_12_universal2.whl (9.1 MB view details)

Uploaded Python 3macOS 10.12+ universal2 (ARM64, x86-64)macOS 10.12+ x86-64macOS 11.0+ ARM64

File details

Details for the file proviz_elekto-0.9.0-py3-none-win_amd64.whl.

File metadata

File hashes

Hashes for proviz_elekto-0.9.0-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 80899edafe0ceba9779ae94dbc9ddaedd8bd2a66410ae3027963b2537edf730f
MD5 3508a8c9222265bbcb6da099c537354d
BLAKE2b-256 cd23592b8e4657c8b3e2b7bf1d60a7ab66186df37faf3224de53df83d4f4b224

See more details on using hashes here.

Provenance

The following attestation bundles were made for proviz_elekto-0.9.0-py3-none-win_amd64.whl:

Publisher: release.yml on JustGui/proviz-elekto

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file proviz_elekto-0.9.0-py3-none-musllinux_1_2_x86_64.whl.

File metadata

File hashes

Hashes for proviz_elekto-0.9.0-py3-none-musllinux_1_2_x86_64.whl
Algorithm Hash digest
SHA256 d9bedd0cbce1b6ba8dd1a58b623723f5f32b3345175f005664f936401ed54bcb
MD5 b8b3daceff49c840205ac3ae32472c88
BLAKE2b-256 02dfe3bf5d0d091a4f1b83b525568d981a38c877ca15a61f43e3f95dc3df5e1f

See more details on using hashes here.

Provenance

The following attestation bundles were made for proviz_elekto-0.9.0-py3-none-musllinux_1_2_x86_64.whl:

Publisher: release.yml on JustGui/proviz-elekto

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file proviz_elekto-0.9.0-py3-none-manylinux_2_36_x86_64.whl.

File metadata

File hashes

Hashes for proviz_elekto-0.9.0-py3-none-manylinux_2_36_x86_64.whl
Algorithm Hash digest
SHA256 e88ee29e93e39ff835dc252182018b174434c8b51c1c6f5d43b54f1ceda5b15b
MD5 f6e2a2fadcccdf4f7f33ebeb149a160a
BLAKE2b-256 55c0d9b1f0e93db3ca9a62a9d75581e5a10280f0acf186e546be8d60d1384711

See more details on using hashes here.

Provenance

The following attestation bundles were made for proviz_elekto-0.9.0-py3-none-manylinux_2_36_x86_64.whl:

Publisher: release.yml on JustGui/proviz-elekto

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file proviz_elekto-0.9.0-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for proviz_elekto-0.9.0-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 74bbaf995a4e096f2d4cfffc9a42ccbdedb2153bcfa001126568dfb4d9b6b2e3
MD5 ee8ea081e509c9c5a8100a21eee9788b
BLAKE2b-256 9a89eb7bf3244e3a660f9c54493ce510d5f6cc7eda8f3c1ceec5215b1331e8f6

See more details on using hashes here.

Provenance

The following attestation bundles were made for proviz_elekto-0.9.0-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on JustGui/proviz-elekto

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file proviz_elekto-0.9.0-py3-none-macosx_10_12_x86_64.macosx_11_0_arm64.macosx_10_12_universal2.whl.

File metadata

File hashes

Hashes for proviz_elekto-0.9.0-py3-none-macosx_10_12_x86_64.macosx_11_0_arm64.macosx_10_12_universal2.whl
Algorithm Hash digest
SHA256 047993147e000fbac97d85c9dd8c463ab7a7e9f3084fd14a8a46659a26970219
MD5 38908c93e514f59e7d29ad2b66a67e8e
BLAKE2b-256 1ebcd40a7d468fdd87999b8d936accec6b6a0bc80b1fa8c70cff321bbcf541ca

See more details on using hashes here.

Provenance

The following attestation bundles were made for proviz_elekto-0.9.0-py3-none-macosx_10_12_x86_64.macosx_11_0_arm64.macosx_10_12_universal2.whl:

Publisher: release.yml on JustGui/proviz-elekto

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page