Skip to main content

PyInferenceManager

Use any LLM. Switch models without rewriting code. Cut inference costs 40-60%.

Intelligently route requests across Claude, GPT-4, Gemini, Llama, Mistral, and more based on cost, speed, or availability. One line of code. Automatic failover. No vendor lock-in.

PyPI Python 3.10+ Tests License: Proprietary


30-Second Start

from pyinferencemanager import Manager

# Create manager (routes across all providers)
mgr = Manager()

# Same code. Different models. Different costs.
response = mgr.chat(
    "Which LLM should I use for this task?",
    prefer="cheapest"  # or "fastest" or "highest-quality"
)

print(response.text)
print(f"Used: {response.model}")  # Which provider was chosen?
print(f"Cost: ${response.cost:.4f}")  # How much did it cost?

Why PyInferenceManager?

The Problem:

  • Each LLM has different APIs (Claude, OpenAI, Google, Anthropic, etc.)
  • Costs vary wildly (GPT-4 is 10x more expensive than Llama)
  • You can't switch models without rewriting your code
  • Provider outages break your application

The Solution:

  • One unified API for all LLM providers
  • Automatic routing based on cost, speed, or quality
  • Provider failover (if Claude is down, switch to GPT-4 automatically)
  • Easy cost comparison and optimization

Key Features

  • 11 Providers: Claude (Anthropic), GPT-4/3.5 (OpenAI), Gemini (Google), Llama (Meta), Mistral, Cohere, PaLM, Falcon, and more
  • Smart Routing: Automatic selection based on cost/speed/quality
  • Cost Tracking: Real-time cost estimation and reporting
  • Failover: Automatic provider switching if one goes down
  • Batch Processing: Process 1000s of requests with automatic optimization
  • Streaming Support: Get responses as they arrive
  • Rate Limiting: Built-in quotas and backoff

Real-World Use Cases

Cost Optimization:

# Cheap tasks use Llama, complex tasks use Claude
response = mgr.chat(prompt, prefer="cheapest")
# Llama for summarization: $0.0001
# Claude for reasoning: $0.001
# Automatic choice based on task difficulty

Reliability:

# If Claude API is down, automatically use GPT-4
response = mgr.chat(prompt, fallback="gpt-4")

Multi-Model Comparison:

# Test a prompt across all providers
for model in ["claude", "gpt-4", "gemini", "llama"]:
    result = mgr.chat(prompt, model=model)
    print(f"{model}: ${result.cost}")

Provider Comparison

Provider Speed Cost Quality Notes
Claude 3 Opus Fast $$ Excellent Best reasoning
GPT-4 Medium $$$ Excellent General purpose
Gemini Fast $ Good Great value
Llama 2 Slow $ Good Local option
Mistral Fast $ Good European option

Installation

pip install pyinferencemanager
# or with uv
uv pip install pyinferencemanager

Set API keys (one time):

export ANTHROPIC_API_KEY=sk-...
export OPENAI_API_KEY=sk-...
export GOOGLE_API_KEY=goog-...

Documentation


License

Proprietary License - Free to use with explicit attribution. See LICENSE.


PyInferenceManager v2.0.0 | Smart LLM routing | Python 3.10+

License

MIT


MCP 2.0 Mega-Platform | v2.0.0 | Wheels-Only Distribution

Release files for pyinferencemanager 1.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for pyinferencemanager 1.1.0
File Interpreter ABI Platform
pyinferencemanager-1.1.0-cp310-abi3-macosx_11_0_arm64.whl CPython 3.10 abi3 macOS 11.0+ ARM64 Details

Release files / pyinferencemanager-1.1.0-cp310-abi3-macosx_11_0_arm64.whl

Download URL pyinferencemanager-1.1.0-cp310-abi3-macosx_11_0_arm64.whl
Size 3.5 MB
Tags CPython 3.10 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
a34c0f6b0635919b806936389c835f3e97b97f871780ce2f8c22aa0157c91cf4
BLAKE2b-256 checksum
How to use checksums
75db028886f9c57869effd66c97b6c54ae37336efbba5bdef050d2a073fc73bf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.6

Release history Release notifications | RSS feed

1.3.0

2 release files

1.2.0

2 release files

1.1.1

1 release file

This release

1.1.0 This release

1 release file

1.0.0

2 release files

0.4.2

1 release file

0.4.1

1 release file

0.4.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page