Skip to main content

Intelligent AI workload orchestrator. Multi-provider routing, cost optimization, semantic caching, dynamic routing with budget enforcement. Production observability included.

Project description

PyInferenceManager

Intelligent AI workload orchestrator. Multi-provider routing, cost optimization, semantic caching, dynamic routing with budget enforcement.

Status Python Tests Distribution License


Product Overview

PyInferenceManager is a proprietary, production-grade AI workload orchestrator. Route inference requests across multiple providers (Anthropic, OpenAI, etc.) with cost optimization and intelligent fallback.

Why Teams Choose This

The Problem:

  • Single-provider lock-in creates vendor risk
  • No visibility into inference costs
  • Failures cascade without intelligent fallback
  • Manual routing decisions waste engineering time

The Solution:

  • Route to any provider with single API change
  • Cost-aware routing with budget enforcement
  • Semantic caching to reduce redundant calls
  • Automatic fallback on provider failures
  • Comprehensive cost tracking and attribution

Result: Reduce inference costs 30-40%, eliminate vendor lock-in, improve reliability.


Installation

pip install pyinferencemanager
# or with uv
uv pip install pyinferencemanager

Requirements

  • Python 3.10+
  • Precompiled wheels for macOS, Linux, Windows

Distribution Model

Proprietary-first distribution:

  • ✅ Wheels-only via PyPI (no source code)
  • ✅ Production-optimized multi-provider routing
  • ✅ 261 comprehensive tests
  • ✅ Used in production ML systems

Quick Start

from pyinferencemanager import OrchestrationEngine, ProviderConfig

# Configure multiple providers
config = OrchestrationEngine.Config(
    providers=[
        ProviderConfig(name='anthropic', api_key='...'),
        ProviderConfig(name='openai', api_key='...'),
    ],
    budget=1000.00,  # Monthly budget
    cost_optimization=True,
)

engine = OrchestrationEngine(config)

# Single API for all providers
response = engine.infer(
    prompt="Analyze this dataset...",
    model='claude-3-sonnet',
    fallback_to=['gpt-4', 'claude-opus'],
    budget_tier='standard',  # or 'economy', 'premium'
)

print(f"Response: {response.text}")
print(f"Cost: ${response.cost:.4f}")
print(f"Provider used: {response.provider}")

Features

  • Multi-Provider Routing: Anthropic, OpenAI, and custom providers
  • Cost Optimization: Automatic provider selection based on cost
  • Semantic Caching: Reduce redundant API calls
  • Budget Enforcement: Hard stop when monthly budget exceeded
  • Intelligent Fallback: Automatic failover to alternate providers
  • Cost Attribution: Track spending by model, user, team
  • Health Monitoring: Provider availability and latency tracking
  • Observability: OpenTelemetry instrumentation

Quality & Testing

  • 261 tests passing
  • Production-grade — used in real ML systems
  • Observability — cost tracking, performance metrics
  • Reliability — intelligent fallback and retry logic

Support

For production deployments: mullassery@gmail.com


Version: 0.3.1
License: Proprietary
Distribution: Wheels-only via PyPI
Python: 3.10+

Built for production AI systems.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pyinferencemanager-0.4.1-cp310-abi3-macosx_11_0_arm64.whl (3.5 MB view details)

Uploaded CPython 3.10+macOS 11.0+ ARM64

File details

Details for the file pyinferencemanager-0.4.1-cp310-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for pyinferencemanager-0.4.1-cp310-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 f2fb6c68debb990bd26869a2daac545a148297fcfd96a696b6ccc81f7ab179c0
MD5 085155569e8d1cd5bca7107e98ed6f3e
BLAKE2b-256 0991dfa31e2d21e78e7f256c0fefc71b616f36fd344fc386aee3dccb1240aba8

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page