Skip to main content

OmniCache

PyPI version License: MIT

OmniCache is a lightweight, local caching proxy for Anthropic (Claude), OpenAI (GPT), and Google (Gemini) APIs.

When developing with AI agents (like Claude Code, Cursor, Aider, or custom LLM scripts), repeated prompts, test runs, and static file queries frequently make duplicate upstream API calls. OmniCache sits between your client and upstream providers to intercept matching requests locally in <1ms, saving API costs and eliminating remote network latency.


Installation

pip install omnicache-proxy

Quickstart

1. Start the Proxy Server

omnicache

By default, the proxy runs on http://localhost:8000. You can change the port with --port:

omnicache --port 8080

2. Connect Your Client

Claude Code (Terminal CLI)

Set the Anthropic base URL environment variable before running claude:

export ANTHROPIC_BASE_URL="http://localhost:8000"
claude

Python (OpenAI SDK)

Route the base_url parameter to the local proxy:

from openai import OpenAI

client = OpenAI(
    api_key="your-api-key",
    base_url="http://localhost:8000/v1"
)

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "How do I reverse a linked list in Python?"}]
)
print(response.choices[0].message.content)

Cursor / VS Code / Other Tools

In your tool's model settings, set the API Base URL to http://localhost:8000/v1.


Key Features

  • Two-Tier Cache Engine:
    • L1 Exact Match (Trie Hash / Redis): Sub-0.05ms lookup for identical payloads.
    • L2 Semantic Match (Multi-Table LSH & FAISS ANN Indexing): Matches semantically equivalent prompts using high-speed multi-table hyperplane locality-sensitive hashing or FAISS HNSW.
    • Pluggable Embedders: Instant zero-dependency 512-d FastHashEmbedder or dense 384-d ONNX embeddings (all-MiniLM-L6-v2).
  • Horizontal Scaling & Redis Clustering:
    • Pluggable storage adapter architecture supporting both zero-dependency standalone mode and distributed multi-worker/multi-replica clusters.
    • Atomic spend tracking (INCRBYFLOAT) and sliding-window Redis RPM rate limiting across all worker processes.
    • Synchronized cluster-wide Circuit Breaker & upstream model failover state.
  • Asynchronous Write-Behind Persistence:
    • Micro-batched non-blocking worker queue writing to embedded SQLite WAL store off the critical path with zero latency impact.
    • Durable Virtual Key and spend budget ledger surviving process cold starts.
  • Remote Authenticated MCP JSON-RPC 2.0 Transport:
    • Native /mcp and /v1/mcp endpoint enabling AI IDEs (Cursor, Claude Code, VS Code) to perform intent-gated caching, vector search, and cache invalidation over HTTP.
  • Agent Stream Replayer: Emulates natural token-streaming for cached responses so interactive CLIs (like Claude Code) stream smoothly without terminal glitches.
  • Request Coalescing (SingleFlight): Deduplicates concurrent in-flight requests for the same prompt, making only one upstream call.
  • Built-in CLI Utilities:
    • omnicache doctor: Checks database state, port bindings, and embedder health.
    • omnicache benchmark: Measures P50, P95, and P99 cache lookup latencies on your machine.
    • omnicache stats: Prints total tokens and cost savings directly to the console.
  • Observability:
    • Web Dashboard: http://localhost:8000/dashboard
    • Prometheus Metrics: http://localhost:8000/metrics
    • CSV Ledger Export: http://localhost:8000/v1/cache/export

Horizontal Multi-Worker Deployment

To run OmniCache with multiple worker processes or in a clustered container environment, simply provide REDIS_URL:

# Multi-worker deployment with Redis distributed state
REDIS_URL="redis://127.0.0.1:6379/0" uvicorn server.gateway:app --host 127.0.0.1 --port 8000 --workers 4

Configuration

OmniCache can be configured via command-line flags or environment variables (in your shell or a local .env file):

Environment Variable Default Description
HOST 127.0.0.1 Host interface to listen on (local-first by default).
PORT 8000 Port to bind the proxy server to.
REDIS_URL "" Redis connection URL (e.g. redis://127.0.0.1:6379/0) for multi-worker state clustering.
CACHE_STORAGE_BACKEND auto Storage engine backend: auto, redis, or memory.
EMBEDDER_BACKEND fast_hash Semantic embedder: fast_hash, onnx, or auto.
ANN_INDEX_ENABLED true Enables sub-millisecond Approximate Nearest Neighbor vector search.
REQUIRE_AUTH false When true, enforces valid API key registration on all requests.
ADMIN_API_KEY "" Master admin secret for managing /v1/enterprise/quotas and data exports.
PRIVACY_SALT (auto-generated) 256-bit cryptographic salt for anonymized PII tokenization.
SEMANTIC_CACHE_TTL_SECONDS 604800 Default time-to-live for cache entries (7 days).
SEMANTIC_SIMILARITY_THRESHOLD 0.92 Minimum cosine similarity required for an L2 semantic cache hit.
OMNICACHE_DB_PATH ~/.omnicache/omnicache.db Path to SQLite persistence database (WAL mode enabled).
ANTHROPIC_API_KEY (Optional) Default upstream Anthropic API key (if not passed in client headers).
OPENAI_API_KEY (Optional) Default upstream OpenAI API key (if not passed in client headers).
GEMINI_API_KEY (Optional) Default upstream Google Gemini API key.

Running Tests

Run the test suite using pytest:

git clone https://github.com/13manmayarai-hash/omnicache-proxy.git
cd omnicache-proxy
pip install -e .
pytest tests/ -v

Documentation


License

MIT License. See LICENSE for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

omnicache_proxy-2.5.1.tar.gz (85.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

omnicache_proxy-2.5.1-py3-none-any.whl (73.3 kB view details)

Uploaded Python 3

File details

Details for the file omnicache_proxy-2.5.1.tar.gz.

File metadata

  • Download URL: omnicache_proxy-2.5.1.tar.gz
  • Upload date:
  • Size: 85.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: python-httpx/0.28.1

File hashes

Hashes for omnicache_proxy-2.5.1.tar.gz
Algorithm Hash digest
SHA256 27b9420517747cec314f59a7cd7275eb765dccd68a870dfc75f3fc429e5ba936
MD5 f32125cc588d5ac4ba684a321f65931b
BLAKE2b-256 db4d42a675c001d399a5b9940ebb743b604046ba365456d3a693f0774e954357

See more details on using hashes here.

File details

Details for the file omnicache_proxy-2.5.1-py3-none-any.whl.

File metadata

File hashes

Hashes for omnicache_proxy-2.5.1-py3-none-any.whl
Algorithm Hash digest
SHA256 4f794525be684e908c7a83dc8b06fa378c56eadfb4c68d8a70a92ce7f7762e1b
MD5 501fbe038fe47c3e455b25c412329937
BLAKE2b-256 307d92213bd58f174f9b236284cc7f1534ff14f7bc649a5423b2e21f89a7276b

See more details on using hashes here.

Release history Release notifications | RSS feed

2.9.0

2 files

2.8.0

2 files

2.7.2

2 files

2.7.1

2 files

2.7.0

2 files

2.6.8

2 files

2.6.7

2 files

2.6.6

2 files

2.6.5

2 files

2.6.4

2 files

2.6.3

2 files

2.6.2

2 files

2.6.1

2 files

2.6.0

2 files

2.5.9

2 files

2.5.8

2 files

2.5.7

2 files

2.5.6

2 files

2.5.5

2 files

2.5.4

2 files

2.5.3

2 files

2.5.2

2 files

This release

2.5.1 This release

2 files

2.5.0

2 files

2.3.0

2 files

2.2.1

2 files

2.2.0

2 files

2.1.4

2 files

2.1.3

2 files

2.1.2

2 files

2.1.1

2 files

2.1.0

2 files

2.0.5

2 files

2.0.4

2 files

2.0.3

2 files

2.0.2

2 files

2.0.1

2 files

2.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page