Skip to main content

🚀 llmcachex-ai

PyPI Python License

Drop-in semantic + exact caching layer for LLM applications (RAG, agents, chatbots)

Save up to 80% LLM cost and reduce latency by avoiding repeated model calls using intelligent caching.


⚡ Installation

pip install llmcachex-ai

✨ Why llmcachex-ai?

Most LLM applications repeatedly call the model for:

  • Slightly rephrased queries
  • Agent/tool loops
  • Chat history variations

This leads to higher latency and unnecessary cost.

llmcachex-ai solves this automatically by caching responses intelligently using exact + semantic matching.


🔥 Features

  • ⚡ Exact cache (Redis-backed)
  • 🧠 Semantic cache (FAISS + embeddings)
  • 🔍 Hybrid retrieval (BM25 + vector search)
  • 🧬 Cross-encoder reranking (high-quality matches)
  • 🤖 Works with agents and tools
  • 🧵 Memory-aware context support
  • 💰 Token usage and cost tracking
  • 🧩 Plug-and-play decorator API

🏗️ How It Works

User Query
   ↓
llm_cache decorator
   ├── Exact Cache (Redis)
   ├── Semantic Engine
   │     ├── FAISS (vector)
   │     ├── BM25 (lexical)
   │     └── CrossEncoder (rerank)
   └── LLM / Agent

🚀 Quick Start

from llm_cachex import llm_cache, CacheConfig

@llm_cache(CacheConfig())
def ask_llm(prompt):
    return llm(prompt)

print(ask_llm("What is AI?"))      # LLM call
print(ask_llm("Explain AI"))       # Semantic cache hit

🤖 Agent Example

Works seamlessly with tools:

@llm_cache(CacheConfig())
def agent(raw_query, full_prompt):

    if "calculate" in raw_query:
        return str(eval(raw_query.replace("calculate", "").strip()))

    if "search" in raw_query:
        return f"[TOOL SEARCH RESULT] {raw_query}"

    return llm(full_prompt)

🧠 Semantic Cache (Why it’s powerful)

"What is AI?"
"Explain artificial intelligence"

Both return the same cached response — no additional LLM call required.


⚙️ Configuration

CacheConfig(
    enable_exact=True,
    enable_semantic=True,
    similarity_threshold=0.7,
    top_k=3,
    model_name="gpt-4o-mini",
    enable_metrics=True,
    enable_token_cost=True
)

📊 Metrics

from llm_cachex import metrics

print(metrics.summary())

Example output:

{
  "hits": 2,
  "misses": 1,
  "hit_rate": 66.67,
  "avg_llm_latency_ms": 2000,
  "avg_cache_latency_ms": 30,
  "total_cost_rupees": 0.01
}

🎯 Use Cases

  • RAG pipelines
  • AI agents & tool execution
  • Chatbots with memory
  • Cost optimization for LLM APIs
  • High-frequency query systems

⚡ Performance Impact

Typical improvements:

  • 2–10x latency reduction
  • 50–80% cost savings

📁 Project Structure

llm_cachex/
├── api/            # decorator layer
├── core/           # cache, metrics, memory
├── semantic/       # hybrid search + reranker
├── embedding/      # embeddings
├── index/          # FAISS index
├── similarity/     # similarity utils
├── utils/          # helpers

🧭 Roadmap

  • Async support
  • Streaming support
  • Batch inference
  • Multi-model caching
  • Pluggable vector DBs (Chroma / Pinecone)
  • Observability dashboard

🤝 Contributing

Contributions are welcome. Open an issue to discuss ideas or submit a PR.


📜 License

MIT License


👤 Author

Himanshu Singh


⭐ Support

If this project helps you, consider giving it a star ⭐ on GitHub.

Metadata

Release files for llmcachex-ai 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llmcachex-ai 0.1.2
File Size Uploaded
llmcachex_ai-0.1.2.tar.gz 13.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llmcachex-ai 0.1.2
File Interpreter ABI Platform
llmcachex_ai-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 28.4 kB

Release files / llmcachex_ai-0.1.2.tar.gz

Download URL llmcachex_ai-0.1.2.tar.gz
Size 13.0 kB
Tags Source
SHA-256 checksum
How to use checksums
8e23340602bae3f7c5ae9c9ea9e1078a1ef8baec6e87a0957d31864001ac06de
BLAKE2b-256 checksum
How to use checksums
1e10984f0af4f93d4be00b503d69a6e60bab2729505e9e0117f40e27f7cb4f7d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.0

Release files / llmcachex_ai-0.1.2-py3-none-any.whl

Download URL llmcachex_ai-0.1.2-py3-none-any.whl
Size 15.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a0421112afdd8682ae0c81180ed86913e752d7b48d08226376198dbdb9535191
BLAKE2b-256 checksum
How to use checksums
679fa7d0633ee3af505c68d3b2d5bdf00621d0c2ae856e4eed2ec3e74f3839fe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.0

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page