🚀 llmcachex-ai
Drop-in semantic + exact caching layer for LLM applications (RAG, agents, chatbots)
Save up to 80% LLM cost and reduce latency by avoiding repeated model calls using intelligent caching.
⚡ Installation
pip install llmcachex-ai
✨ Why llmcachex-ai?
Most LLM applications repeatedly call the model for:
- Slightly rephrased queries
- Agent/tool loops
- Chat history variations
This leads to higher latency and unnecessary cost.
llmcachex-ai solves this automatically by caching responses intelligently using exact + semantic matching.
🔥 Features
- ⚡ Exact cache (Redis-backed)
- 🧠 Semantic cache (FAISS + embeddings)
- 🔍 Hybrid retrieval (BM25 + vector search)
- 🧬 Cross-encoder reranking (high-quality matches)
- 🤖 Works with agents and tools
- 🧵 Memory-aware context support
- 💰 Token usage and cost tracking
- 🧩 Plug-and-play decorator API
🏗️ How It Works
User Query
↓
llm_cache decorator
├── Exact Cache (Redis)
├── Semantic Engine
│ ├── FAISS (vector)
│ ├── BM25 (lexical)
│ └── CrossEncoder (rerank)
└── LLM / Agent
🚀 Quick Start
from llm_cachex import llm_cache, CacheConfig
@llm_cache(CacheConfig())
def ask_llm(prompt):
return llm(prompt)
print(ask_llm("What is AI?")) # LLM call
print(ask_llm("Explain AI")) # Semantic cache hit
🤖 Agent Example
Works seamlessly with tools:
@llm_cache(CacheConfig())
def agent(raw_query, full_prompt):
if "calculate" in raw_query:
return str(eval(raw_query.replace("calculate", "").strip()))
if "search" in raw_query:
return f"[TOOL SEARCH RESULT] {raw_query}"
return llm(full_prompt)
🧠 Semantic Cache (Why it’s powerful)
"What is AI?"
"Explain artificial intelligence"
Both return the same cached response — no additional LLM call required.
⚙️ Configuration
CacheConfig(
enable_exact=True,
enable_semantic=True,
similarity_threshold=0.7,
top_k=3,
model_name="gpt-4o-mini",
enable_metrics=True,
enable_token_cost=True
)
📊 Metrics
from llm_cachex import metrics
print(metrics.summary())
Example output:
{
"hits": 2,
"misses": 1,
"hit_rate": 66.67,
"avg_llm_latency_ms": 2000,
"avg_cache_latency_ms": 30,
"total_cost_rupees": 0.01
}
🎯 Use Cases
- RAG pipelines
- AI agents & tool execution
- Chatbots with memory
- Cost optimization for LLM APIs
- High-frequency query systems
⚡ Performance Impact
Typical improvements:
- 2–10x latency reduction
- 50–80% cost savings
📁 Project Structure
llm_cachex/
├── api/ # decorator layer
├── core/ # cache, metrics, memory
├── semantic/ # hybrid search + reranker
├── embedding/ # embeddings
├── index/ # FAISS index
├── similarity/ # similarity utils
├── utils/ # helpers
🧭 Roadmap
- Async support
- Streaming support
- Batch inference
- Multi-model caching
- Pluggable vector DBs (Chroma / Pinecone)
- Observability dashboard
🤝 Contributing
Contributions are welcome. Open an issue to discuss ideas or submit a PR.
📜 License
MIT License
👤 Author
Himanshu Singh
⭐ Support
If this project helps you, consider giving it a star ⭐ on GitHub.
Metadata
Release files for llmcachex-ai 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llmcachex_ai-0.1.2.tar.gz | 13.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llmcachex_ai-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 28.4 kB
Release files / llmcachex_ai-0.1.2.tar.gz
| Download URL | llmcachex_ai-0.1.2.tar.gz |
|---|---|
| Size | 13.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8e23340602bae3f7c5ae9c9ea9e1078a1ef8baec6e87a0957d31864001ac06de
|
|
BLAKE2b-256 checksum How to use checksums |
1e10984f0af4f93d4be00b503d69a6e60bab2729505e9e0117f40e27f7cb4f7d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.0
|
Release files / llmcachex_ai-0.1.2-py3-none-any.whl
| Download URL | llmcachex_ai-0.1.2-py3-none-any.whl |
|---|---|
| Size | 15.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a0421112afdd8682ae0c81180ed86913e752d7b48d08226376198dbdb9535191
|
|
BLAKE2b-256 checksum How to use checksums |
679fa7d0633ee3af505c68d3b2d5bdf00621d0c2ae856e4eed2ec3e74f3839fe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.0
|