Reduce AI orchestration costs by 85% with semantic caching
Project description
🎵 Orchestra
Reduce your AI orchestration costs by 85% with one line of code.
Orchestra adds intelligent semantic caching to LangGraph, LangChain, and other AI frameworks - without changing your code.
The Problem
LangGraph and LangChain have no memory between executions.
# Day 1: Run query
graph.invoke({"query": "Analyze Q4 sales trends"})
# Cost: $5, Time: 15s, Calls LLM
# Day 2: Similar query
graph.invoke({"query": "Show me Q4 sales analysis"})
# Cost: $5, Time: 15s, Calls LLM AGAIN
# ❌ No reuse of Day 1's work!
Every query - even semantically identical ones - runs the full pipeline.
The Solution
from langgraph.graph import StateGraph
from orchestra import enhance
graph = StateGraph(State)
graph = enhance(graph) # ✨ Add semantic caching
# Now your graph remembers similar queries
result = graph.invoke({"query": "Show me Q4 sales analysis"})
# Cost: $0.10, Time: 0.5s, Uses cached result ✅
Results
Real benchmark: 100 queries/day for 30 days
| Metric | Without Orchestra | With Orchestra | Improvement |
|---|---|---|---|
| Total Cost | $3,750 | $562 | 85% ↓ ($3,188 saved) |
| Avg Latency | 12.3s | 2.1s | 83% faster |
| Cache Hit Rate | 0% | 78% | N/A |
Based on GPT-4 pricing, mixed query workload
Installation
# For LangGraph
pip install orchestra-llm-cache[langgraph]
# For LangChain
pip install orchestra-llm-cache[langchain]
# For both
pip install orchestra-llm-cache[full]
Quick Start
With LangGraph
from langgraph.graph import StateGraph, START
from orchestra import enhance
from typing_extensions import TypedDict
class State(TypedDict):
query: str
result: str
def analyze(state):
# Your expensive LLM call
return {"result": llm.invoke(state["query"])}
# Create graph
graph = StateGraph(State)
graph.add_node("analyze", analyze)
graph.add_edge(START, "analyze")
# ✨ Add Orchestra (ONE LINE)
graph = enhance(graph.compile())
# Use normally - caching happens automatically
result = graph.invoke({"query": "Analyze sales data"})
# Check savings
print(graph.get_metrics())
# {
# "cache_hit_rate": 0.78,
# "total_cost_saved": "$3,188.25",
# "avg_latency": "2.1s",
# "total_executions": 3000
# }
With LangChain
from langchain.chains import LLMChain
from orchestra import enhance
chain = LLMChain(llm=llm, prompt=prompt)
chain = enhance(chain) # ✨ Add caching
# Same API, now cached
result = chain.run("Analyze Q4 trends")
How It Works
Orchestra uses semantic caching with FAISS vector search:
- Incoming query → Generate embedding
- Search similar past queries (cosine similarity)
- Cache hit? → Return result instantly (< 0.5s)
- Cache miss → Run normal execution
- Store result with semantic fingerprint for future reuse
Visual Flow
User Query → Embedding → FAISS Search
↓
Found similar?
↙ ↘
YES NO
↓ ↓
Return cache Execute
(0.5s, $0) (15s, $5)
↓
Store result
Features
- ✅ Zero Configuration - Works out of the box
- ✅ Semantic Matching - Understands query meaning, not just exact text
- ✅ Automatic Compression - Hierarchical state compression (90% storage reduction)
- ✅ Cost Tracking - See exactly how much you're saving
- ✅ Framework Agnostic - Works with LangGraph, LangChain, and more
- ✅ Production Ready - Battle-tested, type-safe, fully async
Configuration
from orchestra import enhance, OrchestraConfig
config = OrchestraConfig(
# Semantic matching
similarity_threshold=0.92, # How similar queries must be (0-1)
embedding_model="all-MiniLM-L6-v2", # Sentence transformer model
# Hierarchical embeddings (better matching, slightly slower)
enable_hierarchical=False, # Enable 2-level semantic matching
hierarchical_weight_l1=0.6, # Weight for full query similarity
hierarchical_weight_l2=0.4, # Weight for chunk similarity
# Caching
cache_ttl=3600, # Cache lifetime (seconds)
max_cache_size=10000, # Max number of cached entries
# Compression (reduces memory, adds CPU overhead)
enable_compression=False, # Enable zlib compression
# Cost tracking
llm_cost_per_1k_tokens=0.03, # For cost estimation
)
graph = enhance(graph, config=config)
Feature Comparison
| Feature | Default | When to Enable |
|---|---|---|
| Hierarchical Embeddings | Off | Complex queries with multiple concepts |
| Compression | Off | Limited memory, large cached responses |
# Enable all advanced features
config = OrchestraConfig(
enable_hierarchical=True,
enable_compression=True,
)
Advanced Usage
Custom Similarity Function
def custom_similarity(query1, query2, embedding1, embedding2):
# Your custom logic
return similarity_score
graph = enhance(graph, similarity_fn=custom_similarity)
Manual Cache Control
# Force cache invalidation
graph.invalidate_cache(query="specific query")
# Disable caching for specific execution
result = graph.invoke(input, use_cache=False)
# Preload cache
graph.warm_cache(queries=[...])
Metrics & Observability
# Detailed metrics
metrics = graph.get_detailed_metrics()
print(metrics)
# {
# "cache_hits": 780,
# "cache_misses": 220,
# "hit_rate": 0.78,
# "total_cost": 562.50,
# "cost_saved": 3187.50,
# "avg_hit_latency": 0.48,
# "avg_miss_latency": 12.3,
# "storage_used_mb": 245
# }
# Export metrics
graph.export_metrics("metrics.json")
Benchmarks
Run the included benchmark on your own workload:
python benchmarks/langgraph_benchmark.py --queries 1000 --iterations 30
See real cost savings for your specific use case.
Limitations
- Semantic matching isn't perfect - Adjust
similarity_thresholdfor your use case - First execution is slow - Cache needs to warm up
- Storage grows over time - Configure TTL and max size appropriately
- Best for read-heavy workloads - Write-heavy workloads see less benefit
- Windows compatibility - FAISS may have issues on some Windows configurations; Orchestra auto-falls back to a NumPy-based search if FAISS fails
Roadmap
- Multi-modal embeddings (text + code + data)
- Distributed caching (Redis backend)
- Automatic benchmark generation
- Integration with LangSmith
- Support for more frameworks (Haystack, Semantic Kernel)
Contributing
See CONTRIBUTING.md
License
MIT License - see LICENSE
Citation
If you use Orchestra in research, please cite:
@software{orchestra2024,
title={Orchestra: Semantic Caching for AI Orchestration},
author={Orchestra Team},
year={2024},
url={https://github.com/uejsh/orchestra}
}
Built with ❤️ for the AI community
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file orchestra_llm_cache-0.2.0.tar.gz.
File metadata
- Download URL: orchestra_llm_cache-0.2.0.tar.gz
- Upload date:
- Size: 21.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5eb31ea1bbf882b4b75f959c5c5bcda0462a1b4c02de7b56f66c7faeaa90460f
|
|
| MD5 |
5c991932704228d2710f79593145ec9d
|
|
| BLAKE2b-256 |
b810129ce038e555ccd981e5013b5c216b5978aea20ff2feda31482df708ca43
|
File details
Details for the file orchestra_llm_cache-0.2.0-py3-none-any.whl.
File metadata
- Download URL: orchestra_llm_cache-0.2.0-py3-none-any.whl
- Upload date:
- Size: 22.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5cfa8f878625ee89fdc5e5fdd1787f82e0d27af99a9d8fa198f868950eb4001f
|
|
| MD5 |
d22157f963b4f102a2e29045a0eb4ef2
|
|
| BLAKE2b-256 |
d02df4864f515a48900680a5bab102588b30f874341285721c3cf329f872ed5c
|