Skip to main content

PromptCache

Reduce your LLM API costs by 30--70% with semantic caching.

PromptCache reuses LLM responses for semantically similar prompts, not just exact string matches.

If two users ask:

  • "Explain Redis in simple terms"

  • "Can you explain Redis simply?"

You shouldn't pay twice.

PromptCache makes sure you don't.


License


The Problem

If you're using OpenAI or any LLM API in production, you're likely paying repeatedly for:

  • The same question phrased differently

  • Similar support requests across users

  • Slight variations in prompts

  • Background job retries

  • RAG pipelines returning near-identical queries

Traditional caching only works for exact matches.

LLMs need semantic caching.


What PromptCache Does

  1. Embeds your prompt into a vector

  2. Searches Redis for similar past prompts

  3. If similarity ≥ threshold → returns cached response

  4. Otherwise → calls the LLM and stores the result

User Prompt
     
Embed  Redis Vector Search
     
Hit?  Return cached answer
Miss?  Call LLM  Store result

10-Second Example

from promptcache import SemanticCache
from promptcache.backends.redis_vector import RedisVectorBackend
from promptcache.embedders.openai import OpenAIEmbedder
from promptcache.types import CacheMeta

embedder = OpenAIEmbedder(model="text-embedding-3-small")

backend = RedisVectorBackend(
    url="redis://localhost:6379/0",
    dim=embedder.dim,
)

cache = SemanticCache(
    backend=backend,
    embedder=embedder,
    namespace="support-bot",
    threshold=0.92,
)

meta = CacheMeta(
    model="gpt-4.1-mini",
    system_prompt="You are a helpful support assistant.",
)

result = cache.get_or_set(
    prompt="How do I reset my password?",
    llm_call=my_llm_call,
    extract_text=lambda r: r.output_text,
    meta=meta,
)

print(result.cache_hit)  # True or False`

That's it.


Example Impact

In a SaaS support assistant:

  • 62% cache hit rate

  • 48% reduction in token usage

  • 44% reduction in API spend

Your mileage depends on workload --- but high-volume, repetitive systems benefit the most.


Production-Ready Design

PromptCache isolates cache entries by:

  • namespace

  • model

  • system_prompt

  • tools_schema

  • embedder

This prevents cross-context contamination.

Additional features:

  • ✅ Redis HNSW vector search (cosine similarity)

  • ✅ TTL support

  • ✅ Hit-rate statistics

  • ✅ Optional cost tracking

  • ✅ In-memory backend (for testing)

  • ✅ Framework-agnostic (no LangChain dependency)


Installation

pip install promptcache-ai

Optional OpenAI embedder:

pip install promptcache-ai[openai]

Redis Setup

PromptCache requires Redis Stack (RediSearch with vector support).

Run locally:

docker run -d --name redis-stack -p 6379:6379 redis/redis-stack:latest

Verify:

redis-cli MODULE LIST

You should see:

search

Stats

Measure impact:

print(cache.stats())

Example:

{
    "hits": 1240,
    "misses": 860,
    "total": 2100,
    "hit_rate_percent": 59.05
}

When It Helps Most

  • Customer support bots

  • Internal copilots

  • FAQ systems

  • Knowledge assistants

  • Deterministic / low-temperature tasks

  • High-volume similar prompts


When It May Not Help

  • Highly personalized prompts

  • Creative high-temperature tasks

  • Frequently changing context


Testing

Run unit tests:

pytest

Run Redis integration tests:

export REDIS_URL="redis://localhost:6379/0"
pytest

Release files for promptcache-ai 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for promptcache-ai 0.2.0
File Size Uploaded
promptcache_ai-0.2.0.tar.gz 13.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for promptcache-ai 0.2.0
File Interpreter ABI Platform
promptcache_ai-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 26.8 kB

Release files / promptcache_ai-0.2.0.tar.gz

Download URL promptcache_ai-0.2.0.tar.gz
Size 13.1 kB
Tags Source
SHA-256 checksum
How to use checksums
a38431d4230b80b36f44cfd34edeef91ec134c0ed5505cea8daaecb4d5358ed5
BLAKE2b-256 checksum
How to use checksums
72bc960b25151b2abd64bb5c5a86e2a43ba59da1787b13308547c2b677876019
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Feb 20, 2026.

Transparency log

Release files / promptcache_ai-0.2.0-py3-none-any.whl

Download URL promptcache_ai-0.2.0-py3-none-any.whl
Size 13.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
07e3306d592b051eba5ec8cd5ea48b7835d4b822a880738e08b5c9b1973b4d2f
BLAKE2b-256 checksum
How to use checksums
28322c4a896fc06c04b10e0defa9c70fbed828e083400fe598399d2f61165b85
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Feb 20, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.5

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page