hotdata-materialized
Materialize expensive Django query results into Hotdata.
On a cache miss your code runs in your environment as usual and the result
is captured into Hotdata as parquet — in the background, after your response
has already returned. On a hit your database is never touched: the data
comes back as a pyarrow.Table you can turn into a dataframe, or query with
SQL server-side.
Think materialized views, not Redis: entries are snapshots of expensive analytical results, with a TTL lifecycle, each living in its own Hotdata managed database. The host application needs no migrations and no Redis — entry metadata lives in a small Hotdata managed database (the registry), created on first use.
Measured against TPC-H on a Neon Postgres (see demo/):
| Direct on Postgres | Miss (perceived) | Hit | |
|---|---|---|---|
| Q1 pricing summary (full scan, 4 rows) | ~1.4 s | ~1.4 s | ~80 ms |
| Revenue by part (50,000 rows) | ~2.1 s | ~2.0 s | ~260 ms |
A miss costs the same as not caching (the persist runs write-behind); a hit is 8–18× faster with zero load on your database.
Install
pip install hotdata-materialized
# settings.py
HOTDATA_MATERIALIZED = {
"API_KEY": env("HOTDATA_API_KEY"), # hd_... key from app.hotdata.dev
"WORKSPACE_ID": "work...",
# optional:
# "API_URL": "https://api.hotdata.dev",
# "REGISTRY_DATABASE": "materialized_registry",
# "INLINE_THRESHOLD_BYTES": 2 * 1024 * 1024,
# "BACKGROUND": True, # write-behind persists (default)
}
Using it in a Django view
Wrap the expensive part in @materialize; call it from your view:
# reports.py
from django.db.models import Count, Sum
from hotdata_materialized import materialize
from .models import Order
@materialize(ttl=3600)
def revenue_by_region():
return (
Order.objects
.filter(status="complete")
.values("region")
.annotate(orders=Count("id"), revenue=Sum("total"))
.order_by("-revenue")
)
# views.py
from django.http import JsonResponse
from .reports import revenue_by_region
def revenue_view(request):
frame = revenue_by_region()
return JsonResponse({"rows": frame.to_pylist(), "cached": frame.cached})
On a hit the function never runs and your database is never touched. On a
miss the function runs exactly as it would uncached, the caller gets the
result immediately, and the persist happens in a background thread. Either
way you get a MaterializedFrame:
frame.arrow()— the data as apyarrow.Table(frame.df()for pandas,frame.to_pylist()for dicts,len(frame)for the row count)frame.sql("SELECT region, revenue FROM this ORDER BY revenue DESC LIMIT 5")— SQL runs server-side against the cached entry;thisnames the data. Needs a persisted entry: available on hits (on a fresh miss the write-behind persist has to land first;background=Falseavoids that)frame.cached— which path served it;frame.entry— the registry record
The wrapped function can return an iterable of dicts (a .values() queryset),
a pyarrow.Table, or a pandas DataFrame. Decorator knobs: ttl= (seconds,
None = never expires), version= (bump to bust the cache on code changes),
key= (human-readable label), key_fn= (stable identity for arguments that
aren't JSON-serializable), background=False (block until persisted).
Fail-open by design: if Hotdata is unreachable, the function runs and its
result is served uncached — the cache degrades to "no cache," never to
"no page."
Search
Declare a search index on the decorator and the entry gets it when it persists (builds run server-side; a freshly missed entry may error on search for a few seconds until its index is ready):
from hotdata_materialized import BM25, Vector, materialize
@materialize(ttl=3600, index=BM25("notes"))
def incidents():
return Incident.objects.values("id", "notes", "severity")
incidents().search("checkout errors outage", column="notes", limit=5)
Vector (semantic) search embeds your text server-side via a workspace embedding provider — pass the indexed source column and a natural-language query:
@materialize(ttl=3600, index=Vector("notes", provider="emb_...", metric="cosine"))
def incidents_semantic():
return Incident.objects.values("id", "notes")
incidents_semantic().vector_search("the site was down", column="notes", limit=5)
Both return the matching rows as a pyarrow.Table with a relevance column
(score / distance). One platform rule to know: an embedding-backed vector
index cannot share an entry with other indexes — the decorator rejects that
combination up front.
Evicting entries
The decorator composes public primitives (fingerprint_call /
fingerprint_queryset, Registry, EntryStore) that you can use directly —
for example to drop an entry explicitly:
from hotdata_materialized import fingerprint_call
from hotdata_materialized.decorator import get_runtime
registry, store = get_runtime()
store.evict(fingerprint_call(revenue_by_region))
Expired entries stop being served immediately (the is_expired check) and
their databases carry a server-side expiry backstop; a sweep command that
deletes them proactively is on the roadmap below.
How it works
One managed database (the registry) holds one row per entry: fingerprint, status, TTL, the entry's database id, and — for small results — the data itself as an inline Arrow payload, so a chart-sized hit is served in a single API round trip. Larger results live in a per-entry managed database and are fetched as an Arrow IPC stream via a result id minted at persist time (no re-query on hits). Failures are loud and fail toward "no cache": a failed persist leaves no registry row, and the next request simply misses again.
See DESIGN.md for the architecture and the accepted trade-offs.
Roadmap
- Content-addressed fingerprinting (querysets and function calls)
- Remote registry — no migrations or local state in the host app
- Entry store with write-behind persists
- Arrow-native reads (inline payloads, persisted-result fetch)
-
@materializedecorator andMaterializedFrame(.arrow()/.df()/.to_pylist()/.sql()) - TPC-H demo and benchmark
- CI: pytest, mypy, flake8
- Stale-while-revalidate refresh (rebuild protocol)
- Sweep command for expired entries
- Chainable queryset facade on the frame (
.filter()/.order_by()) - Vector/BM25 index declarations,
frame.search()andframe.vector_search() - Async view support (
await frame.aarrow()) - PyPI release
Demo
demo/ is a small Django project that benchmarks TPC-H queries on a Neon Postgres directly vs. through the cache:
cd demo
python manage.py compare # Q1: tiny result, inline hit
python manage.py compare --scenario parts # 50k rows: remote Arrow hit
Development
uv venv && uv pip install -e '.[test]'
uv run pytest
Metadata
Release files for hotdata-materialized 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hotdata_materialized-0.1.0.tar.gz | 28.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hotdata_materialized-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 51.1 kB
Release files / hotdata_materialized-0.1.0.tar.gz
| Download URL | hotdata_materialized-0.1.0.tar.gz |
|---|---|
| Size | 28.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
123fe4eb8e8b20369fd0ec329cc397edc11163e1bcec942ed59bcd9ddd6ed1a5
|
|
BLAKE2b-256 checksum How to use checksums |
205eb0e86b8a46e55e57d722e12a8f1ee828943ad2df25640e510828b19f9e69
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 22, 2026.
Transparency logRelease files / hotdata_materialized-0.1.0-py3-none-any.whl
| Download URL | hotdata_materialized-0.1.0-py3-none-any.whl |
|---|---|
| Size | 22.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bf36406dec88c94ada3d9fe8f547586631fdc189997e1a226c4a52acc45c5ae4
|
|
BLAKE2b-256 checksum How to use checksums |
61cbbbf7eb9308333f2134557fb73caffae551bd6af3807a300ab2dca023184f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 22, 2026.
Transparency log