hearth-engine
Deterministic model residency, in Python. Keep declared models warm, and tell the truth about which ones are.
Native bindings to hearth's Rust core, built with PyO3 and maturin. No GPU required to use this package — it is the decision procedure, not the runtime.
pip install hearth-engine
import hearth
One abi3 wheel per platform covers Python 3.9+, so a new Python release does not need a new wheel.
Why you would want this
Your inference call hangs, then fails. Three completely different things cause that, and they arrive looking identical:
- the runtime evicted the model to free VRAM,
- the host detached the GPU and gave it to another tenant,
- a 32B model was simply still loading.
One is a capacity problem you own. One is your provider's, and no configuration you write will touch it. One is not a problem at all. A timeout cannot tell you which — so all three get "fixed" repeatedly and none of them go away.
Will this card hold this roster?
Answered before anything loads.
from hearth import GIB, declare, plan
p = plan(48 * GIB, [
declare("muse-local:latest", 20 * GIB, GIB),
declare("deepseek-r1:32b", 20 * GIB, GIB),
declare("gemma4:26b", 16 * GIB, GIB),
])
print(p["explain"])
# 2 of 3 admitted, 42.0 GiB committed of 44.2 GiB usable
# REJECTED gemma4:26b — needs 17.0 GiB, 2.2 GiB free, short by 14.8 GiB
for r in p["rejected"]:
print(f"{r['model']}: short by {r['short_bytes'] / GIB:.1f} GiB")
# gemma4:26b: short by 14.8 GiB
Declare five 20 GiB models on a 48 GiB card and no runtime will error. It loads, evicts, loads, evicts, forever, and presents to everyone as "the models got slow." This refuses the model that does not fit and tells you the shortfall.
Two rules in the planner are deliberate:
- Declaration order is priority order. First fit, never best fit — reordering to squeeze in one more model would silently demote whatever you listed first, and on a serving box first means most important.
- The reserve is never planned into (default 8%). Weights are not the whole cost: KV cache grows with context and parallelism, each CUDA context is hundreds of megabytes, and fragmentation is real on a card that has been up for weeks.
A live fleet, and an answer you can act on
from hearth import GIB, Fleet, declare
fleet = Fleet(48 * GIB, [declare("muse-local:latest", 20 * GIB, GIB)])
fleet.set_endpoint("muse-local:latest", "127.0.0.1:8090")
fleet.observe("muse-local:latest", "load_started")
r = fleet.route("muse-local:latest")
if r["ready"]:
send_to(r["endpoint"])
elif r["try_elsewhere"]:
retry_elsewhere(score_down=r["operator_fault"])
route() gives you three booleans answering three different questions:
| field | question |
|---|---|
ready |
can I send this request here, right now |
try_elsewhere |
should I go find another node |
operator_fault |
is this the operator's fault — False for a detached GPU |
That last one is the one nothing else reports. A reputation system fed the wrong answer slowly deletes its own honest operators.
The route key names which case you are in:
{"route": "unknown", "ready": False, "try_elsewhere": True, "operator_fault": False}
{"route": "warming", "for_ms": 20000, "ready": False, "try_elsewhere": False}
{"route": "ready", "endpoint": "127.0.0.1:8090", "ready": True}
{"route": "lost", "reason": "gpu_detached", "operator_fault": False, "try_elsewhere": True}
{"route": "lost", "reason": "evicted", "operator_fault": True, "try_elsewhere": True}
{"route": "not_admitted", "short_bytes": 19155554136, "try_elsewhere": True}
{"route": "not_declared", "try_elsewhere": True}
warming is the one every stack gets wrong: wait or route around, but do not fault this node. Routing to a model that is still coming up, then calling the inevitable timeout an error, is the most common way a serving stack lies about itself.
Key naming: every key this package returns is
snake_case, and every key it accepts is too. The Node package returns camelCase for the same data — each is idiomatic for its own language, so do not copy key names between the two.
Recording what you observed
fleet.observe(model, kind, detail=None, now=None)
kind is one of load_started · probe_ok · probe_failed · process_exited · load_failed · stop. now defaults to now_ms().
The single most important field you will ever pass here is gpu_present on a probe_failed:
fleet.observe("muse-local:latest", "probe_failed", {
"gpu_present": False, # the card is GONE — not this operator's fault
"detail": "no CUDA device",
})
Omit it and it reads as True — "the card was still there" — so a missing field can never quietly exonerate an operator. Absence has to be positively observed.
Either spelling works —
gpu_presentorgpuPresent, andvram_bytesorvramBytesonprobe_ok. Up to and including 0.3.2 this parser read only the snake_case spelling and an unrecognised key fell back to "the GPU was present", so{"gpuPresent": False}reportedevictedwithoperator_fault: True— the opposite of what the caller meant, silently. The mapping now lives inhearth-coreand is tested once, so the two bindings cannot drift apart again.
You pass facts, never conclusions. What state those produce is the core's job, decided by one state machine tested once in Rust rather than three times in three languages.
Everything, as one block
print(fleet.report())
# 0.0 / 44.2 GiB held (1 declared, 1 admitted)
# muse-local:latest loading for 5s
Same text hearth status prints — one truth in two places is how they stop matching.
Run the example
python examples/python/the_night.py
Replays the night hearth was built for: a model warms up, serves for an hour, the host takes the card away — and the router is told, in words, that it was not the operator's fault.
Also available
| package | registry |
|---|---|
@interchained/hearth |
npm |
hearth-core |
crates.io — the pure logic |
hearth-serve |
crates.io — the supervisor and hearth CLI |
Built by Vex × Interchained
© Interchained LLC · BUSL-1.1 (converts to Apache-2.0 on 2030-08-27)
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hearth_engine-0.4.6.tar.gz.
File metadata
- Download URL: hearth_engine-0.4.6.tar.gz
- Upload date:
- Size: 59.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b83ee399b9023df67594c03950383b1a9d4a83f2fd057004fbec343b0ff622ad
|
|
| MD5 |
52c1f0c6f4bea04a333747ec28d212a3
|
|
| BLAKE2b-256 |
88519a00136c235b0f70014dd8398f934243150c07e38ed40f4f92cc86ab676d
|
File details
Details for the file hearth_engine-0.4.6-cp39-abi3-win_amd64.whl.
File metadata
- Download URL: hearth_engine-0.4.6-cp39-abi3-win_amd64.whl
- Upload date:
- Size: 205.8 kB
- Tags: CPython 3.9+, Windows x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
16150616761da80c77f5f9646e9dd6ac772d5a3fefa6070c7c1e6a255727a00a
|
|
| MD5 |
b9ad14374f22ae2e93418739b25be22f
|
|
| BLAKE2b-256 |
0748e83159ad2cd3fc7d82f7f5b8552b826afd6d1daed4fd8f3024f85e9ab0b7
|
File details
Details for the file hearth_engine-0.4.6-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.
File metadata
- Download URL: hearth_engine-0.4.6-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
- Upload date:
- Size: 351.3 kB
- Tags: CPython 3.9+, manylinux: glibc 2.17+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2cfc48b80d388123df38b1635e54515733168501f3922ddc5dcee2a80b249917
|
|
| MD5 |
a31fc8701d682d09ac90f2284ce20f8c
|
|
| BLAKE2b-256 |
1a8b19a43d089577fa403154c7a4f655d2ba637484673afba13b3029fb13f94d
|
File details
Details for the file hearth_engine-0.4.6-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.
File metadata
- Download URL: hearth_engine-0.4.6-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
- Upload date:
- Size: 339.5 kB
- Tags: CPython 3.9+, manylinux: glibc 2.17+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
97841606558dc0a496ca1dd5d724cb52b50950def27f060422fdf65b3cf95825
|
|
| MD5 |
eaba2b1e81812da5389a1a817e9726d6
|
|
| BLAKE2b-256 |
e8b86904cbbd74f7d2041fe157b0bac78894d7aefaccfc7afba2f0a5803a7da7
|
File details
Details for the file hearth_engine-0.4.6-cp39-abi3-macosx_11_0_arm64.whl.
File metadata
- Download URL: hearth_engine-0.4.6-cp39-abi3-macosx_11_0_arm64.whl
- Upload date:
- Size: 301.5 kB
- Tags: CPython 3.9+, macOS 11.0+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bfd3fa40c594126ca41c1e49c9ee07d4f48092d569fc83a853b9d4b02d7db1aa
|
|
| MD5 |
ed80b8495b9bc1fed42b85707eae6e80
|
|
| BLAKE2b-256 |
af3f7012472d459e530d8017cf940ac55f3f3998fae9166f5bce440ce4b17b74
|
File details
Details for the file hearth_engine-0.4.6-cp39-abi3-macosx_10_12_x86_64.whl.
File metadata
- Download URL: hearth_engine-0.4.6-cp39-abi3-macosx_10_12_x86_64.whl
- Upload date:
- Size: 309.5 kB
- Tags: CPython 3.9+, macOS 10.12+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
maturin/1.15.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1006916f483ee8aa3e1781a8952fa4c2fc112c7561343412d8d16c5137636a61
|
|
| MD5 |
ec7b8c2c352e438dc7f424b6099b6ba7
|
|
| BLAKE2b-256 |
2cd732b6aeb3518bc166a11f09b599cad9494fc4cf07cafd279253912e77ac7c
|