Skip to main content

hearth

Deterministic model residency. Keep declared models warm, and tell the truth about which ones are.

Not an inference engine. llama.cpp and vLLM have spent years on kernels, samplers and tokenizers, and none of that is the problem. The problem is that no serving stack will promise a model stays loaded, and none of them can say why one stopped being.

The night this came from

A PIN operator on one rented RTX A6000 stopped answering. Requests hung, then failed. Four pull requests landed against the streaming path in a single evening and none of them were the cause — because the cause was never visible. Three completely different failures were arriving as the same timeout:

  • the runtime evicted a model to free VRAM,
  • the host detached the GPU and gave it to another tenant,
  • a 32B model was simply still loading.

One of those is a capacity problem you own. One is your provider's and no configuration will touch it. One is not a problem at all. A timeout cannot tell you which, so all three got "fixed" repeatedly and none of them went away.

Worse, the operator was being scored down for a card their host reclaimed. A reputation system fed that kind of data slowly deletes its own honest operators.

What it does

$ hearth status
42.0 / 44.2 GiB held  (5 declared, 2 admitted)
  muse-local:latest            resident for 14203s
  deepseek-r1:32b              loading for 47s
  gemma4:26b                   not admitted — short by 14.8 GiB
  qwen3.6:27b                  not admitted — short by 15.8 GiB
  gemma4-extract:31b           not admitted — short by 17.8 GiB

Three things, none of which you can get today:

1. Residency is a named state with a named reason.

Unknown    never probed
Loading    weights materializing — with elapsed, so it can say how long
Resident   loaded AND answering AND accounted for
Lost       with a reason: Evicted · GpuDetached · ProcessExited · Unhealthy
Failed     won't load, and what the runtime said
Stopped    unloaded on purpose

Evicted and GpuDetached are the two states nothing else reports, and they were the two that mattered. They are distinguished by one bit — whether the GPU was still present when the probe failed — and that bit is the entire diagnosis.

2. The card's size is arithmetic, checked before anything loads.

Declare four 20 GiB models on a 48 GiB card and no runtime errors. It loads, evicts, loads, evicts, forever, and presents as "the models got slow." hearth refuses the fifth model at declare time and tells you it was short by 17.8 GiB. Nothing is ever evicted to make room for a load — if it doesn't fit, the honest answer is that it doesn't fit.

3. Routers get an answer they can act on.

answer what a router should do
Ready send it
Warming{for_ms} wait, or try elsewhere — but do not fault this node
Lost{GpuDetached} try elsewhere, and do not score this operator down
Lost{Evicted} try elsewhere; this box is over-committed
NotAdmitted{short} stop asking — it will never fit here
Unknown we genuinely don't know yet, and we say so

Design rules

Nothing fails on a clock. A 32B materializing over a network fabric can legitimately spend minutes before its first token, and killing it at an arbitrary deadline turns a slow success into a fast failure. Loading reports how long it has been loading; deciding what to do about that belongs to whoever knows if a human is waiting. Progress is reported — patience is a policy, not a constant.

Loading is not Ready. The most common way a serving stack lies is routing to something still coming up and calling the inevitable timeout an error.

Declaration order is priority order. First fit, never best fit. Reordering to squeeze in one more model would silently demote whatever the operator listed first, and on a serving box first means most important. A planner that outsmarts the operator surprises them at 3am.

The reserve is never planned into. Weights aren't the whole cost — KV cache grows with context and parallelism, the CUDA context is hundreds of megabytes, and fragmentation is real on a card that's been up for weeks.

Status

hearth-core — state machine, VRAM planner, fleet routing. 33 tests, all pure logic, no GPU required to run them.

Next: supervisor over llama-server children · NVML probe · HTTP surface (OpenAI-compatible + /residency) · CLI · napi + PyO3 bindings · NEDB event log.

That last one is the interesting one. Every state transition becomes an event in NEDB, which is bi-temporal — so "what was resident as of 03:14?" is a real query against a real causal chain. When a model goes cold at 3am you get the answer instead of a theory.

Integration

hearth speaks OpenAI-compatible, so pin-clientd works with it today: set apiMode: "openai" and point inferenceUri at hearth. /residency then adds the truth that the OpenAI shape has no way to express.

Build

cargo test          # 33 tests, no GPU needed
cargo run --example a6000

© Interchained LLC · BUSL-1.1 (converts to Apache-2.0 on 2030-08-27)

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hearth_engine-0.2.0.tar.gz (26.6 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

hearth_engine-0.2.0-cp39-abi3-win_amd64.whl (208.3 kB view details)

Uploaded CPython 3.9+Windows x86-64

hearth_engine-0.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (349.3 kB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ x86-64

hearth_engine-0.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (337.5 kB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ ARM64

hearth_engine-0.2.0-cp39-abi3-macosx_11_0_arm64.whl (299.7 kB view details)

Uploaded CPython 3.9+macOS 11.0+ ARM64

hearth_engine-0.2.0-cp39-abi3-macosx_10_12_x86_64.whl (307.7 kB view details)

Uploaded CPython 3.9+macOS 10.12+ x86-64

File details

Details for the file hearth_engine-0.2.0.tar.gz.

File metadata

  • Download URL: hearth_engine-0.2.0.tar.gz
  • Upload date:
  • Size: 26.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: maturin/1.15.0

File hashes

Hashes for hearth_engine-0.2.0.tar.gz
Algorithm Hash digest
SHA256 0e67871ddbc0e7f369580d50548edcf4bac6fec3d2db150e404871c7ad292a9c
MD5 0afa054400deb1ddcbd2a71246674054
BLAKE2b-256 a01e17232be61fd0a78e2b1c6d13c147dbb09a5d152060e9cf2daf486dea7dee

See more details on using hashes here.

File details

Details for the file hearth_engine-0.2.0-cp39-abi3-win_amd64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.2.0-cp39-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 7d054ec368d62835ae0b86884c9de7ebb26fff5a252e002bac8338f37e1c7056
MD5 033d31dcc4edf39dfd4923e0f6e3cde2
BLAKE2b-256 42ef11f78d5aa589c9f8c4401e84c4c941146da38435c1a85c5a2b927c12c47c

See more details on using hashes here.

File details

Details for the file hearth_engine-0.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.2.0-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 8e97623f3b17210f8a934f953e0b0dec417f67c9511efa8bcb938afa2017a031
MD5 897860470a468a386b389c6451d6fac6
BLAKE2b-256 ad4dc3d5c6f6b4d37c4a36be190b7af0d9c2f1a3d24fe05ec6cc9f2d70d91908

See more details on using hashes here.

File details

Details for the file hearth_engine-0.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.2.0-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 54060a6411f8865d15a0994dbfc0ae434f7e94e8c2b433c0f6c496f33a0a11a0
MD5 a73258fdd171e745c11864fc594f8c05
BLAKE2b-256 2e17e5f4f9723c4edfea251f7fc897af82cb76b3778927f2852c20c0070b2b7c

See more details on using hashes here.

File details

Details for the file hearth_engine-0.2.0-cp39-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.2.0-cp39-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 df23506bc50483173c3555fa13f7b9d57e45be798dd5ab602766eb1d1bdf9764
MD5 60492d702e34c0b9c7e0cda9e30ec6f5
BLAKE2b-256 78bc6a885d21db179477fbe78d346296ecc90e5e26194e28fdceea3d5d5742ef

See more details on using hashes here.

File details

Details for the file hearth_engine-0.2.0-cp39-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.2.0-cp39-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 75c7a0a612b027e1b747789439cb215d2b644a25887568ecd66eb19100dce09b
MD5 e0e2603700c5d260e6f59959d71a47bb
BLAKE2b-256 6d23150f3d3aedf71693cfe02a74076103c2c0ed8797632fc2e4b5a582d00490

See more details on using hashes here.

Release history Release notifications | RSS feed

0.4.7

6 files

0.4.6

6 files

0.4.5

6 files

0.4.3

6 files

0.4.2

6 files

0.4.1

6 files

0.4.0

6 files

0.3.9

6 files

0.3.8

6 files

0.3.7

6 files

0.3.6

6 files

0.3.5

6 files

0.3.4

6 files

0.3.3

6 files

0.3.2

6 files

0.3.1

6 files

0.3.0

6 files

This release

0.2.0 This release

6 files

0.1.2

6 files

0.1.1

6 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page