Skip to main content

hearth

Deterministic model residency. Keep declared models warm, and tell the truth about which ones are.

Not an inference engine. llama.cpp and vLLM have spent years on kernels, samplers and tokenizers, and none of that is the problem. The problem is that no serving stack will promise a model stays loaded, and none of them can say why one stopped being.

The night this came from

A PIN operator on one rented RTX A6000 stopped answering. Requests hung, then failed. Four pull requests landed against the streaming path in a single evening and none of them were the cause — because the cause was never visible. Three completely different failures were arriving as the same timeout:

  • the runtime evicted a model to free VRAM,
  • the host detached the GPU and gave it to another tenant,
  • a 32B model was simply still loading.

One of those is a capacity problem you own. One is your provider's and no configuration will touch it. One is not a problem at all. A timeout cannot tell you which, so all three got "fixed" repeatedly and none of them went away.

Worse, the operator was being scored down for a card their host reclaimed. A reputation system fed that kind of data slowly deletes its own honest operators.

What it does

$ hearth status
42.0 / 44.2 GiB held  (5 declared, 2 admitted)
  muse-local:latest            resident for 14203s
  deepseek-r1:32b              loading for 47s
  gemma4:26b                   not admitted — short by 14.8 GiB
  qwen3.6:27b                  not admitted — short by 15.8 GiB
  gemma4-extract:31b           not admitted — short by 17.8 GiB

Three things, none of which you can get today:

1. Residency is a named state with a named reason.

Unknown    never probed
Loading    weights materializing — with elapsed, so it can say how long
Resident   loaded AND answering AND accounted for
Lost       with a reason: Evicted · GpuDetached · ProcessExited · Unhealthy
Failed     won't load, and what the runtime said
Stopped    unloaded on purpose

Evicted and GpuDetached are the two states nothing else reports, and they were the two that mattered. They are distinguished by one bit — whether the GPU was still present when the probe failed — and that bit is the entire diagnosis.

2. The card's size is arithmetic, checked before anything loads.

Declare four 20 GiB models on a 48 GiB card and no runtime errors. It loads, evicts, loads, evicts, forever, and presents as "the models got slow." hearth refuses the fifth model at declare time and tells you it was short by 17.8 GiB. Nothing is ever evicted to make room for a load — if it doesn't fit, the honest answer is that it doesn't fit.

3. Routers get an answer they can act on.

answer what a router should do
Ready send it
Warming{for_ms} wait, or try elsewhere — but do not fault this node
Lost{GpuDetached} try elsewhere, and do not score this operator down
Lost{Evicted} try elsewhere; this box is over-committed
NotAdmitted{short} stop asking — it will never fit here
Unknown we genuinely don't know yet, and we say so

Design rules

Nothing fails on a clock. A 32B materializing over a network fabric can legitimately spend minutes before its first token, and killing it at an arbitrary deadline turns a slow success into a fast failure. Loading reports how long it has been loading; deciding what to do about that belongs to whoever knows if a human is waiting. Progress is reported — patience is a policy, not a constant.

Loading is not Ready. The most common way a serving stack lies is routing to something still coming up and calling the inevitable timeout an error.

Declaration order is priority order. First fit, never best fit. Reordering to squeeze in one more model would silently demote whatever the operator listed first, and on a serving box first means most important. A planner that outsmarts the operator surprises them at 3am.

The reserve is never planned into. Weights aren't the whole cost — KV cache grows with context and parallelism, the CUDA context is hundreds of megabytes, and fragmentation is real on a card that's been up for weeks.

Status

hearth-core — state machine, VRAM planner, fleet routing. Pure logic, no GPU required. hearth-resolve — one reference syntax over the Ollama registry and HuggingFace GGUFs, verified against the live registries. hearth-storethe NEDB spine. Every residency transition is a bi-temporal, causally-linked, tamper-evident event in an embedded NEDB. Not a log on the side: the supervisor's memory IS the database. hearth-serve — the supervisor. llama-server children, honest health probes, gpu_present via nvidia-smi (an honest None on CPU boxes), SIGTERM records unloaded and reaps children. Plus the hearth CLI.

Proven end-to-end on a real model: hearth serve brought a GGUF resident under a real llama-server (warmup measured, not guessed), served real completion tokens, took a SIGTERM, and a fresh process then read the whole story back off disk:

$ hearth why stories
● seq     3  unloaded       {}
└─ seq     2  resident       {"endpoint":"127.0.0.1:18080","warmup_ms":256}
└─ seq     1  loading        {"endpoint":"127.0.0.1:18080","pid":10121}
└─ seq     0  declared       {"admitted":true,"vram_bytes":1073741824}

$ hearth as-of stories 2
as of seq 2: stories was `resident` {"endpoint":"127.0.0.1:18080","warmup_ms":256}

$ hearth verify
verify ok — 4 nodes checked, history intact

What was resident as of seq 2 is a real query against a real causal chain. When a model goes cold at 3am you get the answer instead of a theory — and verify proves nobody rewrote it. No other serving stack can print that.

Next: hearth pull wired to the resolvers + blob store · HTTP surface (OpenAI-compatible proxy + /residency) · multi-model fleets under one supervisor · napi + PyO3 bindings for store/serve.

Integration

hearth speaks OpenAI-compatible, so pin-clientd works with it today: set apiMode: "openai" and point inferenceUri at hearth. /residency then adds the truth that the OpenAI shape has no way to express.

Build

cargo test --workspace   # no GPU needed
cargo run --example a6000

# serve a model under supervision (needs llama.cpp's llama-server)
hearth serve --model muse --gguf ./muse.gguf --port 8080
hearth status && hearth why muse && hearth verify

© Interchained LLC · BUSL-1.1 (converts to Apache-2.0 on 2030-08-27)

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hearth_engine-0.3.2.tar.gz (35.4 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

hearth_engine-0.3.2-cp39-abi3-win_amd64.whl (205.3 kB view details)

Uploaded CPython 3.9+Windows x86-64

hearth_engine-0.3.2-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (349.9 kB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ x86-64

hearth_engine-0.3.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (338.2 kB view details)

Uploaded CPython 3.9+manylinux: glibc 2.17+ ARM64

hearth_engine-0.3.2-cp39-abi3-macosx_11_0_arm64.whl (300.3 kB view details)

Uploaded CPython 3.9+macOS 11.0+ ARM64

hearth_engine-0.3.2-cp39-abi3-macosx_10_12_x86_64.whl (308.3 kB view details)

Uploaded CPython 3.9+macOS 10.12+ x86-64

File details

Details for the file hearth_engine-0.3.2.tar.gz.

File metadata

  • Download URL: hearth_engine-0.3.2.tar.gz
  • Upload date:
  • Size: 35.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: maturin/1.15.0

File hashes

Hashes for hearth_engine-0.3.2.tar.gz
Algorithm Hash digest
SHA256 512f1468ca9ddc15eb238e30469ca4cc63718e98337e478d6b04aa1869635534
MD5 f4aac337b585aa4074de9c7dd1f92c37
BLAKE2b-256 196d90c6b8a0af9722b512b74c46942ebb08f7e1b86a5bacd289b548da922d39

See more details on using hashes here.

File details

Details for the file hearth_engine-0.3.2-cp39-abi3-win_amd64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.3.2-cp39-abi3-win_amd64.whl
Algorithm Hash digest
SHA256 2f197f9c5cf53e95ec970bc97d62f3c0df0016e8343cffe182adf9b63750d2da
MD5 a16e0691917cec0c4d8157046d411f38
BLAKE2b-256 b87fe29aa1f788bdf507b9e7b4d19b0477688e94fbc0575b4e9c65120b14d911

See more details on using hashes here.

File details

Details for the file hearth_engine-0.3.2-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.3.2-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 e5e1d202033cc8ca8da76056d79d0242d6afe2b491dd86d501b9e26b4db91cd6
MD5 a4d32face977aa7f19aa5a5b60f902f4
BLAKE2b-256 a182fe80824dea7f55b5b34259cf44ef7991f2fe299ce3f13d8392c3830edd5a

See more details on using hashes here.

File details

Details for the file hearth_engine-0.3.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.3.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 73aa71f31091fa67d16ee055df7f5bea29a2ede43e6ef3194c59c3d143db3572
MD5 5423d01691874e45cd316b8bff2deaf5
BLAKE2b-256 7b5cfaae0c8a54a7b71c613b90be08a326273e8bd215af5e1879ac73b634a5d7

See more details on using hashes here.

File details

Details for the file hearth_engine-0.3.2-cp39-abi3-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.3.2-cp39-abi3-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 4ce3af81ab5c2e2583895d9f5d327cf3a3d5ae12e3d3c44dfd6e7a936566a321
MD5 d6c0ef61706ffcad9c70e9d68c0663a1
BLAKE2b-256 595c2f4422c650d85d91bf3ba557323de2aeab95d235dfe2f5496607191c0926

See more details on using hashes here.

File details

Details for the file hearth_engine-0.3.2-cp39-abi3-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for hearth_engine-0.3.2-cp39-abi3-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 5905899ef109e9aff1cfe6ab09de8165dabc367873d488ebcc06d4e48b9f9705
MD5 214888688327da20f65e0995083f48b9
BLAKE2b-256 dcd6be01e0e61c7c760b17b373047e2db6e54051cb589c5a4925bcd5ef0c0c1b

See more details on using hashes here.

Release history Release notifications | RSS feed

0.4.7

6 files

0.4.6

6 files

0.4.5

6 files

0.4.3

6 files

0.4.2

6 files

0.4.1

6 files

0.4.0

6 files

0.3.9

6 files

0.3.8

6 files

0.3.7

6 files

0.3.6

6 files

0.3.5

6 files

0.3.4

6 files

0.3.3

6 files

This release

0.3.2 This release

6 files

0.3.1

6 files

0.3.0

6 files

0.2.0

6 files

0.1.2

6 files

0.1.1

6 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page