llm-preflight
Fail-safe client discipline for local LLMs on memory-constrained machines.
Small local models fail unattended and long-session work in predictable, preventable ways: OOM kills mid-job, swap-thrash that turns a 20-second call into a 5-minute timeout, hidden reasoning tokens silently eating your budget, crashed sessions losing everything. This library is the client-side layer that catches each failure at the right seam — with any OpenAI-compatible server (omlx, llama.cpp, LM Studio, Ollama, FreeToken, vLLM).
The model stays small. The system stops being fragile.
Why this exists
The pain is real and documented in the wild — e.g. openclaw#65551: "Local MLX/LM Studio models get terminated on RAM pressure + no graceful handling" in cron jobs. Local LLM reliability isn't a model problem; it's an operational problem, and nobody ships the operational layer.
Everything in this library was built to run real production cron jobs on a base M4 Mac Mini (24 GB), then generalized. The measured numbers behind every default are in docs/MEASURED.md.
The five protections
| # | Protection | What it prevents |
|---|---|---|
| 1 | Preflight memory check — measures the pool your model actually allocates from (unified RAM on Apple Silicon; VRAM + RAM independently on discrete-GPU systems, gating on whichever is tighter) | Swap-thrash: the silent failure where a starved call burns minutes then times out, and a naive retry burns them again |
| 2 | Starvation-aware retry — attempts that die after slow_death_s are never retried |
Doubling the cost of a doomed call |
| 3 | Thinking-mode-off — chat_template_kwargs: {"enable_thinking": false} by default |
The hidden-reasoning tax: 4.7× more tokens, 4.1× slower on tasks that don't need it (measured; prompt-level /no_think begging does not work) |
| 4 | Token budgets — hard input cap, bounded output | 60k-token prompts into a server that takes minutes to prefill |
| 5 | Typed failures — MemoryPressureError (defer, don't retry) vs LocalModelUnavailable (alert/fallback) |
Silent garbage delivery |
Quick start
pip install local-llm-preflight # stdlib-only core; zero dependencies
from llm_preflight import PreflightClient, ClientConfig, MemoryPressureError
client = PreflightClient(ClientConfig(
base_url="http://127.0.0.1:8000/v1",
model="your-served-model-id", # from: curl $BASE/v1/models
))
try:
text, usage = client.chat(
system="You are a concise summarizer. Output only what is asked.",
user=long_input_text,
)
except MemoryPressureError:
defer_or_fallback() # RAM/VRAM starved — do NOT retry now
Or configure per-hardware with a TOML file:
from llm_preflight.config import load_config
client = PreflightClient(load_config()) # reads ./llm-preflight.toml
# llm-preflight.toml
[server]
base_url = "http://127.0.0.1:8000/v1"
model = "your-served-model-id" # exact id from your server's /v1/models
[memory]
min_system_mb = 2500
min_vram_mb = 1500 # gates independently on NVIDIA GPUs (pynvml, optional)
cold_system_mb = 7000 # extra headroom when the model isn't loaded yet
[retry]
slow_death_s = 90.0 # raise this for interactive long generations
Long interactive sessions (experimental)
The distinct problem: preflight catches "don't start while starved", not "started fine, ran out of memory 40 turns in." Long sessions die from context growth — exactly what happens running 35B-class MoE models on gaming GPUs. The experimental session manager bounds it:
from llm_preflight.session import Session, SessionConfig
sess = Session(client_config, SessionConfig(
compact_at_frac=0.75, # compact at 75% of server max context
strategy="truncate", # anchors + sliding window (see docs for why
# summarize-and-replace is opt-in, not default)
))
sess.seed("You are a coding assistant.")
reply, usage = sess.send("Explain this error: ...")
# Crash? Resume from the last turn-boundary checkpoint:
sess2 = Session.resume(client_config, session_id=sess.sid, session_config=SessionConfig())
Checkpoints are the portable seam: we persist the message list at turn boundaries and let the server rebuild its own KV cache on resume. We never touch engine-internal KV state — there is no portable API for it, and pretending otherwise couples you to one engine.
⚠️ The session API is experimental (v0.2 preview) and may change. The single-shot client is the stable, production-proven core.
Check your server first
"OpenAI-compatible" servers vary in what they actually implement (usage fields, thinking-mode kwargs, real context caps). The probe answers the three questions our design depends on:
llm-preflight-probe [base_url] [model]
# or: python -m llm_preflight.probe http://127.0.0.1:8000/v1
What this library does NOT fix
Honest limits (see docs/MEASURED.md for the full accounting): this fixes availability and format failures — crashes, timeouts, malformed output, memory exhaustion. It does nothing for correctness failures. A 9B/4-bit model will still confidently hallucinate facts inside perfectly valid JSON. Validation catches syntax, not truth. Route high-stakes calls to bigger models or the cloud, and keep the script owning the facts while the model owns the formatting.
Development
git clone https://github.com/manulaggarwal/llm-preflight
cd llm-preflight
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest tests/ -q # unit suite (no server needed)
.venv/bin/python tests/integration_live.py # against your running server
The unit suite monkeypatches platform memory readers — it passes identically on macOS, Linux, and Windows CI without depending on the host's RAM state.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file local_llm_preflight-0.1.1.tar.gz.
File metadata
- Download URL: local_llm_preflight-0.1.1.tar.gz
- Upload date:
- Size: 27.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0897feb5075732fe97a94754f3da13775cff4bfcfd7fea552522d2659b33f824
|
|
| MD5 |
5ce45c95d9154a90e4423b23af1bc6dc
|
|
| BLAKE2b-256 |
1e699bccfd1853f9f809d569b2ef648ff5aed570c35010fd7611f452ff75422c
|
Provenance
The following attestation bundles were made for local_llm_preflight-0.1.1.tar.gz:
Publisher:
publish.yml on manulaggarwal/llm-preflight
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
local_llm_preflight-0.1.1.tar.gz -
Subject digest:
0897feb5075732fe97a94754f3da13775cff4bfcfd7fea552522d2659b33f824 - Sigstore transparency entry: 2583507943
- Sigstore integration time:
-
Permalink:
manulaggarwal/llm-preflight@38a571afb7a2d863a35a652409c5eb722d224425 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/manulaggarwal
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@38a571afb7a2d863a35a652409c5eb722d224425 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file local_llm_preflight-0.1.1-py3-none-any.whl.
File metadata
- Download URL: local_llm_preflight-0.1.1-py3-none-any.whl
- Upload date:
- Size: 20.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3927fe966a600518eade205d7904dbaa0db02a9f33f7a0a8c4d313dfbc5ea8ed
|
|
| MD5 |
742bdc4e8b03332e40e3cb77eefaee70
|
|
| BLAKE2b-256 |
a8cef6a8dc110cf84c5aea7398e727a448880ae891f86de5f8328f5567a2477e
|
Provenance
The following attestation bundles were made for local_llm_preflight-0.1.1-py3-none-any.whl:
Publisher:
publish.yml on manulaggarwal/llm-preflight
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
local_llm_preflight-0.1.1-py3-none-any.whl -
Subject digest:
3927fe966a600518eade205d7904dbaa0db02a9f33f7a0a8c4d313dfbc5ea8ed - Sigstore transparency entry: 2583507955
- Sigstore integration time:
-
Permalink:
manulaggarwal/llm-preflight@38a571afb7a2d863a35a652409c5eb722d224425 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/manulaggarwal
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@38a571afb7a2d863a35a652409c5eb722d224425 -
Trigger Event:
workflow_dispatch
-
Statement type: