lithe
A reusable, storage-free ReAct agent kernel: the engine, framework and
protocols that drive a tool-calling agent, plus optional capability bundles. It
knows nothing about how runs are stored (or even whether they are) — persistence
is an optional EventSink a host plugs in. Any application (a thesis-writing
platform, a coding agent, ...) implements its own tools + prompt + storage and
reuses this kernel.
Core (zero I/O, zero business)
runtime—AgentRuntimeReAct loop + event stream. Supports cancellation (run(..., stop=Event|callable)→cancelledevent,status="cancelled"; checked between streaming deltas too, closing the in-flight model stream instead of paying for a response nobody wants) and streaming (LLMConfig(stream=True)→assistant_deltaevents while the model generates; a final fullassistantevent always follows). Everytool_callevent is announced before any of the step's tools execute, so a slow tool never hides what is pending on the frontend; result events still follow in model order. Every model call emits ausageevent — prompt/completion/total tokens, that call's cost, and context fullness (context_tokens,context_chars,context_window,context_percentwhenLLMConfig(context_window=...)is set) — so a frontend can show live token/context gauges;RunStatsand the hostdoneevent carry the cumulative breakdown. Run budgets (max_cost/max_total_tokens) cut a runaway run short withstatus="budget_exceeded"— no further tool execution, no further model calls — and a host may seedstatswith prior-conversation totals to budget across runs. A stuck model reissuing the identical (tool, args) call more thanrepeat_call_limit(default 3) times gets an inline nudge inside that tool's result, feeding self-correction (Nonedisables). Malformed tool-call arguments come back to the model as failed tool results instead of executing with empty/wrong args. A run that exhausts its step budget gets a forced toolless wrap-up call (tool_choice="none") so it ends with the model's summary, reportingstatus="max_steps"(and never executing the stubborn model's further tool calls).LLMConfig(temperature=..., max_tokens=...)are forwarded on every call, and a generation cut off by the token cap (finish_reason="length") is marked — the notice rides in the final text, the per-callusageevent carriesfinish_reason, and a tool_call whose arguments JSON was truncated gets a "cut off by max_tokens, re-issue the call" error instead of a bare "invalid JSON" that invites a byte-identical retry. An empty response (no text, no tool calls) is retried once and otherwise ends the run withstatus="empty_response"instead of an empty "success". A missingtool_callid is synthesized at one point (streamed and non-streamed alike), so a gateway that omits ids never poisons the next request withtool_call_id: null— and streamed calls no longer mint collidingcall_0-style ids across steps. Synthetic run endings (max-step / budget notices) are recorded like any assistant turn, so sinks and the next run's replay see why the run stopped. Tool calls execute in model order with consecutive READ tools parallel and every WRITE/META tool alone (no write races);ToolSpec(timeout=...)cancels a hung call. A mid-run context budget (context_budget, default 400k chars) shrinks old tool results head+tail so long runs don't blow the context window. Sinks are error-isolated (a broken display/record sink logs instead of killing the run;strict_records=Trueopts record failures back into fatal).AgentRuntime(envelope=True)emits therun_start/doneenvelope itself for hosts driving the runtime directly, andAgentRuntime(http_client=...)(forwarded byAgentHost) reuses one host-ownedhttpx.AsyncClientacross runs — connection pooling, limits, proxy/verify config for service hosts.llm— OpenAI-compatible chat client (retry / jittered backoff / response-shape validation /Retry-After-aware 429 handling / fatal-4xx fail-fast) + SSE streaming helpers; transports normalize usage toprompt_tokens/completion_tokens/total_tokensacross chat-completions and Responses shapestools—ToolRegistry(register / unregister / dispatch / mode filtering / argument validation / per-tool timeout) + pre-dispatch middleware viaadd_middleware(audit / quota / human-in-the-loop confirmation of write tools)actions—Action+UndoEngine(pure, storage-free undo; reverters may be sync or async)memory—replay_messages/recap_text/window_with_recap/run_timeline(reconstruct a stored turn's ordered event timeline) +MemoryProvider; replay reconciliation is two-sided and window-safe (a tool row whose calling assistant fell outside the window is dropped, not sent as an orphan first message)events—EventSink(display + record channels) + SSE serialization (to_ssedegrades non-JSON values viastr()— adatetimeinside a tool'suipayload can't crash the host's SSE layer)context/modes—AgentContext+AgentMode/ToolCategory; host-defined modes viaregister_mode(name, categories)(unknown modes raise instead of silently degrading to read-only).AgentContext.sharedis per-run state shared by reference with subagent contexts — the vehicle for cross-context coordination (the workspace stale-file guard's revision map, the run's cancellation handle)
Optional bundles (lithe.bundles)
-
host—AgentHosthost-adapter framework: message assembly, run envelope (run_start/done), error funneling,StoreSink(persist via a store),undo_run(zero-config: bundled tools register reverters keyed by their action kinds, and reverters see the host'sextracontext; a reversion whose status mark fails surfaces in the report instead of being swallowed),DictToolAdapter(wrap a dict-based tool system — specs without a handler log a warning instead of silently vanishing, and areverters=map wires custom action kinds into undo). Eliminates the per-host boilerplate round 1 left behind.host.run(..., stop=...)forwards cancellation to the runtime;AgentHost(max_cost=..., max_total_tokens=..., repeat_call_limit=...)forwards the runtime's budget / repeat guards to every run it builds;AgentHost(http_client=...)shares one HTTP client across runs, andAgentHost(capture_actions=False)turns offStoreSink's automatic action capture for hosts that persist actions themselves. -
store—RunStore/ConversationStore/BlobStoreProtocols +JsonlRunStore(default, zero-database JSONL + content-addressed blob spillover). Hosts wanting a DB implement the Protocols; the default needs none. Action values larger thanspill_threshold(default 8KB) externalize to blobs and rehydrate transparently on read; id counters are in-memory so appends don't rescan the stream.list_actions(subagent=...)filters by subagent during the fold (unselected rows skip blob rehydration — the subagent engine's snapshots stay O(one worker) on long runs); torn lines log a warning and count indropped_linesinstead of vanishing silently;fsync=Trueflushes each append for crash-durability. -
subagents—SubagentEngine+SubagentRoster/SubagentSpec+delegatetool: isolated worker agents the orchestrator hands subtasks to, tagged so undo still reverts them while the orchestrator's context stays lean.delegate_parallelfans independent tasks out concurrently (bounded bymax_parallel, failures isolated per agent, duplicate agents in one batch rejected — they would claim each other's actions);SubagentEngine(on_subagent_event=...)streams livesubagent_progressheartbeats to a host push channel (text, args and summaries capped — a 50KBwrite_filepayload never rides the callback);SubagentSpec(transport=...)routes a subagent to a different endpoint.make_delegate_tool(engine, timeout=...)caps one delegation's wall time so a hung worker cannot hold the orchestrator's step forever. Budget caps apply per subagent run (each delegation gets its ownmax_cost/max_total_tokens) while each delegation's spend folds into the parent run'sdoneevent, stats and store row (with asubagent_*breakdown); cancellation of the orchestrating run propagates into an in-flight subagent (its stop handle rides inctx.shared); a delegation's own mutations are identified by an id-membership snapshot, so stores with opaque (non-integer) action ids count them correctly. -
admin—tool_categories/list_tools_admin/list_tool_packages_admin/check_packagesover a registry + display packages (admin-panel source; packages are grouping only, tools stay per-name togglable). -
patch—apply_patch: line-oriented multi-file edits via the Codex*** Begin Patchenvelope (add/update/move/delete files,@@chunks of context/-/+lines,*** End of Fileanchoring). Chunk location runs a four-pass fuzzy ladder (exact → trailing-ws → strip → Unicode-punctuation fold);*** End of Filechunks run the tail-anchored ladder at full strength before any forward match, so a whitespace-mismatched tail beats an exact look-alike earlier in the file instead of silently editing the wrong site; two same-position insertions keep document order. Application is all-or-nothing against an in-memory overlay (later sections of the same file chain; create-then-edit works), so a bad hunk leaves the workspace untouched. Samefile_changeevents + one new undo kind (file_delete). -
workspace— sandboxed file I/O + read/write/edit/list/search/glob tools + file undo reverters. Directory walks (list_files/search_files/glob_files) never follow symlinks — code executed byrun_codecould otherwise plant a link to a host file and read it back through a walk, bypassing the path guard (symlinks list as an opaquesymlinktype;list_files(dirs=...)routes the requested folders through the same path guard).read_fileprefixes every line with its 1-based number and supportsoffset/limitline-window reads.edit_filerefuses ambiguousold_text(multi-match) unlessreplace_allis set, returns a diff, and falls back to a whole-line fuzzy match (shared ladder, trailing-blank aligned likeapply_patch) when the exact substring misses. All write tools guard against stale content: a file read this run and changed out-of-band is refused with "re-read it" instead of being clobbered; the revision map lives inctx.shared, so the guard spans the orchestrator and every (including parallel) subagent of the run.search_filesgreps content by regex (dir / glob filters) with a 30s tool timeout, a 10s scan budget and a 10k-char line cap (catastrophic-backtracking patterns and huge minified lines can't hang the loop);glob_filesmatches paths by pattern. -
todos— a per-scope task list the agent plans against:TodoStore(pure container) +JsonTodoStore(atomic tmp+rename JSON persistence; malformed items are dropped on load instead of crashing the system prompt) +update_todos/list_todostools (todo_replacereverter) +todos_blockfor splicing the list into a system prompt. -
images— tool-tier image perception, no kernel changes.image_info: stdlib-only header probe (PNG/JPEG/GIF/BMP/WEBP — dimensions, dpi, color mode) answering deterministic questions with zero model calls; it stats first and reads only a bounded header window off the event loop.analyze_image: ONE vision-model call (OpenAIimage_urldata-URL block + the question), registered only when aLLMConfigis passed — the host's main config reuses the main model, a dedicated one routes vision elsewhere. Size is checked by stat before reading (a 2GB upload is refused, not loaded), file reads run off the event loop, anddetail="auto"shares a cache key with an omitted detail (the API treats them identically — no double billing). Images never enter the main conversation (answers are memoized per file hash + question), so context budget / trimming / replay stay untouched. -
sandbox— Python code execution (bubblewrap or passthrough backend); output truncation keeps head+tail so tracebacks at the end stay visible.run_code/run_fileare WRITE-classified (executing model-written code can mutate the workspace): invisible in read-only modes and never run in parallel with other tool calls. Their side effects produce no undo records (the tool descriptions say so) — durable edits belong inwrite_file/edit_file/apply_patch. Note the contract difference:register_code_toolstakesworkspace_for(ctx) -> root path(astr), while the workspace/images bundles takeworkspace_for(ctx) -> Workspace. -
download—download_file(url, path?): stream one HTTP(S) resource into the workspace (defaultdownloads/), the controlled ingress the network-isolated sandbox deliberately lacks — fetching is a bounded, audited tool while executing model code stays offline. SSRF guard: http/ https only, redirects followed manually with every hop re-validated, and all resolved addresses of every hop must be globally routable (ipaddress.is_globalrejects loopback/private/link-local/CGN/reserved/ multicast — cloud metadata included); checking all A/AAAA records up front narrows DNS rebinding to the TTL window. Size cap byContent-Lengthpre-check plus streaming cutoff (a lying header gets cut mid-stream and the partial.partfile removed; the file lands atomically via rename). Existing targets are refused (the model picks a new name), so nothing undoable is mutated. Defaults: 64MB cap, 120s total budget (kernelToolSpectimeout as backstop), 5 redirects; WRITE-classified like the other workspace-mutating tools. -
skills— markdown skill libraries (SkillLibraryflat dir; package-awareSkillPackageswithRemoteSkillSourceregistry mirrors — refresh failures clean their staging dir, log, and keep the previous cache) +load_skilltool. Skills do file/network I/O, hence a bundle —import lithestays zero-I/O (deprecatedlithe.skillsalias kept; it warns on import). -
mcp—MCPManager+MCPServerConfig+parse_servers: bridge external MCP servers into the registry — stdio transport (command=[...], e.g.npx -y @z_ai/mcp-server) or streamable-http (url=+headers=for auth,Mcp-Session-Idhandled automatically).tools/listpagination (nextCursor) is followed, so paginated servers don't silently lose half their tools. Spawned servers get only a safe env allow-list plus the configuredenv(PATH/locale/ HOME/TMPDIR) — never all ofos.environwith its secrets — unlessinherit_env=Trueopts back into the legacy behavior. URLs log as scheme://host only (credentials in the query or userinfo never reach a log line). Sessions outlive registry rebuilds and lazily self-heal: a dead session is restarted on the next tool call (revive), andattach/ensurere-register tools on any fresh registry (ensureis the idempotent variant for hosts that cache registries across runs).tool_allowlist/tool_blocklistfilter by the server-side tool name.readOnlyHintannotations map to the READ category; MCP tools carry no undo reverters. Stdlib-only (newline-delimited JSON-RPC / plain POST + SSE), secrets stay host-side:from lithe.bundles import MCPManager, MCPServerConfig manager = MCPManager([MCPServerConfig( name="zai", command=["npx", "-y", "@z_ai/mcp-server"], env={"Z_AI_API_KEY": "...", "Z_AI_MODE": "ZHIPU"}, default_category="read")]) await manager.attach(registry) # registry gains zai__* tools
Install
pip install lithe
# local dev (editable + test/lint deps):
pip install -e ".[dev]"
Storage model
The kernel stores nothing. A host provides:
- an
EventSink(write side) — persists message records however it likes (DB / file / nowhere); - a
MemoryProvider(read side, optional) — replays prior turns; - an
UndoEnginefed from wherever the host kept actions.
Minimal host sketch
from lithe import AgentContext, LLMConfig, ToolRegistry
from lithe.bundles import AgentHost, JsonlRunStore
from lithe.bundles.workspace import Workspace, register_file_tools
registry = ToolRegistry()
# workspace_for(ctx) -> Workspace (a sandboxed root per user):
register_file_tools(registry, lambda ctx: Workspace(f"/data/{ctx.user_id}"))
store = JsonlRunStore("/var/lib/myapp/agent") # zero-DB default
host = AgentHost(registry,
LLMConfig(model=..., base_url=..., api_key=...),
store, build_system_prompt=my_prompt_builder)
ctx = AgentContext(run_id=rid, user_id=uid)
async for event in host.run(ctx, task, history=prior_turns):
... # forward run_start / step / tool_call / tool_result / done to your frontend
A host supplies only its tools, system prompt, and (optionally) a store
backend — the engine, persistence, run envelope, undo, and (via subagents)
delegation are all reused.
More
examples/minimal_host.py— a runnable, offline minimal host (tools → run → events → undo); runs in CI.examples/live_host.py— the same flow against a real OpenAI-compatible endpoint (setBETA_API_KEY/BETA_BASE_URL/BETA_MODEL; no default endpoint, it never spends tokens by accident).CHANGELOG.md— what changed and when.
License
MIT
Release files for lithe 0.9.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| lithe-0.9.4.tar.gz | 205.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| lithe-0.9.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 345.6 kB
Release files / lithe-0.9.4.tar.gz
| Download URL | lithe-0.9.4.tar.gz |
|---|---|
| Size | 205.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8243d0cb85a52d8e3fe58ea45b1ef7e05b44c8fb3a5e3ba125cf614436920291
|
|
BLAKE2b-256 checksum How to use checksums |
660e30326cfcc057615e10c2af6949267ee29e7334603df7ba4c95caff614d03
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / lithe-0.9.4-py3-none-any.whl
| Download URL | lithe-0.9.4-py3-none-any.whl |
|---|---|
| Size | 139.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0ca6208a50c9e2e4a1a3e55a442c0aeaabe155820a2aae81d758831cb27a0bb0
|
|
BLAKE2b-256 checksum How to use checksums |
c3bbd9bfd9941f9b635ed99131314c9453fb1235febe835c00aa55547f1fe324
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|