Skip to main content

betaloop

A reusable, storage-free ReAct agent kernel: the engine, framework and protocols that drive a tool-calling agent, plus optional capability bundles. It knows nothing about how runs are stored (or even whether they are) — persistence is an optional EventSink a host plugs in. Any application (a thesis-writing platform, a coding agent, ...) implements its own tools + prompt + storage and reuses this kernel.

Core (zero I/O, zero business)

  • runtime — AgentRuntime ReAct loop + event stream. Supports cancellation (run(..., stop=Event|callable) → cancelled event, status="cancelled"; checked between streaming deltas too, closing the in-flight model stream instead of paying for a response nobody wants) and streaming (LLMConfig(stream=True) → assistant_delta events while the model generates; a final full assistant event always follows). Every tool_call event is announced before any of the step's tools execute, so a slow tool never hides what is pending on the frontend; result events still follow in model order. Every model call emits a usage event — prompt/completion/total tokens, that call's cost, and context fullness (context_tokens, context_chars, context_window, context_percent when LLMConfig(context_window=...) is set) — so a frontend can show live token/context gauges; RunStats and the host done event carry the cumulative breakdown. Run budgets (max_cost / max_total_tokens) cut a runaway run short with status="budget_exceeded" — no further tool execution, no further model calls — and a host may seed stats with prior-conversation totals to budget across runs. A stuck model reissuing the identical (tool, args) call more than repeat_call_limit (default 3) times gets an inline nudge inside that tool's result, feeding self-correction (None disables). Malformed tool-call arguments come back to the model as failed tool results instead of executing with empty/wrong args. A run that exhausts its step budget gets a forced toolless wrap-up call (tool_choice="none") so it ends with the model's summary, reporting status="max_steps" (and never executing the stubborn model's further tool calls). LLMConfig(temperature=..., max_tokens=...) are forwarded on every call, and a generation cut off by the token cap (finish_reason="length") is marked — the notice rides in the final text, the per-call usage event carries finish_reason, and a tool_call whose arguments JSON was truncated gets a "cut off by max_tokens, re-issue the call" error instead of a bare "invalid JSON" that invites a byte-identical retry. An empty response (no text, no tool calls) is retried once and otherwise ends the run with status="empty_response" instead of an empty "success". A missing tool_call id is synthesized at one point (streamed and non-streamed alike), so a gateway that omits ids never poisons the next request with tool_call_id: null — and streamed calls no longer mint colliding call_0-style ids across steps. Synthetic run endings (max-step / budget notices) are recorded like any assistant turn, so sinks and the next run's replay see why the run stopped. Tool calls execute in model order with consecutive READ tools parallel and every WRITE/META tool alone (no write races); ToolSpec(timeout=...) cancels a hung call. A mid-run context budget (context_budget, default 400k chars) shrinks old tool results head+tail so long runs don't blow the context window. Sinks are error-isolated (a broken display/record sink logs instead of killing the run; strict_records=True opts record failures back into fatal). AgentRuntime(envelope=True) emits the run_start/done envelope itself for hosts driving the runtime directly, and AgentRuntime(http_client=...) (forwarded by AgentHost) reuses one host-owned httpx.AsyncClient across runs — connection pooling, limits, proxy/verify config for service hosts.
  • llm — OpenAI-compatible chat client (retry / jittered backoff / response-shape validation / Retry-After-aware 429 handling / fatal-4xx fail-fast) + SSE streaming helpers; transports normalize usage to prompt_tokens/completion_tokens/total_tokens across chat-completions and Responses shapes
  • tools — ToolRegistry (register / unregister / dispatch / mode filtering / argument validation / per-tool timeout) + pre-dispatch middleware via add_middleware (audit / quota / human-in-the-loop confirmation of write tools)
  • actions — Action + UndoEngine (pure, storage-free undo; reverters may be sync or async)
  • memory — replay_messages / recap_text / window_with_recap / run_timeline (reconstruct a stored turn's ordered event timeline) + MemoryProvider; replay reconciliation is two-sided and window-safe (a tool row whose calling assistant fell outside the window is dropped, not sent as an orphan first message)
  • events — EventSink (display + record channels) + SSE serialization (to_sse degrades non-JSON values via str() — a datetime inside a tool's ui payload can't crash the host's SSE layer)
  • context / modes — AgentContext + AgentMode / ToolCategory; host-defined modes via register_mode(name, categories) (unknown modes raise instead of silently degrading to read-only). AgentContext.shared is per-run state shared by reference with subagent contexts — the vehicle for cross-context coordination (the workspace stale-file guard's revision map, the run's cancellation handle)

Optional bundles (betaloop.bundles)

  • host — AgentHost host-adapter framework: message assembly, run envelope (run_start/done), error funneling, StoreSink (persist via a store), undo_run (zero-config: bundled tools register reverters keyed by their action kinds, and reverters see the host's extra context; a reversion whose status mark fails surfaces in the report instead of being swallowed), DictToolAdapter (wrap a dict-based tool system — specs without a handler log a warning instead of silently vanishing, and a reverters= map wires custom action kinds into undo). Eliminates the per-host boilerplate round 1 left behind. host.run(..., stop=...) forwards cancellation to the runtime; AgentHost(max_cost=..., max_total_tokens=..., repeat_call_limit=...) forwards the runtime's budget / repeat guards to every run it builds; AgentHost(http_client=...) shares one HTTP client across runs, and AgentHost(capture_actions=False) turns off StoreSink's automatic action capture for hosts that persist actions themselves.

  • store — RunStore/ConversationStore/BlobStore Protocols + JsonlRunStore (default, zero-database JSONL + content-addressed blob spillover). Hosts wanting a DB implement the Protocols; the default needs none. Action values larger than spill_threshold (default 8KB) externalize to blobs and rehydrate transparently on read; id counters are in-memory so appends don't rescan the stream. list_actions(subagent=...) filters by subagent during the fold (unselected rows skip blob rehydration — the subagent engine's snapshots stay O(one worker) on long runs); torn lines log a warning and count in dropped_lines instead of vanishing silently; fsync=True flushes each append for crash-durability.

  • subagents — SubagentEngine + SubagentRoster/SubagentSpec + delegate tool: isolated worker agents the orchestrator hands subtasks to, tagged so undo still reverts them while the orchestrator's context stays lean. delegate_parallel fans independent tasks out concurrently (bounded by max_parallel, failures isolated per agent, duplicate agents in one batch rejected — they would claim each other's actions); SubagentEngine(on_subagent_event=...) streams live subagent_progress heartbeats to a host push channel (text, args and summaries capped — a 50KB write_file payload never rides the callback); SubagentSpec(transport=...) routes a subagent to a different endpoint. make_delegate_tool(engine, timeout=...) caps one delegation's wall time so a hung worker cannot hold the orchestrator's step forever. Budget caps apply per subagent run (each delegation gets its own max_cost / max_total_tokens) while each delegation's spend folds into the parent run's done event, stats and store row (with a subagent_* breakdown); cancellation of the orchestrating run propagates into an in-flight subagent (its stop handle rides in ctx.shared); a delegation's own mutations are identified by an id-membership snapshot, so stores with opaque (non-integer) action ids count them correctly.

  • admin — tool_categories / list_tools_admin / list_tool_packages_admin / check_packages over a registry + display packages (admin-panel source; packages are grouping only, tools stay per-name togglable).

  • patch — apply_patch: line-oriented multi-file edits via the Codex *** Begin Patch envelope (add/update/move/delete files, @@ chunks of context/-/+ lines, *** End of File anchoring). Chunk location runs a four-pass fuzzy ladder (exact → trailing-ws → strip → Unicode-punctuation fold); *** End of File chunks run the tail-anchored ladder at full strength before any forward match, so a whitespace-mismatched tail beats an exact look-alike earlier in the file instead of silently editing the wrong site; two same-position insertions keep document order. Application is all-or-nothing against an in-memory overlay (later sections of the same file chain; create-then-edit works), so a bad hunk leaves the workspace untouched. Same file_change events + one new undo kind (file_delete).

  • workspace — sandboxed file I/O + read/write/edit/list/search/glob tools + file undo reverters. Directory walks (list_files / search_files / glob_files) never follow symlinks — code executed by run_code could otherwise plant a link to a host file and read it back through a walk, bypassing the path guard (symlinks list as an opaque symlink type; list_files(dirs=...) routes the requested folders through the same path guard). read_file prefixes every line with its 1-based number and supports offset/limit line-window reads. edit_file refuses ambiguous old_text (multi-match) unless replace_all is set, returns a diff, and falls back to a whole-line fuzzy match (shared ladder, trailing-blank aligned like apply_patch) when the exact substring misses. All write tools guard against stale content: a file read this run and changed out-of-band is refused with "re-read it" instead of being clobbered; the revision map lives in ctx.shared, so the guard spans the orchestrator and every (including parallel) subagent of the run. search_files greps content by regex (dir / glob filters) with a 30s tool timeout, a 10s scan budget and a 10k-char line cap (catastrophic-backtracking patterns and huge minified lines can't hang the loop); glob_files matches paths by pattern.

  • todos — a per-scope task list the agent plans against: TodoStore (pure container) + JsonTodoStore (atomic tmp+rename JSON persistence; malformed items are dropped on load instead of crashing the system prompt) + update_todos/list_todos tools (todo_replace reverter) + todos_block for splicing the list into a system prompt.

  • images — tool-tier image perception, no kernel changes. image_info: stdlib-only header probe (PNG/JPEG/GIF/BMP/WEBP — dimensions, dpi, color mode) answering deterministic questions with zero model calls; it stats first and reads only a bounded header window off the event loop. analyze_image: ONE vision-model call (OpenAI image_url data-URL block + the question), registered only when a LLMConfig is passed — the host's main config reuses the main model, a dedicated one routes vision elsewhere. Size is checked by stat before reading (a 2GB upload is refused, not loaded), file reads run off the event loop, and detail="auto" shares a cache key with an omitted detail (the API treats them identically — no double billing). Images never enter the main conversation (answers are memoized per file hash + question), so context budget / trimming / replay stay untouched.

  • sandbox — Python code execution (bubblewrap or passthrough backend); output truncation keeps head+tail so tracebacks at the end stay visible. run_code/run_file are WRITE-classified (executing model-written code can mutate the workspace): invisible in read-only modes and never run in parallel with other tool calls. Their side effects produce no undo records (the tool descriptions say so) — durable edits belong in write_file/edit_file/apply_patch. Note the contract difference: register_code_tools takes workspace_for(ctx) -> root path (a str), while the workspace/images bundles take workspace_for(ctx) -> Workspace.

  • skills — markdown skill libraries (SkillLibrary flat dir; package-aware SkillPackages with RemoteSkillSource registry mirrors — refresh failures clean their staging dir, log, and keep the previous cache) + load_skill tool. Skills do file/network I/O, hence a bundle — import betaloop stays zero-I/O (deprecated betaloop.skills alias kept; it warns on import).

  • mcp — MCPManager + MCPServerConfig + parse_servers: bridge external MCP servers into the registry — stdio transport (command=[...], e.g. npx -y @z_ai/mcp-server) or streamable-http (url= + headers= for auth, Mcp-Session-Id handled automatically). tools/list pagination (nextCursor) is followed, so paginated servers don't silently lose half their tools. Spawned servers get only a safe env allow-list plus the configured env (PATH/locale/ HOME/TMPDIR) — never all of os.environ with its secrets — unless inherit_env=True opts back into the legacy behavior. URLs log as scheme://host only (credentials in the query or userinfo never reach a log line). Sessions outlive registry rebuilds and lazily self-heal: a dead session is restarted on the next tool call (revive), and attach/ensure re-register tools on any fresh registry (ensure is the idempotent variant for hosts that cache registries across runs). tool_allowlist/tool_blocklist filter by the server-side tool name. readOnlyHint annotations map to the READ category; MCP tools carry no undo reverters. Stdlib-only (newline-delimited JSON-RPC / plain POST + SSE), secrets stay host-side:

    from betaloop.bundles import MCPManager, MCPServerConfig
    
    manager = MCPManager([MCPServerConfig(
        name="zai", command=["npx", "-y", "@z_ai/mcp-server"],
        env={"Z_AI_API_KEY": "...", "Z_AI_MODE": "ZHIPU"},
        default_category="read")])
    await manager.attach(registry)   # registry gains zai__* tools
    

Install

pip install betaloop
# local dev (editable + test/lint deps):
pip install -e ".[dev]"

Storage model

The kernel stores nothing. A host provides:

  • an EventSink (write side) — persists message records however it likes (DB / file / nowhere);
  • a MemoryProvider (read side, optional) — replays prior turns;
  • an UndoEngine fed from wherever the host kept actions.

Minimal host sketch

from betaloop import AgentContext, LLMConfig, ToolRegistry
from betaloop.bundles import AgentHost, JsonlRunStore
from betaloop.bundles.workspace import Workspace, register_file_tools

registry = ToolRegistry()
# workspace_for(ctx) -> Workspace (a sandboxed root per user):
register_file_tools(registry, lambda ctx: Workspace(f"/data/{ctx.user_id}"))
store = JsonlRunStore("/var/lib/myapp/agent")                       # zero-DB default
host = AgentHost(registry,
                 LLMConfig(model=..., base_url=..., api_key=...),
                 store, build_system_prompt=my_prompt_builder)

ctx = AgentContext(run_id=rid, user_id=uid)
async for event in host.run(ctx, task, history=prior_turns):
    ...  # forward run_start / step / tool_call / tool_result / done to your frontend

A host supplies only its tools, system prompt, and (optionally) a store backend — the engine, persistence, run envelope, undo, and (via subagents) delegation are all reused.

More

  • examples/minimal_host.py — a runnable, offline minimal host (tools → run → events → undo); runs in CI.
  • examples/live_host.py — the same flow against a real OpenAI-compatible endpoint (set BETA_API_KEY / BETA_BASE_URL / BETA_MODEL; no default endpoint, it never spends tokens by accident).
  • CHANGELOG.md — what changed and when.

License

MIT

Release files for betaloop 0.8.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for betaloop 0.8.0
File Size Uploaded
betaloop-0.8.0.tar.gz 189.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for betaloop 0.8.0
File Interpreter ABI Platform
betaloop-0.8.0-py3-none-any.whl Python 3 none any Details

Total release size: 319.1 kB

Release files / betaloop-0.8.0.tar.gz

Download URL betaloop-0.8.0.tar.gz
Size 189.7 kB
Tags Source
SHA-256 checksum
How to use checksums
7869822a7593adcc21218f4a472c4e1b1d6c70ae078dcb7a7c55cafed7e5e7c5
BLAKE2b-256 checksum
How to use checksums
ebae453d2334d5eb0e2ef99c57176e102521a8403762ef2d4706ca32f259492f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / betaloop-0.8.0-py3-none-any.whl

Download URL betaloop-0.8.0-py3-none-any.whl
Size 129.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
07bf872c18371c835fde42f5cf0ffac100b83416521a0449778537222365134d
BLAKE2b-256 checksum
How to use checksums
53a41fe519b94e6da1fbab3bc327b9a13612e476c1bbf24a81b62a892e83842c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

0.9.4

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

This release

0.8.0 This release

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page