Skip to main content

abstract_toolserver

The abstract_* ecosystem exposed as an API-callable AI toolset — every tool is a plain Python function turned into a self-describing HTTP endpoint by abstract_flask. A portable tool layer any model (Claude, hugpy, …) can drive over HTTP instead of being bound to one runtime's tool harness.

Part of the hugpy orbit

                    ┌──────────────────────────── hugpy (fleet) ───────────────────────────┐
                    │ central + workers: platform · engine · fleet · server · media · …     │
                    │ OpenAI-compatible /v1 — every local model, incl. B (Qwen3-Coder-Next) │
                    └───────▲───────────────────────▲──────────────────────────▲───────────┘
                            │ inference             │ inference                │ B reductions
   ┌────────────────────────┴──┐   ┌────────────────┴──────────┐   ┌───────────┴───────────────┐
   │ hugpy-station             │   │ hugpy-agent               │   │ abstract-toolserver       │
   │ desktop + headless console│──▶│ agent runtime · TUI ·     │◀─▶│ comms · ledgers · boards ·│
   │ tmux seats per locus      │   │ OpenCode/qwen seats       │   │ exchanges · MCP · b_ask   │
   └────────────┬──────────────┘   └────────────┬──────────────┘   └───────────▲───────────────┘
                │ keeper/codex seats            │ --serve                      │ tools (MCP/HTTP)
   ┌────────────▼──────────────┐   ┌────────────▼──────────────┐               │
   │ abstract-gpt (Codex seat) │   │ abstract-claude serve ────┼───────────────┘
   │ abstract-claude (Claude)  │   │  └ abstract-serve-core    │
   └───────────────────────────┘   └───────────────────────────┘
          everything ships through abstract-pypit → PyPI (+ GitHub)
Package Role PyPI
hugpy (14 lockstep dists) the self-hosted LLM fleet: central, workers, engine, media, server hugpy
hugpy-station Electron desktop + headless backend; tmux seats, prompt composer, loop/bug scan deb via central install links
hugpy-agent agent runtime on the fleet; hugpy-agent tui over abstract-claude serve hugpy-agent
abstract-claude Claude Code launch/session/rollover + abstract-claude serve (roles keeper/chat/worker/local) abstract-claude
abstract-serve-core the HTTP routes abstract-claude serve actually runs (queue, relay, rollover sweeps) abstract-serve-core
abstract-gpt Codex/ChatGPT seat counterpart of abstract-claude abstract-gpt
abstract-toolserver one tool service per host: comms, ledgers, boards, exchanges, MCP bridge, B on call abstract-toolserver
abstract-pypit one-command publisher: bump → build → PyPI → GitHub push abstract-pypit

Run it

pip install abstract_toolserver          # + the extras you want to expose
python -m abstract_toolserver            # HOST/PORT/DEBUG from TOOLSERVER_* env
from abstract_toolserver import get_toolserver_app
app = get_toolserver_app()               # a normal Flask/WSGI app

Self-describing surface

The app auto-mounts introspection endpoints (from abstract_flask):

Endpoint What it gives an LLM
GET /prefixes the tool categories (/fs, /db, /ui, …)
GET /endpoints every tool as {endpoint, url, methods}
GET /<cat>/<tool>?help=true that tool's signature/help

Call a tool with JSON; unknown keys are pruned to the function signature, and the reply is {"result": ...} (or {"error": ...}). Discover-and-dispatch from a client is already provided by abstract_apis.make_endpoint_call.

curl -s localhost:5000/fs/count_tokens -d '{"text":"hello world"}'
# {"result": 2}
curl -s localhost:5000/db/schema           # {"result": {table: [cols...]}}
curl -s 'localhost:5000/db/query?help=true'

Tool categories

Prefix Tools Backend
/fs search, read_span, extract, read_file, write_file, read_json, find_keys, find_paths, glob, imports, find_content abstract_search, abstract_utilities, abstract_paths
/text count_tokens, chunk, detect_language abstract_utilities
/web text, links, attributes abstract_webtools
/media ocr_image, pdf_text, summarize, keywords, transcribe abstract_ocr, abstract_pandas, media_intelligence
/ai query abstract_ai
/db tables, schema, columns, fetch, query abstract_database
/sys run_cmd stdlib (gated)
/ui capture, monitors, ocr, windows, click_verify abstract_clicks, abstract_windows
/browser surfaces, lab, shot, look, read, locate, click, move, drag, scroll, type, key, go, wait, console_save — a real, non-WebDriver browser driven by screen only, on a libvirt guest (on=vm:<name>), a pool guest (on=pool:<name>) or the desktop the toolserver runs on / delegates to (`on=screen monitor:
/loci list, pointers (the hugpy-station distribution feed), register, archive central locus registry (Postgres)
/handoff request, list, claim, station (seat-API probe) jump-in seats via hugpy-station
/session identity, pull, spin, list, release pull a live Claude Code session into a hugpy-station seat
/instructions tree, read, add composition guides plus create-only caller contributions
/b ask — B (the fleet's local model) reduces text / a ledger / a file to the context a question needs hugpy fleet /v1/chat/completions
/comms ping, send, inbox, poll, claim, ack, reply, delivered central board rows + live push to serves
/ledger put, get, list, template — structured handoff state per locus/task Postgres
/todo, /board add, batch, update, done, remove, list · board list/summary Postgres (scoped write path)
/exchange, /assess record, ingest_transcript, archive, sessions, usage · rolling state, focus, roll Postgres + rolling_reduce
/canvas, /issue, /vl, /vm, /image, /claude, /gpt design docs · issue memory · fleet vision models · VMs · images · Claude/Codex seat config (from abstract-claude / abstract-gpt when installed) various

Visual browser (/browser/*)

A browser nobody automated: no WebDriver, no debugging port, no page script. The tools see the screen and move the pointer, so they work on anything that draws a browser window.

browser_surfaces                         where can I drive?  (guests, the desktop, pool guests, the lab)
browser_lab(action="start")              a disposable clone of the controlled lab image -> on="vm:<clone>"
browser_go(url, on=...)                  Ctrl+L, type, Enter, wait until the page stops painting
browser_look(on=...)                     the screen as text: [{id, text, x, y}]  (no model)
browser_click(text="Sign in")            exact visible words (OCR)            | each click walks the
browser_click(id="t12")                  an item of the last look             | pointer there, presses,
browser_click(target="the gear icon")    a description (hugpy's VL model)     | and diffs the screen:
browser_click(cell="C4.B3")              a cell of browser_shot(grid=true)    | changed_at_click |
browser_click(x=412, y=388)              a pixel of the screenshot            | changed_elsewhere | no_change
browser_type(text="...", into_text="Email", submit=true)     browser_key("ctrl+l")
browser_move / browser_drag / browser_scroll / browser_wait(text=...) / browser_read
  • Surfaces (on=): vm:<domain> (libvirt guest on this host: virsh screenshot + QMP tablet/keyboard), pool:<name> (a guest on the delegated VM-pool node), screen | monitor:<n> | window:<title> | region:x,y,w,h (the desktop this toolserver runs on — mss + pyautogui — or the node ui is delegated to). Every coordinate is a pixel of the screenshot; origins, HiDPI scale and the tablet range are the surface's business.
  • Aiming (aim.py). text= is OCR over two readings (the page at 2x and every text line normalised to dark-on-light) and refuses to guess: absent words are an error, repeated words list their places (nth=). target= asks the VL model: a grounding model (Qwen2.5-VL) answers with a box on a wire image sized for it (sides in multiples of 28, >= 1024 image tokens), a second look on a magnified crop must confirm it — framed around any on-screen text the description names — and tiny results are snapped onto the element. Any other VL model aims by grid-image inference: a labelled grid drawn at wire size, the chosen cell magnified and gridded again, the last cell snapped to the element inside it. A bot that reads images itself uses the same grid: browser_shot(grid=true) -> browser_shot(grid=true, cell="C4") -> browser_click(cell="C4.B3").
  • Results are data a bot can branch on: {ok, point: {x, y, method, confidence}, verdict, regions, ...}; a target that is not on the screen is never clicked: {ok: false, error: {code: "not_visible"}, hint}.
  • The lab: hugpy-browser-lab is a shut-off libvirt image (Ubuntu GNOME, Firefox with a fresh profile and fixed policies, home page = a precision test page, no NIC). browser_lab boots/destroys disposable clones of it, so a test starts from the same pixels every time.

Measured 2026-10-06 (233 labelled targets on five mock screens; real Chrome on Xvfb; Firefox in a lab clone): pointer error 0 px on both surface kinds; exact-text aims reach 81 of 82 visible labels with none wrong (the previous exact-text path: 36 of 82, 4 wrong); description aims with Qwen2.5-VL-7B land on the element for 85% of 40 targets in a single look, and on the hardest kind — small icons, checkboxes and buttons told apart only by a row label or card title elsewhere on the screen — for 14 of 14 with the row-framed second look (median error 1 px). Guide: instructions/hugpy/hugpy-browser.

Operating instructions

Tool schemas describe individual calls; the built-in instruction tree documents how calls compose into repeatable workflows:

  • GET /tools/toolserver/ — MCP configuration, discovery, runtime-neutral channel comms, and optional runtime-specific wake-up adapters.
  • GET /instructions/ae/solcatcher/ — Solcatcher-specific entry point.
  • GET /instructions/a-brain/alpha/ — Alpha's capability-channel instructions.
  • MCP: instructions_tree, then instructions_read, through ts_call.

Authenticated MCP callers may create a new document with instructions_add at an instructions/... path. Creation is durable and immediately readable, but MCP intentionally exposes no update or delete tool. Those operations remain local to server administration, and built-in documents are immutable.

Safety gates

Backends load lazily, so the server boots on a headless box and a missing backend errors only when its tool is called. Beyond that:

  • /db/query — read-only gate: rejects anything that isn't a single SELECT/WITH, blocks stacked statements and data-modifying keywords. Use /db/fetch (identifier-composed, params-not-SQL) as the default read path.
  • /sys/run_cmd — disabled unless TOOLSERVER_CMD_ALLOWLIST=ls,grep,… is set; only allowlisted binaries run.
  • /fs/read_file · /fs/write_file — local-only; the underlying SSH/remote kwargs are never exposed at the boundary.
  • /session/pull — the agent's own MCP bridge (abstract-claude mcp) fills the calling session's identity (session id, user@host, cwd); the toolserver registers machine + session loci, records a pull handoff, and asks the station (HANDOFF_SPAWN_URL) to seat claude --resume <id> --fork-session there. Without a resume-capable station it stays pending (/session/spin retries) — never a silent fresh seat. A pulled session carries its OWN name in the station (name= → tmux seat and session locus; default sess-<id8>).
  • default loci — the station service user (vm_mgr) and the host are always in the /loci/pointers distribution: seeded once into the registry, re-merged into the feed even before the table exists (TOOLSERVER_DEFAULT_LOCI, TOOLSERVER_STATION_USER).
  • /ui/click_verify — the click→observe→verify loop: locate (text or image template) → click → re-capture → report whether the screen (or a region) changed. The half most tool APIs lack.

Authentication — the operator token is the ONLY gate

Every route requires TOOLSERVER_OPERATOR_TOKEN, sent as X-Operator-Token: <token> or Authorization: Bearer <token> (the MCP bridge sends both). There is no IP allow-list and no loopback bypass: a LAN, WireGuard or 127.0.0.1 caller without the header gets 401 {"error":"unauthorized"} exactly like the public internet (operator ruling 2026-09-29 — the nginx allow 192.168.x/deny all block that used to front toolserver.hugpy.ai was a second, redundant gate and is gone). With the env var unset the server fails closed (every gated route 401s; startup logs an error).

Open by design (they carry their own credential or expose nothing):

Path Why
GET /healthz liveness probe, {"ok": true} only
GET /endpoints?access=<TOOLSERVER_ENDPOINTS_TOKEN> read-only catalog capability for a browser link
/ch/<id>?t=<token> shareable comms link — per-channel token (channels.py)
/clients/heartbeat · /clients/work · /clients/result per-client token (clients.py)
GET / /console /ui the static console page where the operator types the token (wsgi.py)

TOOLSERVER_REQUIRE_TOKEN (the old opt-in blueprint gate) is deprecated: parsed, logged as ignored, never enforced.

Configuration

Env var Purpose
TOOLSERVER_OPERATOR_TOKEN required — the only access gate (see Authentication)
TOOLSERVER_ENDPOINTS_TOKEN optional read-only capability for GET /endpoints?access=
TOOLSERVER_HOST / TOOLSERVER_PORT / TOOLSERVER_DEBUG bind + debug
TOOLSERVER_CMD_ALLOWLIST comma-separated binaries /sys/run_cmd may run
SOLCATCHER_POSTGRESQL_* DB connection (via abstract_database)
HANDOFF_SPAWN_URL / HANDOFF_SPAWN_TOKEN hugpy-station seat API (/api/handoff/spawn) + its X-Console-Token
TOOLSERVER_DEFAULT_LOCI name=user@host[:port][|goal],… — loci every station inherits (default: vm_mgr + this login on this host)
TOOLSERVER_STATION_USER station service user for the fallback default locus (default vm_mgr)

canvas.* — per-locus ◳ design / flow documents (2026-09-03)

The station's ◳ canvas tab (⬚ design = wireframe.v1, ⋔ flow = flow.v1) and every seat share ONE copy per (locus, kind) in the canvas table:

  • POST /canvas/get {locus, kind} → {state|null, rev, by, note, updated}
  • POST /canvas/put {locus, kind, state, by?, note?, notify?} — whole document, validated fail-closed, stored verbatim; a flow's rev bumps on every changed put; notify=true also posts a [canvas] high-priority request on the locus board (a deliberate hand-off — never for autosave).
  • POST /canvas/list {locus?} → which loci hold which kinds (no bodies).

Writes fire on the locus_change bus (table canvas, id = kind) so open drawers reload live. Through abstract-claude mcp these are the Claude Code tools canvas_get / canvas_put / canvas_list.

B on call — b_ask (2026-10-02)

b_ask(question, text= | ledger_locus=,ledger_task= | path=, model=) sends the source plus the question to B, the fleet's local model, and returns {answer, excerpts[], model, source, chars, clipped}. Excerpts are verbatim passages from the source; B is told never to invent content. This is how an agent reads an oversized ledger or file without loading it whole.

  • Model: TOOLSERVER_B_MODEL → HANDOFF_POLISH_MODEL → HANDOFF_JUDGE_MODEL → default Qwen3-Coder-Next-GGUF. Transport: HUGPY_BASE + HUGPY_API_KEY, the same fleet JSON chat the roller uses.
  • TOOLSERVER_B_MAX_CHARS (48000) caps the source sent; clipped: true says it was cut. TOOLSERVER_B_TIMEOUT (120 s).
  • Latency: ~18 s warm on Coder-Next for a 33 KB file; the first call after an idle period includes the model load (~100 s).
  • Ledgers are read straight from the DB, so a b_ask over a ledger is never hit by the MCP result governor below.

Comms

comms_ping writes a durable [ping] board row, then (0.0.47+) makes a best-effort live push to the target locus's registered pointer.serve_url (POST <serve>/api/session/message, to=keeper, wait=false, 5 s, never blocks; result.push reports pushed to …). The lookup ignores the registry's archive status. from_/to ids must match [a-z0-9][a-z0-9_-]{0,62} — no colons. kind=message rows are drained into the target Station's mail tab and auto-closed.

Ledgers, handoffs, boards

  • Ledgers (ledger_put/get/list/template) are the handoff state of record per locus/task: goal, rulings, decisions, world state, in-flight step, open questions, pointers. A put replaces the whole doc; a sections-only put fails (500).
  • Handoffs (handoff_request/pull/claim) hold the seat handoff; the old file-pointer handoffs are retired.
  • Boards: todo_update refuses unknown ids (no todo <id>); board-item format is checked per TOOLSERVER_BOARD_FORMAT=warn|enforce.
  • instructions_add needs both path and text.

Exchange recording — every operator message is on record (2026-10-06)

exchange_ingest_transcript (rolling.py prime_transcript) turns a Claude Code transcript into exchanges rows, one per operator message:

  • A message the operator sends while the agent is working (Claude Code's queued_command attachment, origin.kind = human) is its own row, holding what answered it; pointers.mid_turn = true. Before this, those messages were dropped entirely.
  • The incremental floor is per session, not per locus. Parallel sessions on one locus used to push the locus max(ts) past a turn that was still running, and that turn was never ingested. since_ts=0 still re-primes a whole transcript (idempotent per turn uuid).
  • Recording still happens at the Stop hook (end of turn). Prompt-first recording (write the prompt before the model sees it) is the open next step.

Rolling state

rolling_reduce.py is deterministic — no model call. It folds primed exchange rows into objective / done / in_progress / blockers / next_steps / key_paths / open_questions / init_prompt (each list capped at 30). Since 0.0.48, tool-failure blockers expire once newer turns arrive, prose blockers drop when the objective changes, and nothing expires on an empty slice. Station's banner shows blockers[0..3] of /api/frontier/state. An optional model polish (HANDOFF_POLISH=1, HANDOFF_POLISH_MODEL, ≤30 s, skipped if busy) and the legacy judge fold (HANDOFF_REDUCER=judge) remain available.

MCP bridge and the result governor

abstract-toolserver-mcp exposes every tool as mcp__toolserver__<name>; ts_call reaches any tool by name, including ones added after the MCP client loaded its tool list. session_message accepts the cross-locus target <locus>:keeper.

Every MCP result over AC_MCP_RESULT_MAX_LINES (200) or AC_MCP_RESULT_MAX_BYTES (16384) is cut to 75 % head + 25 % tail, the full text is archived (exchange_archive_get locus=… source=mcp-result:<tool>:<ts>), and a [truncated: …] marker is appended. AC_MCP_GOVERNOR=0 disables it; AC_MCP_GOVERNOR_EXEMPT lists tools never cut (default: the fs_read_* tools and exchange_archive_get). Prefer b_ask over reading a truncated result.

Two things the bridge does for tools that take many optional arguments or return pictures (since the 2026-10-06 release):

  • Schemas come from signatures. ?help cannot tell "no default" from "default None", so the bridge used to mark every optional-None parameter required. After discovery it now overlays the server's own catalog (POST /mcp?mode=flat): only parameters without a default are required, and annotated types are kept.
  • Images are images. A result value that is an image data URL (image, data_url, ...) or raw base64 under image_b64 / overlay_b64 is lifted into an MCP image content block and replaced in the JSON by a one-line note, instead of reaching the model as a megabyte of base64 the governor then cuts. AC_MCP_IMAGES=0 disables it.

Deployment on a host

One Flask service per host: unit 7004_hugpy_toolserver (gunicorn as vm_mgr, 127.0.0.1:7004) running an editable install of the dev tree, so a service restart picks up source edits without a publish. Discovery writes endpoint.json. Token resolution order: HUGPY_TOOLSERVER_TOKEN → TOOLSERVER_OPERATOR_TOKEN → TOOLSERVER_TOKEN → HUGPY_OPERATOR_TOKEN → STATION_CONSOLE_TOOLSERVER_TOKEN → env files. Auth fails closed when no token is set; /ch/ channels and per-client token paths are exempt; the /vms/ page uses cookie auth.

Attention-worthy

  • The MCP tool list is fixed when a client session starts: a new tool needs a service restart and a new session to appear as mcp__toolserver__* (use ts_call meanwhile).
  • ledger_get returns the doc twice (doc + sections), so a 17 KB ledger trips the 16 KB governor — read big ledgers through b_ask.
  • The roller's background exchange_ingest_transcript logs no transcript at (none found) when a locus has no transcript; harmless.
  • Fleet usage.prompt_tokens can under-report on the hugpy route.

Metadata

Release files for abstract-toolserver 0.0.67

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for abstract-toolserver 0.0.67
File Size Uploaded
abstract_toolserver-0.0.67.tar.gz 499.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for abstract-toolserver 0.0.67
File Interpreter ABI Platform
abstract_toolserver-0.0.67-py3-none-any.whl Python 3 none any Details

Total release size: 914.9 kB

Release files / abstract_toolserver-0.0.67.tar.gz

Download URL abstract_toolserver-0.0.67.tar.gz
Size 499.5 kB
Tags Source
SHA-256 checksum
How to use checksums
397d2b264959878745dce24232197d2b2a35eb93883845ecc16ca22ebcb0e07d
BLAKE2b-256 checksum
How to use checksums
86b6666913c33d37ef7ea5f747e9abc2ffb4c3e3766e83c9bd01f6d27ab25b18
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / abstract_toolserver-0.0.67-py3-none-any.whl

Download URL abstract_toolserver-0.0.67-py3-none-any.whl
Size 415.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
28d67e1c4bc176e68b41925056adbf6b60c9871f8d48038f48c0c9b5212a6e9d
BLAKE2b-256 checksum
How to use checksums
e801f5fd51493d872704487c69f3b3cab297d03d80882a5242f9f086e4016196
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

This release

0.0.67 This release

2 release files

0.0.55

1 release file

0.0.54

1 release file

0.0.53

1 release file

0.0.52

1 release file

0.0.51

1 release file

0.0.50

1 release file

0.0.49

1 release file

0.0.48

1 release file

0.0.47

1 release file

0.0.46

1 release file

0.0.45

1 release file

0.0.44

1 release file

0.0.43

1 release file

0.0.31

2 release files

0.0.30

2 release files

0.0.29

2 release files

0.0.28

1 release file

0.0.17

2 release files

0.0.16

2 release files

0.0.15

2 release files

0.0.14

2 release files

0.0.13

2 release files

0.0.12

2 release files

0.0.11

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page