title: forge · Forging ideas into action type: index status: 已完成 owner: 待指定 created: 2026-08-29 updated: 2026-09-20 tags: [索引] evidence_level: ⚠️ 间接信号 source: "internal knowledge base" dikw_exempt: true dikw_exempt_reason: "索引页;D/I/K/W 分层由被索引的文档承担"
forge
本页为索引:D/I/K/W 分层由被索引的文档承担,本页只做汇总与导航。
A ReAct agent you can actually read — the step loop is ~235 lines of straight-line Python, and the other 31 modules are opt-in layers you can skip.
Zero heavy dependencies (stdlib + openai + pyyaml). Every mechanism is spelled out instead of hidden: retry · fallback · circuit breaker · approval gates · command sandbox · structured logs · four-role Computer Use loop. Runs against any OpenAI-compatible endpoint — DeepSeek, Qwen, vLLM, Ollama, local models.
中文版 README · docs/ · Releases · CHANGELOG · 373 tests, all offline
pip install "git+https://github.com/musokean/forge.git" # or: clone && pip install -e .
export DEEPSEEK_API_KEY=sk-xxx
forge # interactive REPL
forge "帮我算 (3+5)*2" # one-shot question
forge --web # browser chat UI (zero-dependency HTTP server)
Why forge?
Most agent projects fall into two camps:
- Production-grade giants (OpenHands, AutoGen, MetaGPT): tens of thousands of lines, powerful ecosystems — but you can't read them to understand how agents work.
- Minimal demos (smolagents-style ~1k lines, tutorials): readable, but they stop at "it runs" — no engineering foundation.
forge sits in the gap: the ~150-line ReAct loop is implemented directly, every mechanism (retry / fallback / circuit breaker / rolling summary / approval gates) is explainable, and it ships with a complete engineering foundation — multi-agent orchestration, a self-contained knowledge base, long-term memory, golden-set evaluation, and a zero-dependency philosophy.
Zero heavy dependencies: standard library + openai SDK + pyyaml. Works with any OpenAI-compatible endpoint — DeepSeek, Qwen, vLLM, Ollama, local models. Chinese-first, domestic-model friendly.
Quick start
# install from GitHub (or clone the repo and run: pip install -e .)
pip install "git+https://github.com/musokean/forge.git"
# set your API key (env var, picked up automatically)
export DEEPSEEK_API_KEY=sk-xxx
forge # interactive REPL
forge "帮我算 (3+5)*2" # one-shot question
forge --web # browser chat UI (zero-dependency HTTP server)
forge --serve --port 8080 # HTTP API service (pip install "handcraft-agent[server]" first)
First run auto-generates a default config/models.yaml (if missing) — no config file, no crash. Edit it (or /config in the REPL) to switch models / roles / endpoints. Model registry → roles → debate lineup → routing → knowledge base path, all config-driven, no code changes.
Installed with pip? It goes to ~/.forge/config/models.yaml (or set FORGE_CONFIG to
point at your own) — a config in the current directory always wins.
Feature highlights
| Area | What you get |
|---|---|
| Core loop | Direct ReAct loop implementation with loop-guard, streaming output, reasoning display |
| Engineering | Tool read-only tiers, write-operation approval gates, exponential-backoff retry, model fallback, circuit breaker, token accounting, per-step trace |
| Multi-agent | Parallel task fan-out, multi-role debate (pro/con/judge), supervisor plan→execute→merge, automatic task routing |
| Context | Token-budget truncation, rolling summary via cheap model, tool-output clipping |
| Knowledge | Self-contained SQLite+FTS5 knowledge base (the index is the source), Chinese trigram search, one-key ingest/sync/export |
| Memory | Cross-session user profile auto-recalled per query |
| Reliability | Golden-set regression (/eval, keyword-hit + LLM-as-judge), model-failure resilience, endpoint self-check on startup |
| UX | Sky-blue theme, interrupt/redirect generation (Esc / type a steer), auto tasks, Web UI |
| Service (#14) | HTTP API (forge --serve): multi-session persistence, API-key auth (loopback-only by default), per-caller rate limiting, Swagger docs at /docs |
| Client executor (#17) | Drive remote PCs: a light executor on each machine dials out (long poll, no inbound port) and exposes shell / files / screenshot / GUI input behind two policy layers. A four-role Computer Use loop (planner → executor → evaluator → supervisor) keeps one model from being brain, hand and judge at once |
| Voice (#11) | Cascade voice pipeline with streaming transcription, sentence-level synthesis and barge-in; audio source and playback are injectable, so the whole mechanism is tested in CI without a microphone or a sound |
| Hardware (#16) | Serial / MQTT real link behind a control plane: asset registry, staged policy, command state machine (Created→Sent→Accepted→Applied) with timeout, retries and rollback, plus agent-side temperature/runtime guards. device_sim.py speaks the same protocol, so the whole link is testable with no hardware |
| Safety (#4) | Command sandbox: Docker isolation when available (no network, read-only mount, memory/CPU/PID caps, non-root), hardened local fallback, dangerous-command blocking. Host environment is never handed to child processes — a command can no longer read your API keys |
| Logging (#7) | Structured JSONL logs with rotation, retention and secret redaction; per-run correlation ids (role/model/steps/tokens/latency); HTTP request log; /logs to inspect |
Commands
/reset /usage /trace /kb /export /key /model /config /circuit
/skill /memory /remember /task /eval /web /serve /logs /sandbox /device /executor /help /exit
/key sk-xxx — paste a key, auto-assigns to the main model. /config — guided panel, no YAML hand-editing needed.
Project layout
handcraft-agent/
├── config/models.yaml # all configuration (models/roles/debate/router/kb)
├── config/golden.yaml # golden-set eval cases
├── forge/
│ ├── agent.py # ReAct loop + context mgmt + status bar + approval
│ ├── llm.py # openai gateway + retry + fallback + streaming + breaker
│ ├── tools.py # 20 tools + read-only tiers + KB tools
│ ├── orchestrator.py # parallel / debate / supervisor
│ ├── router.py # rule-first task routing (0ms for common intents)
│ ├── knowledge.py # SQLite+FTS5 knowledge base
│ ├── memory.py # cross-session user profile
│ ├── eval.py # golden-set evaluation
│ ├── web.py # zero-dependency web chat
│ ├── server.py # #14 HTTP API service (sessions + auth + rate limit)
│ ├── hwproto.py # #16 hardware protocol v1 (line-JSON + CRC + seq/ack/state)
│ ├── hwtransport.py # #16 transports: serial (pyserial) / MQTT (paho) / memory
│ ├── hwcontrol.py # #16 control plane: assets + policy + command state machine
│ ├── executor_hub.py # #17 executor hub (registry + policy + command queue)
│ ├── executor.py # #17 client executor (capabilities + path jail + long poll)
│ ├── cua.py # #17 four-role Computer Use loop (planner/executor/evaluator/supervisor)
│ └── executor_agent.py # #17 entry point that runs on the controlled PC
│ ├── sandbox.py # #4 command sandbox (Docker isolation / hardened local)
│ ├── logging_setup.py # #7 structured JSONL logs (rotation, retention, redaction)
│ ├── keypress.py # interrupt/steer during generation
│ └── ...
├── main.py # CLI entry
└── test_*.py # milestone + stress + module tests (all mock, no network)
Server mode (HTTP API)
Turn forge into an HTTP service with sessions, auth and rate limiting (module #14):
pip install "handcraft-agent[server]" # optional extra: fastapi + uvicorn
forge --serve --port 8080 # or "/serve 8080" inside the REPL
- Sessions — every conversation is persisted in SQLite (
data/sessions.db), each with its own Agent context: clients can disconnect and resume later, or keep several threads apart. - Auth —
Authorization: Bearer <key>orX-API-Key: <key>. Keys come fromserver.api_keysinconfig/models.yaml, or theFORGE_API_KEYenv var (comma-separated,env:VARindirection supported). No keys configured → loopback-only (convenient locally; never expose that to the internet). - Rate limit —
server.rate_limit_per_min(default 60) per caller; beyond it you get429+Retry-After. - Write safety — a service has no interactive approval channel, so write tools are rejected by default; keep those in the CLI. Override with
server.approve_modeonly for controlled deployments. - Interactive docs:
http://127.0.0.1:8080/docs.
| Method | Path | Purpose |
|---|---|---|
| GET | /healthz |
liveness probe (no auth) |
| GET | /api/status |
model, session count, auth mode, rate limit |
| POST | /api/chat |
{"message": "...", "session_id": "optional"} → reply + token usage |
| POST | /api/sessions |
create a session ({"title": "optional"}) |
| GET | /api/sessions |
list sessions |
| GET | /api/sessions/{id} |
session + full message history |
| DELETE | /api/sessions/{id} |
delete a session |
curl -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"message": "hello"}' http://127.0.0.1:8080/api/chat
Client executor (#17)
Drive other PCs. Each controlled machine runs a light executor that only dials out (HTTP long poll - no inbound port, no firewall change):
# centre
forge --serve
# controlled PC (one per machine)
pip install "handcraft-agent[executor]" # optional: adds screenshot + GUI input
python executor_agent.py --center http://<centre>:8080 --token <KEY> --id pc-01 --root D:/work
# centre REPL
/executor # hub overview + what is online
/executor run pc-01 "whoami" # run a command over there
/executor cua pc-01 "open notepad and type hello"
Capabilities are declared by the client and filtered twice: at the hub (allow-list, staged release
readonly/low_risk/approval/closed_loop, timeout, size caps) and again on the client (allow-list,
path jail, size caps, and shell goes through that machine's own sandbox). Without the GUI extra
the client simply does not declare screenshot/input instead of pretending.
Four-role Computer Use. cua_task splits the loop so no single model is brain, hand and judge:
planner → executor → (dispatch → raw evidence) → evaluator → supervisor on repeated failure. The
evaluator only sees raw evidence, never the executor's own explanation; a completion gate catches the
common case where the executor never says "done". See docs/executor.md.
Hardware (#16)
Devices are tools. tools.py exposes device_status / device_power / device_level /
device_reset; what sits behind them depends on config:
device.enabled / transport |
What the tools talk to |
|---|---|
false or sim (default) |
Phase 0 in-process simulator (fake_device.py) |
true + serial |
real serial / UART (COM5, /dev/ttyUSB0, or socket://host:port) |
true + mqtt |
MQTT: down {prefix}/cmd/{device}, up {prefix}/up/{device} |
Control plane. Talking to hardware is easy; not mis-controlling it is the hard part. Before
a command reaches the wire it passes an asset registry (unknown device aliases are refused, not
guessed), a policy engine (staged release readonly / low_risk / approval / closed_loop,
level/temperature/runtime limits, write cooldown, remote endpoints unwritable by default) and a
command state machine:
Created ──sent──> Sent ──ack.ok──> Accepted ──state──> Applied
└─ ack rejected ─> Rejected
└─ timeout, retries exhausted ─> Timeout (then rollback)
An ack only means the device took the request; only a state snapshot counts as applied. Every
command is audited, and the agent side enforces its own over-temperature / runtime guards instead of
trusting the device alone.
No hardware needed to verify it. The device-side simulator speaks the same protocol over TCP:
# terminal 1 — simulated device (real protocol, real CRC)
python device_sim.py --transport socket --port 9009 --test-hooks
# terminal 2
forge # then:
/device mode serial socket://127.0.0.1:9009
/device connect
/device # device identity, policy, recent audit
/device audit 10
That path exercises real pyserial, real framing, real policy, real state machine — only the physical
component is simulated. docs/hardware.md has the protocol spec plus a reference ESP32 firmware
(hardware/esp32_beauty_device.ino) for the real thing.
Voice (#11)
Talk to forge. Cascade pipeline (STT → agent → TTS) with streaming and barge-in:
pip install "handcraft-agent[voice]" # sounddevice + numpy + edge-tts + openai-whisper
forge --voice # talk, and interrupt it mid-answer
- Streaming (Phase 2) — while you speak the captured audio is re-transcribed every second and the draft appears before you finish; the answer is split into sentences as it generates, and each finished sentence is synthesised and played immediately, so you hear sentence one while sentence two is still being written
- Barge-in (Phase 3) — the microphone keeps listening while forge thinks and talks. Start speaking and it stops mid-sentence, cancels the running generation and takes the half-sentence you already said as the next turn — no repeating yourself
- Testable by design — audio source and playback are injectable (real microphone / an audio file standing in for one / scripted synthetic audio; real ffplay or a recorder that makes no sound), and VAD + sentence splitting are pure state machines. So segmentation, interruption timing and streaming order are covered in CI: no microphone, no model, no sound
- No microphone needed to try it:
forge --voice --audio-source file:question.wav --voice-sink nullruns the whole chain (real Whisper, real model, real edge-tts) silently - Acoustic echo cancellation —
forge --voice --aeckeeps the microphone live while the answer plays and subtracts the agent's own voice using the audio being played as the reference, so you can interrupt hands-free: no muting, no push-to-talk key. Pure-numpy NLMS by default (no C extension); it switches to in-process playback, sinceffplaycannot hand over the samples it is playing. - Speakers work too — you do not have to wear headphones:
forge --voice --half-duplexmutes the microphone while the answer plays (it still hears you while it is thinking, where there is no echo to confuse it), andforge --voice --pttonly captures while you hold space — press to stop it mid-sentence, release to send that utterance. Full duplex with headphones still gives the smoothest barge-in.
Metadata
Release files for handcraft-agent 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| handcraft_agent-0.5.0.tar.gz | 199.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| handcraft_agent-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 410.9 kB
Release files / handcraft_agent-0.5.0.tar.gz
| Download URL | handcraft_agent-0.5.0.tar.gz |
|---|---|
| Size | 199.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
76f84753c850bb2a906673ba07d5bd218d23c225d04185c662088b237f1c4cb0
|
|
BLAKE2b-256 checksum How to use checksums |
d63d10aad0e4a341e1e2a569e5614b268ddae509b57a6e1c275487faed14f516
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.1 {"installer":{"name":"uv","version":"0.12.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / handcraft_agent-0.5.0-py3-none-any.whl
| Download URL | handcraft_agent-0.5.0-py3-none-any.whl |
|---|---|
| Size | 211.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
63a0ef8e9bf7bd567c0ad693ed35cef01f017485fd54101907b4ffcb00d727a2
|
|
BLAKE2b-256 checksum How to use checksums |
0cba74a7224b1542108b8b1c0089d5f872397eaf1e9e8b36261149a8f8ba7cf1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.1 {"installer":{"name":"uv","version":"0.12.1","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|