agent-forge
A production-shaped multi-agent orchestration skeleton on LangGraph — with tracing, retry/fallback policies, and token budgets treated as first-class concerns, not afterthoughts.
I built this because most agent demos I reviewed while designing production systems stopped at "chain some prompts together." The hard 20% — observability, cost control, failure policy — was always left as an exercise for the reader. This repo is my opinionated take on what that missing 20% looks like, extracted into a small, reusable framework skeleton.
Architecture
flowchart LR
Q[User query] --> AF[AgentForge.run]
AF --> GB[GraphBuilder]
GB --> LG[(LangGraph StateGraph)]
LG --> N1[plan node]
N1 --> N2[execute node]
N2 --> N3[synthesize node]
N2 --> TR[ToolRegistry]
TR --> T1[CalculatorTool]
TR --> T2[WebSearchTool stub]
N1 -. wrapped .-> OBS[OTel tracing]
N2 -. wrapped .-> OBS
N3 -. wrapped .-> OBS
N1 -. metered .-> BUD[BudgetTracker]
N2 -. metered .-> BUD
N3 -. metered .-> BUD
N1 & N2 & N3 -. retried .-> POL[Retry / fallback policies]
Every node that enters the graph passes through three cross-cutting wrappers, applied in this order:
- Tracing — an OpenTelemetry span is opened per node execution (name, duration, token usage as span attributes). Export is your choice: OTel Collector, Jaeger, or LangSmith via env vars.
- Budget — token and cost usage recorded against a per-run
RunBudget. Exceeding the budget raisesBudgetExceeded, which halts the run cleanly instead of producing a surprise invoice. - Policy — nodes that call an LLM go through a tenacity-based retry decorator; model calls go through a fallback chain (
primary → secondary → ...) so a provider outage degrades instead of failing.
Features
- Graph builder with guard rails — register typed nodes, wire edges, add conditional routing; validates the graph (dangling edges, missing entry point) before LangGraph ever sees it.
- Tool plugin interface — one
BaseToolcontract (name,description,run()); ships with a safe-AST calculator and an honestly-stubbed web-search tool with a clear extension point. - Retry & fallback policies — a tenacity-based decorator factory, plus a
FallbackChainthat routes model calls across providers. - Per-run token/cost budgets — a
BudgetTrackerwith aggregate and per-node caps, scoped per run via acontextvarstoken so concurrent runs don't leak usage into each other. - OpenTelemetry tracing — spans around every node execution; optional LangSmith-compatible env configuration.
- Deterministic mock LLM — the example runs end-to-end with no API key, so CI and first-time contributors get a green run out of the box.
Quickstart (60 seconds)
pip install -e ".[dev]"
python examples/research_agent.py
Expected output (abridged):
== agent-forge example: research agent ==
Query: What is the capital of France, and what is 17 * 24?
--- plan ---
[1] answer-geography: Identify the capital of France.
[2] answer-math: Compute 17 * 24.
--- execute ---
answer-math -> 408
...
--- synthesize ---
The web search tool is not configured in this demo environment. 408.
Usage: input=300 output=150 total=450
Nodes visited: plan -> execute -> synthesize
Point it at a real model by setting OPENAI_API_KEY and swapping the mock client in examples/research_agent.py — the extension point is marked with a comment.
Tech stack
| Layer | Choice | Why |
|---|---|---|
| Orchestration | LangGraph | Stateful, cyclical graphs; checkpointing story; not a prompt-chaining framework |
| Data contracts | Pydantic v2 | Runtime-validated state, budgets, and tool results |
| Resilience | Tenacity | Battle-tested retry semantics (wait, stop, retry-on) |
| Observability | OpenTelemetry SDK | Vendor-neutral traces; LangSmith compatible via env config |
| LLM client | OpenAI SDK | Kept behind a thin client protocol so other providers slot in |
| Tooling | Ruff, mypy (strict), pytest | Fast lint + strict typing; tests mock all LLM boundaries |
Design decisions
The rationale behind the big calls lives in docs/adr/. The short version:
- LangGraph over CrewAI (ADR-0001): I need control over state shape, routing, and checkpoints — CrewAI's role abstraction hides exactly the seams I care about.
- OpenTelemetry for tracing (ADR-0002): agent traces are most valuable next to your existing infra traces, not in a separate silo. OTel spans carry token usage as first-class attributes.
- Budget as a first-class concern (ADR-0003): cost control is a cross-cutting runtime policy (like auth), not something each node should remember to do. The tracker lives in a context var; nodes report usage, the framework enforces limits.
Full system design write-ups: HLD and LLD.
Repository layout
src/agent_forge/
agent.py # AgentForge: public facade tying everything together
graph.py # GraphBuilder, ForgeState, node wrapper (budget + tracing + policy)
budget.py # RunBudget, BudgetTracker, BudgetExceeded
policies.py # retry decorator factory, FallbackChain
tools/ # BaseTool, ToolRegistry, calculator + web-search stub
tracing.py # OTel span wrappers, LangSmith env config
examples/
research_agent.py # plan -> execute -> synthesize, mock LLM, no API key needed
docs/
hld.md lld.md adr/
tests/
Roadmap
- Checkpointing / durable execution via LangGraph checkpointers
- Streaming intermediate node outputs to the caller
- Cost tables per model (tokens → USD) loaded from config
- Human-in-the-loop interrupt nodes (approval gates)
- Async node execution for I/O-bound tool calls
- A second example: tool-using customer-support agent with fallback model routing
Contributing
See CONTRIBUTING.md. PRs welcome; CI runs ruff, mypy, and pytest on Python 3.10–3.12.
License
MIT — see LICENSE.
Metadata
Release files for agent-forge-svkmsr6 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agent_forge_svkmsr6-0.1.0.tar.gz | 26.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agent_forge_svkmsr6-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 49.9 kB
Release files / agent_forge_svkmsr6-0.1.0.tar.gz
| Download URL | agent_forge_svkmsr6-0.1.0.tar.gz |
|---|---|
| Size | 26.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8dafcbc32aa05067dc625c15a18cff99e3371121df90352cf7cefc8fdff8d255
|
|
BLAKE2b-256 checksum How to use checksums |
3239fc411503df7caa9af3741273dc0e5e9a4f158423e632d0b09a8dd49822b5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|
Release files / agent_forge_svkmsr6-0.1.0-py3-none-any.whl
| Download URL | agent_forge_svkmsr6-0.1.0-py3-none-any.whl |
|---|---|
| Size | 23.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b69d644a36f7a3ada195778cc2a46687fc8d982a3de0fc986a10631c80fd36c5
|
|
BLAKE2b-256 checksum How to use checksums |
a6ff0d10824ad93c2b55fc99c612c9506f384dd5ad0ceec1d9845499ac05271a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|