Skip to main content

agent-forge

A production-shaped multi-agent orchestration skeleton on LangGraph — with tracing, retry/fallback policies, and token budgets treated as first-class concerns, not afterthoughts.

I built this because most agent demos I reviewed while designing production systems stopped at "chain some prompts together." The hard 20% — observability, cost control, failure policy — was always left as an exercise for the reader. This repo is my opinionated take on what that missing 20% looks like, extracted into a small, reusable framework skeleton.

Architecture

flowchart LR
    Q[User query] --> AF[AgentForge.run]
    AF --> GB[GraphBuilder]
    GB --> LG[(LangGraph StateGraph)]
    LG --> N1[plan node]
    N1 --> N2[execute node]
    N2 --> N3[synthesize node]
    N2 --> TR[ToolRegistry]
    TR --> T1[CalculatorTool]
    TR --> T2[WebSearchTool stub]
    N1 -. wrapped .-> OBS[OTel tracing]
    N2 -. wrapped .-> OBS
    N3 -. wrapped .-> OBS
    N1 -. metered .-> BUD[BudgetTracker]
    N2 -. metered .-> BUD
    N3 -. metered .-> BUD
    N1 & N2 & N3 -. retried .-> POL[Retry / fallback policies]

Every node that enters the graph passes through three cross-cutting wrappers, applied in this order:

  1. Tracing — an OpenTelemetry span is opened per node execution (name, duration, token usage as span attributes). Export is your choice: OTel Collector, Jaeger, or LangSmith via env vars.
  2. Budget — token and cost usage recorded against a per-run RunBudget. Exceeding the budget raises BudgetExceeded, which halts the run cleanly instead of producing a surprise invoice.
  3. Policy — nodes that call an LLM go through a tenacity-based retry decorator; model calls go through a fallback chain (primary → secondary → ...) so a provider outage degrades instead of failing.

Features

  • Graph builder with guard rails — register typed nodes, wire edges, add conditional routing; validates the graph (dangling edges, missing entry point) before LangGraph ever sees it.
  • Tool plugin interface — one BaseTool contract (name, description, run()); ships with a safe-AST calculator and an honestly-stubbed web-search tool with a clear extension point.
  • Retry & fallback policies — a tenacity-based decorator factory, plus a FallbackChain that routes model calls across providers.
  • Per-run token/cost budgets — a BudgetTracker with aggregate and per-node caps, scoped per run via a contextvars token so concurrent runs don't leak usage into each other.
  • OpenTelemetry tracing — spans around every node execution; optional LangSmith-compatible env configuration.
  • Deterministic mock LLM — the example runs end-to-end with no API key, so CI and first-time contributors get a green run out of the box.

Quickstart (60 seconds)

pip install -e ".[dev]"
python examples/research_agent.py

Expected output (abridged):

== agent-forge example: research agent ==
Query: What is the capital of France, and what is 17 * 24?

--- plan ---
[1] answer-geography: Identify the capital of France.
[2] answer-math: Compute 17 * 24.

--- execute ---
answer-math -> 408
...

--- synthesize ---
The web search tool is not configured in this demo environment. 408.
Usage: input=300 output=150 total=450
Nodes visited: plan -> execute -> synthesize

Point it at a real model by setting OPENAI_API_KEY and swapping the mock client in examples/research_agent.py — the extension point is marked with a comment.

Tech stack

Layer Choice Why
Orchestration LangGraph Stateful, cyclical graphs; checkpointing story; not a prompt-chaining framework
Data contracts Pydantic v2 Runtime-validated state, budgets, and tool results
Resilience Tenacity Battle-tested retry semantics (wait, stop, retry-on)
Observability OpenTelemetry SDK Vendor-neutral traces; LangSmith compatible via env config
LLM client OpenAI SDK Kept behind a thin client protocol so other providers slot in
Tooling Ruff, mypy (strict), pytest Fast lint + strict typing; tests mock all LLM boundaries

Design decisions

The rationale behind the big calls lives in docs/adr/. The short version:

  • LangGraph over CrewAI (ADR-0001): I need control over state shape, routing, and checkpoints — CrewAI's role abstraction hides exactly the seams I care about.
  • OpenTelemetry for tracing (ADR-0002): agent traces are most valuable next to your existing infra traces, not in a separate silo. OTel spans carry token usage as first-class attributes.
  • Budget as a first-class concern (ADR-0003): cost control is a cross-cutting runtime policy (like auth), not something each node should remember to do. The tracker lives in a context var; nodes report usage, the framework enforces limits.

Full system design write-ups: HLD and LLD.

Repository layout

src/agent_forge/
  agent.py       # AgentForge: public facade tying everything together
  graph.py       # GraphBuilder, ForgeState, node wrapper (budget + tracing + policy)
  budget.py      # RunBudget, BudgetTracker, BudgetExceeded
  policies.py    # retry decorator factory, FallbackChain
  tools/         # BaseTool, ToolRegistry, calculator + web-search stub
  tracing.py     # OTel span wrappers, LangSmith env config
examples/
  research_agent.py   # plan -> execute -> synthesize, mock LLM, no API key needed
docs/
  hld.md  lld.md  adr/
tests/

Roadmap

  • Checkpointing / durable execution via LangGraph checkpointers
  • Streaming intermediate node outputs to the caller
  • Cost tables per model (tokens → USD) loaded from config
  • Human-in-the-loop interrupt nodes (approval gates)
  • Async node execution for I/O-bound tool calls
  • A second example: tool-using customer-support agent with fallback model routing

Contributing

See CONTRIBUTING.md. PRs welcome; CI runs ruff, mypy, and pytest on Python 3.10–3.12.

License

MIT — see LICENSE.

Metadata

Release files for agent-forge-svkmsr6 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-forge-svkmsr6 0.1.0
File Size Uploaded
agent_forge_svkmsr6-0.1.0.tar.gz 26.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-forge-svkmsr6 0.1.0
File Interpreter ABI Platform
agent_forge_svkmsr6-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 49.9 kB

Release files / agent_forge_svkmsr6-0.1.0.tar.gz

Download URL agent_forge_svkmsr6-0.1.0.tar.gz
Size 26.7 kB
Tags Source
SHA-256 checksum
How to use checksums
8dafcbc32aa05067dc625c15a18cff99e3371121df90352cf7cefc8fdff8d255
BLAKE2b-256 checksum
How to use checksums
3239fc411503df7caa9af3741273dc0e5e9a4f158423e632d0b09a8dd49822b5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / agent_forge_svkmsr6-0.1.0-py3-none-any.whl

Download URL agent_forge_svkmsr6-0.1.0-py3-none-any.whl
Size 23.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b69d644a36f7a3ada195778cc2a46687fc8d982a3de0fc986a10631c80fd36c5
BLAKE2b-256 checksum
How to use checksums
a6ff0d10824ad93c2b55fc99c612c9506f384dd5ad0ceec1d9845499ac05271a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page