Skip to main content

Self-hosted, mobile-friendly observability for LangGraph — traces, run history, cost, and a living graph view.

Project description

Windhover — a hovering kestrel

Windhover

PyPI CI MIT

Windhover — the old poetic name for the kestrel, the falcon that hangs motionless in the wind, watching everything below. This tool does the same for your agent graphs.

Self-hosted, mobile-friendly observability for LangGraph. Trace depth like LangSmith (LLM prompts, tokens, cost, latency — plus retrievers and human-in-the-loop interrupts), run history, a timing waterfall, per-node stats, error forensics down to the throwing source line — and a living graph view that auto-updates when your code's topology changes. Point it at any compiled graph, or trace runs in from your own app. No LangSmith account, no cloud tunnel, no fragile websocket. HTTP + SSE, MIT.

Nothing about your graph's domain is baked in. Topology, the input form, and run outputs all come from the graph itself. Windhover observes — it never edits your graph.

Living graph (parallel fan-out) Trace drawer — retrievers, LLM calls, cost, state
Graph view Trace drawer
Runs — search, tags, sessions, interrupts Dashboards — per-day, per-model
Runs table Stats

Quick start

pip install windhover langgraph
WINDHOVER_GRAPH=windhover.demo_graph:graph windhover   # -> :8090

Open http://<host>:8090. New run (input pre-filled from the graph's schema) → watch it execute → Runs for history, span trees, and replay → Stats for cost/latency. Edit the graph file while it runs and the canvas updates itself.

Your own graph: WINDHOVER_GRAPH="myapp.graphs:g" WINDHOVER_GRAPH_DIR=/path python -m windhover.server

Trace runs from any app

from windhover import WindhoverTracer
graph.invoke(input, config={"callbacks": [WindhoverTracer("http://HOST:8090")]})

Node spans, LLM calls (model/prompt/response/tokens/cost), and tools show up in Runs — wherever your app runs. Non-blocking, best-effort; never raises into your graph.

Sessions and tags use standard LangChain config — no Windhover imports needed beyond the tracer:

graph.invoke(input, config={
    "callbacks": [WindhoverTracer("http://HOST:8090")],
    "metadata": {"windhover_session": "chat-42", "windhover_tags": ["prod"]},
    "tags": ["also-captured"],          # langgraph-internal tags are filtered out
})

Features

  • Any graph — topology from graph.get_graph(); input form from its state schema.
  • Full trace tree — nodes → nested LLM / tool / retriever spans: prompts, responses, tokens, cost, latency, retrieved documents with their metadata.
  • Clickable graph — tap a node for health, latency, wiring, its source code, and recent executions with payloads.
  • Error forensics — failed runs show the full traceback; the failing node turns red on the graph, and the node's source renders with the throwing line highlighted.
  • Human-in-the-loop console — a paused graph shows an amber interrupted status with the question it's asking; answer it (Command(resume=…)), redirect it (Command(goto=…)), set static breakpoints per run (interrupt_before), edit state at any checkpoint (update_state), or fork a thread from any historical checkpoint — all from the UI, all pure LangGraph primitives.
  • State evolution — every trace shows which state keys each node wrote, in order.
  • X-ray — graphs with subgraphs get a canvas toggle that expands composite nodes (get_graph(xray=True)).
  • Search & filters — full-text over prompts/payloads/errors (FTS5, LIKE fallback), status/tag/session filters, bookmarks, pagination, CSV/JSON export.
  • Sessions — group runs into threads/batches; roll-up tokens, cost, errors.
  • Scores — attach numeric evals to runs (API or UI): eval harnesses, LLM-as-judge, human review.
  • Live tail — open a running run and watch spans arrive — including the model typing (streamed tokens flush into the span twice a second); nodes push progress via get_stream_writer().
  • Call configs — every LLM span records temperature/max-tokens/stream and the tools the model was offered; conditional-edge branch labels and add_node(metadata=…) render on the graph and node pane; graphs with a context schema get a runtime-context box on New run.
  • Custom eventsdispatch_custom_event("name", {...}) anywhere in your app lands as an event marker in the trace, parented to the node that fired it.
  • Retries + TTFT — tenacity retries badge the span (↻2); streaming LLM calls record time-to-first-token; cache-read / reasoning token details show on the model line.
  • Memory browser — graphs compiled with a LangGraph Store get a Memory view: browse namespaces and search long-term memory items.
  • Time-travel — checkpointed graphs get a per-thread checkpoint browser: state, writes, and next-nodes at every superstep (get_state_history).
  • Run diff — compare any two runs node-by-node: identical vs differing outputs, duration and token deltas.
  • Datasets / batch eval — store golden input sets, run the graph over them, and get an expected_match score per item (see Datasets on the Stats page).
  • Run history + replay — SQLite; runs persist even if the browser closes (worker thread).
  • Living graph — file watcher re-extracts topology in a subprocess and pushes it to the UI.
  • Dashboards — runs/tokens per day, per-model usage and latency, per-node latency, error rate.
  • Mobile-first PWA, light/dark. Fully local (FastAPI + Cytoscape.js).

Datasets API

curl -X POST :8090/api/datasets -H 'Content-Type: application/json' -d '{
  "name": "golden", "items": [
    {"input": {"n": 2},  "expected": 6},
    {"input": {"n": 40}, "expected": "big"}]}'
curl -X POST :8090/api/datasets/golden/run   # -> runs land in an eval:golden:<ts> session

Scores API

curl -X POST :8090/api/runs/RUN_ID/scores -H 'Content-Type: application/json' \
     -d '{"name": "accuracy", "value": 0.92, "comment": "vs golden set"}'

Config (env)

WINDHOVER_GRAPH (module:attr; unset = ingest-only) · WINDHOVER_GRAPH_DIR · WINDHOVER_DB · WINDHOVER_HOST/WINDHOVER_PORT (0.0.0.0/8090) · WINDHOVER_WATCH (1) · WINDHOVER_PRICING · WINDHOVER_RETENTION_DAYS (0 = keep forever; else prune older runs on startup + every 6h) · WINDHOVER_TOKEN (set to require Authorization: Bearer <token> — or ?token= — on all /api routes; the UI prompts once and remembers it). Edit windhover/pricing.json for your models' $/1M rates (unknown model → cost null).

Notes

Runs use the imported graph (restart to run new code); the view always reflects current-on-disk topology. All frontend assets are vendored — no CDN, works fully offline. Deep links: #runs, #sessions, #stats, #run=<id>.

License

MIT.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

windhover-0.11.0.tar.gz (732.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

windhover-0.11.0-py3-none-any.whl (307.8 kB view details)

Uploaded Python 3

File details

Details for the file windhover-0.11.0.tar.gz.

File metadata

  • Download URL: windhover-0.11.0.tar.gz
  • Upload date:
  • Size: 732.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for windhover-0.11.0.tar.gz
Algorithm Hash digest
SHA256 7d4a6c32bd0f9a92e37f19a63cacbeccba8ce781064912f90b74692f235dbc2f
MD5 9c0a44307de5275338a2b3584a7c7462
BLAKE2b-256 35923f67787633dc70896cf75c4eba628d6499442cf7552bd321c533afbc53e7

See more details on using hashes here.

File details

Details for the file windhover-0.11.0-py3-none-any.whl.

File metadata

  • Download URL: windhover-0.11.0-py3-none-any.whl
  • Upload date:
  • Size: 307.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for windhover-0.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 10e126ce15cb86a6cde659af038a1af98c1864a394842599eb6910f013a0fb95
MD5 d61566e05b2dbc108f40cf7c1042fbf0
BLAKE2b-256 4f7b828f03ddad62eaddd50e9e6093bfae1433dce99932a4e66567b64db3839a

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page