Make sense of what your AI agents did — local-first, MCP-native observability, trace replay & trajectory diff.
Project description
agentsense
Make sense of what your AI agents did. A local-first, MCP-native observability & debugging tool: a transparent MCP proxy traces every protocol message with zero code change to the agent, PII is redacted deterministically on the write path, and traces are stored whole (no field whitelist) in local SQLite — then replayed and diffed to see how a different model would have decided.
Apache-2.0 · local-first · no cloud signup
Compare view — the same run replayed against two models; the first point where their decisions diverge is highlighted.
Aim
Ship an open-source, MCP-native agent observability & debugging tool that combines four properties no existing tool has together: local-first, MCP-native (instruments inside the Model Context Protocol, not wrapped via OpenTelemetry), true trace replay + trajectory diff, and open source. The goal is a public release with a working demo and documented architecture — a debugger for AI agents that runs entirely on your machine.
Why it's useful
When an AI agent misbehaves — wrong tool, bad decision, runaway cost — you are usually guessing. agentsense gives you ground truth:
- Protocol-level tracing with zero code change. Point the agent's MCP config at the proxy and every tool call, input/output, latency, and cost is captured — even for agents you can't instrument.
- "Would a different model fix this?" Replay a recorded trace against another model/prompt, feeding back stored tool results, and diff the trajectory to see exactly where behaviour diverges — no live tool calls, no API cost, no side effects.
- Compliance-ready audit trails. PII is redacted deterministically before it is ever stored, and the redaction itself is logged — the audit-trail story for EU AI Act enforcement (from August 2026).
- Local-first. No cloud signup, fully self-hostable.
What's here (v0)
| Component | Module | Status |
|---|---|---|
| MCP proxy (stdio, transparent) | proxy/ |
✅ built first |
| Deterministic PII redaction | redaction/ |
✅ write-path |
| SQLite trace store (whole objects) | store/ |
✅ |
| Mocked replay + trajectory diff | replay/ |
✅ pluggable model client |
| Python capture SDK | sdk/ |
✅ OTel GenAI conventions |
| Local trace-explorer UI | ui/ |
✅ read-only, FastAPI |
Design rules (proven in Phase 0, enforced here)
- Proxy forwards raw bytes unchanged; only a copy is parsed for tracing.
- Logs go to stderr / file, never stdout (stdout is the JSON-RPC channel).
- Trace store preserves unknown/vendor fields — whole objects, no whitelist.
- Redaction is deterministic (hash-derived tokens) so replay aligns.
Install
pip install agentsense-ai # distribution name; the CLI and import are `agentsense`
# extras: pip install "agentsense-ai[ui,replay]" # web UI + live-replay model clients
The install/import names differ (like scikit-learn → sklearn): the package is
agentsense-ai, but you run agentsense … and import agentsense.
Quick start
uv sync --group dev
# Point your MCP client's server config at this command instead of the real server.
# The proxy forwards transparently and traces to SQLite.
uv run agentsense proxy --db traces.db -- \
npx -y @modelcontextprotocol/server-filesystem /tmp
Tests
uv run pytest # unit tests run offline; the proxy has no model dependency
uv run ruff check .
test_proxy_transparency.py is an integration test that drives server-filesystem
through the proxy with a real MCP client; it self-skips if npx is unavailable.
UI (local trace explorer)
A read-only web app over the trace store — no build step, no Node. It serves a JSON API and a single static page (trace list → span tree + timeline + redaction badges).
uv sync --extra ui
uv run agentsense ui --db traces.db # opens http://127.0.0.1:8000
The Compare tab renders the side-by-side trajectory diff: pick two captured traces and see their decision sequences aligned, with the first point of divergence highlighted.
Live replay — on any captured trace, pick a model and re-run it right there: agentsense rebuilds the recording, replays it against your chosen model (OpenAI-compatible/Ollama or AWS Bedrock), and diffs the result against what the agent actually did. Ephemeral — no store writes.
Endpoints: GET /api/traces, GET /api/traces/{id}/spans, GET /api/diff?a=&b=,
POST /api/replay. The frontend is plain HTML/JS in ui/static/ — swappable for React
later behind the same API.
Capture SDK
For agents you can instrument, the SDK records reasoning-level detail the proxy can't
see. It writes through the same store — and therefore the same redaction path — as the
proxy, using OpenTelemetry GenAI attribute naming (gen_ai.*) for interop.
from agentsense.sdk import Tracer
from agentsense.store import SpanStore
tracer = Tracer(SpanStore("traces.db"))
with tracer.session("booking-agent") as s:
s.step("plan", reasoning=..., model="claude-opus-4-8")
s.llm_call("claude-opus-4-8", messages=[...], response={...},
usage={"input_tokens": 1200, "output_tokens": 80})
s.tool_call("search_flights", args={...}, result={...}, cost=0.002)
Spans form a shallow tree (session → steps). llm_call stores the whole conversation
(messages, tools, response) so a captured run can be reconstructed into a replayable
Recording. Demo: uv run python examples/sdk_capture.py.
Replay + trajectory diff
Re-drive an agent run against a different model, injecting recorded tool results (the engine cannot call a live tool — the guarantee holds by construction), then diff the two decision trajectories to find the first point of divergence.
# Offline, no creds — two scripted "models" over the same recorded results:
uv run python examples/replay_scripted.py
# → diverge at decision 0: haiku-4.5 → call get_weather(...) | opus-4.8 → call get_forecast(...)
The model client is pluggable behind one adapter (replay/adapters/): ScriptedAdapter
(offline/tests), BedrockAdapter (Converse, dev), OpenAICompatAdapter (OSS release). The
engine speaks only provider-neutral types — swapping providers never touches it. Recorded
tool results can be supplied directly or pulled from a captured proxy trace via
Recording.from_trace_store(...).
Capture → replay (the full loop)
An SDK-captured run replays end-to-end: Recording.from_sdk_trace(store, trace_id)
reconstructs the question, tools, model, and recorded tool results, and
captured_trajectory(store, trace_id) reconstructs what the agent actually did. Diff
the two to answer "would a different model have decided differently than my agent did?"
uv run python examples/capture_then_replay.py
# → diverge at decision 0: claude-haiku-4-5 → call get_weather(...) | rushed-model → final answer
Model access (live replay only — not needed for the proxy)
Live replay uses AWS Bedrock (Converse API) for dev, OpenAI-compatible for the OSS release. Before running the Bedrock example:
aws sso login --profile <your-aws-profile> # temporary SSO creds
export AWS_PROFILE=<your-aws-profile> AWS_REGION=eu-west-1
uv run --extra replay python examples/replay_bedrock.py
Contributing
Issues and PRs welcome — see CONTRIBUTING.md. All tests run offline
(uv run pytest); the proxy has no model dependency.
License
Apache License 2.0 — includes an explicit patent grant.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file agentsense_ai-0.1.0.tar.gz.
File metadata
- Download URL: agentsense_ai-0.1.0.tar.gz
- Upload date:
- Size: 290.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
71566f94dd2ff76ed6179a3b015ca49bb8978354619c1bb4b479a364f0d96216
|
|
| MD5 |
3e76e963cfe70fd90fe03b4e3e73f956
|
|
| BLAKE2b-256 |
ec292062cdaf118ddd9d4d297a76eae632434686c1c10466f0f34e32e1cc9187
|
File details
Details for the file agentsense_ai-0.1.0-py3-none-any.whl.
File metadata
- Download URL: agentsense_ai-0.1.0-py3-none-any.whl
- Upload date:
- Size: 45.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
db0c2b491ad7d317f6f27459714db85f4d450251eee74b6db3fb6bdd3d91f550
|
|
| MD5 |
e22aab4a46313df939dcb4ce10ac272d
|
|
| BLAKE2b-256 |
7f96f419802ffee41462c9a967ccd6feb5d841a4865b22cb2e5f3e1a65b3f1a0
|