Skip to main content

traceback-ai

LLM Agent Execution Tracer with Failure Attribution — strace for LLM pipelines.

Your LLM agent failed. The final answer was hallucinated, truncated, or incomplete, but your pipeline ran five tool calls, two retrieval steps, and three prompt transforms. Which step actually caused the failure?

traceback-ai instruments any LLM or agent pipeline, records execution spans in local SQLite storage, evaluates step-level health metrics, and deterministically attributes root-cause failures to the exact offending step.


📦 Installation

pip install agent-blame

# Optional extras:
pip install "agent-blame[semantic]"    # Sentence-transformers semantic embeddings
pip install "agent-blame[gemini]"      # Google Gemini SDK instrumentation
pip install "agent-blame[anthropic]"   # Anthropic SDK instrumentation
pip install "agent-blame[openai]"      # OpenAI SDK instrumentation
pip install "agent-blame[langchain]"   # LangChain callbacks
pip install "agent-blame[all]"         # All integrations & extras

⚡ Quickstart

Instrument your functions with @trace. Decorate your pipeline entrypoint with @trace(pipeline=True):

from tracebackai import trace

@trace(step_type="retrieval")
def retrieve_docs(query: str) -> list[str]:
    # Returns relevant passages from your database or index
    return ["Retrieval-augmented generation combines search with LLMs."]

@trace(step_type="prompt")
def build_prompt(query: str, docs: list[str]) -> str:
    return f"Context:\n{chr(10).join(docs)}\n\nQuestion: {query}"

@trace(step_type="llm")
def generate_answer(prompt: str) -> str:
    # Call Claude, GPT-4, or any custom model
    return "RAG combines search retrieval with generative language models."

@trace(pipeline=True)
def answer_pipeline(query: str) -> str:
    docs = retrieve_docs(query)
    prompt = build_prompt(query, docs)
    return generate_answer(prompt)

if __name__ == "__main__":
    answer_pipeline("What is RAG?")

🖥️ CLI Inspection & Failure Attribution

Every execution is persisted to ~/.traceback/traces.db (configurable via TRACEBACK_DB_PATH).

1. Inspect Execution Spans (traceback show)

$ traceback show abc123def

Run: abc123def  |  Pipeline: answer_pipeline  |  2026-08-25 14:32:01
──────────────────────────────────────────────────────────────────────
[0] retrieve_docs      retrieval    312ms  tokens=180   score=0.42 ⚠ WEAK RETRIEVAL
    input:  "What is RAG?"
    output: ["Baking sourdough bread requires wild yeast..."]
[1] build_prompt       prompt       8ms    tokens=1843
[2] generate_answer    llm          941ms  tokens=2011  cost=$0.003  score=0.91 ✓
──────────────────────────────────────────────────────────────────────
Total: 1261ms  |  Cost: $0.0030  |  Final Output: "Sourdough bread..."

2. Attribute Root-Cause Failure (traceback blame)

$ traceback blame abc123def

Analyzing run abc123def (answer_pipeline, 3 steps)...

🔴 BLAME: Step [0] retrieve_docs  (retrieval)
   Score:       0.42  (threshold: 0.55)
   Blame score: 0.81  (high confidence)
   Reason:      Retrieval relevance score was 0.42 (below threshold 0.55).
                Top chunk similarity was 0.42. Downstream steps received low-quality context.

Co-blame: none
Other steps: generate_answer (0.91 ✓), build_prompt (unscored)

3. Compare Pipeline Runs (traceback diff)

$ traceback diff abc123def def456ghi

Comparing abc123def → def456ghi
Pipeline: answer_pipeline

STEP                 SCORE_A    SCORE_B    DELTA      STATUS
retrieve_docs        0.42       0.88       +0.46      ↑ improved
generate_answer      0.91       0.48       -0.43      ↓ REGRESSED ← highest delta

Verdict: REGRESSION in generate_answer
Details: Significant quality regression in 'generate_answer' (delta: -0.43).

🔬 How Scoring Works

Before traces are written to SQLite, each step is evaluated by registered step scorers:

  • Retrieval Scorer (step_type="retrieval"): Measures semantic cosine similarity between the query input and retrieved chunks using sentence-transformers (all-MiniLM-L6-v2), clamped to [0.0, 1.0]. Automatically falls back to pure-Python BM25 term overlap if dependencies are absent. Uses method-aware thresholding (flags weak retrieval when score $< 0.55$ for semantic embeddings, or $< 0.33$ for BM25 term overlap).
  • LLM Scorer (step_type="llm"): Evaluates a composite of response completeness (scaled against an absolute floor of 20 tokens so concise correct answers to large RAG prompts are never penalized), refusal detection (detects refusal strings like "I cannot", "As an AI" and zeroes the score), and self-consistency (pairwise ROUGE-L across multi-sample completions).
  • Tool Scorer (step_type="tool"): Evaluates exceptions, validates non-empty payloads, validates against optional expected_type metadata, and dynamically penalizes based on historical failure rates in SQLite.

🎯 How Blame Attribution Works

traceback blame ranks candidate failure steps deterministically without expensive or recursive LLM calls:

$$\text{blame_score}(\text{step}) = (1.0 - \text{step.score}) \times \text{weight}(\text{step_type}) \times \text{recency_weight}(\text{index}, \text{total})$$

  • Type Weights: retrieval: 1.4 (retrieval failures compound downstream), llm: 1.2, tool: 1.1, prompt: 0.9, generic: 0.8.
  • Recency Multiplier: $1.0 + 0.3 \times \left(1.0 - \frac{\text{index}}{\text{total}}\right)$, giving higher impact to upstream steps.
  • Unscored Step Safety: Generic steps (score is None) are excluded from blame candidacy so they never corrupt attribution. If all steps are unscored, blame identifies the slowest execution bottleneck.

📊 Benchmark Accuracy

traceback-ai is continuously evaluated across 19 realistic failure and healthy scenarios spanning retrieval, LLM, tool, conversational BM25, and cascading degradations (benchmarks/blame_accuracy.py):

  • Top-1 Attribution Accuracy: 100.0% (14/14 failure scenarios correctly attributed)
  • False-Positive Rate: 0/4 (0.0%) on healthy traces (blame score remains $< 0.45$)
  • Benchmark Runtime: < 0.2s (100% offline, zero network access, zero external API keys)
Category Correct Total Accuracy
retrieval 3 3 100.0%
llm 4 4 100.0%
tool 3 3 100.0%
cascading 2 2 100.0%
fallback 2 2 100.0%

🚀 CI Eval Gate Integration

Add automated failure-attribution gates to pull requests in GitHub Actions:

name: Eval Gate
on: [pull_request]

jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.11" }
      - run: pip install -e .
      - run: traceback run examples/simple_rag.py --input examples/test_cases.json --fail-on-blame 0.7

🔌 SDK Integrations

Integration Usage Traced Metrics
Google Gemini from tracebackai.integrations.gemini import TracedGemini model, input_tokens, output_tokens, latency_ms
Anthropic from tracebackai.integrations.anthropic import TracedAnthropic, patch_anthropic model, input_tokens, output_tokens, stop_reason
OpenAI from tracebackai.integrations.openai import TracedOpenAI, patch_openai model, input_tokens, output_tokens, finish_reason
LangChain from tracebackai.integrations.langchain import TracebackCallbackHandler on_llm_*, on_retriever_*, on_tool_*

🗺️ Roadmap

  • Phase 1: Core tracer, SQLite persistence, @trace, Click CLI (list, show)
  • Phase 2: Retrieval (cosine/BM25), LLM, and Tool scorers
  • Phase 3: Blame attribution algorithm, cross-run diffing, explanation generator
  • Phase 4: Anthropic / OpenAI / LangChain integrations, CI eval gates, PyPI packaging
  • Post-Ship: Local Web UI dashboard (traceback serve), Async tracer support, OpenTelemetry export

📄 License & Contributing

Licensed under the MIT License. See CONTRIBUTING.md for local development guidelines.

Metadata

Release files for agent-blame 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-blame 0.1.1
File Size Uploaded
agent_blame-0.1.1.tar.gz 79.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-blame 0.1.1
File Interpreter ABI Platform
agent_blame-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 136.0 kB

Release files / agent_blame-0.1.1.tar.gz

Download URL agent_blame-0.1.1.tar.gz
Size 79.1 kB
Tags Source
SHA-256 checksum
How to use checksums
07e31c864cdc2b4e8cc05695c1ac4786ba30590f7c4ba6fb56176813054ad4c1
BLAKE2b-256 checksum
How to use checksums
02ae0d5a75c2ac61b5575becd9d77d03ad4c21bd27f918083b0b684141581e88
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / agent_blame-0.1.1-py3-none-any.whl

Download URL agent_blame-0.1.1-py3-none-any.whl
Size 56.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
74f33af6c35e9a28ba4c8ede66aeccc9ad8086981b73fc57d270cb19fce98d27
BLAKE2b-256 checksum
How to use checksums
bd52a7f857b361f88703eeb4b18f964f34935d6de0734ab12a9b092e338d1feb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page