Skip to main content

Local-first evidence-chain debugger for RAG and AI agent claim grounding, citation checks, root-cause diagnosis, and regression tests.

Project description

ContextTrace

Local-first evidence-chain forensics for RAG and AI agents.

ContextTrace is a Python SDK and CLI for tracing a failed answer from the user query through retrieved context, answer claims, citations, verdicts, root cause, repair guidance, and CI regression tests.

query -> retrieved context -> answer claims -> citations -> verdicts -> root cause -> regression test

Use it when a RAG or agent score is not enough: ContextTrace points at the unsupported or contradicted claim, the evidence span and citation involved, why the failure likely happened, and how to keep it from coming back. It is not a hosted dashboard. Traces, reports, judge cache, and SQLite state stay local by default.

Install

pip install contexttrace
contexttrace init

Quickstart

contexttrace verify-demo unsupported_claim --report
contexttrace demo --dataset refund_policy
contexttrace report --last --open

Default local storage:

.contexttrace/contexttrace.db

Verify A RAG Trace

Create a portable trace with a query, answer, retrieved contexts, and optional citations:

{
  "query": "How long does refund processing take?",
  "answer": "Refunds are processed within 5 business days.",
  "contexts": [
    {
      "id": "policy",
      "text": "Customers may request refunds within 30 days of purchase."
    }
  ]
}

Run local evidence checks:

contexttrace inspect trace.json
contexttrace verify trace.json --report
contexttrace diagnose trace.json --report
contexttrace qa trace.json --corpus docs/ --report
contexttrace repair trace.json --corpus docs/ --out repair_plan.md

ContextTrace classifies each claim as supported, partially_supported, unsupported, unverifiable, or contradicted, then exposes separate statuses for support, truth, source freshness, citation quality, and likely fix.

Important: supported means grounded by the selected evidence span. It does not mean independently true, current, or authoritative.

Diagnose An Agent Trace

diagnose also accepts agent step traces and localizes tool/final-answer failures:

{
  "goal": "Book a meeting with Alex",
  "steps": [
    {
      "type": "tool_call",
      "tool": "calendar.search",
      "args": {"date": "Friday"},
      "result": "No availability"
    },
    {
      "type": "final_answer",
      "content": "I booked it for Friday."
    }
  ]
}
contexttrace diagnose examples/diagnose_agent_trace.json --report --fail-on high_risk

The diagnosis flags tool_result_contradicted_by_final_answer and suggests gating final-answer generation on tool-result status.

Turn that diagnosis into a CI regression test:

contexttrace diagnose examples/diagnose_agent_trace.json \
  --generate-test \
  --test-out tests/contexttrace/test_calendar_agent_diagnosis.py

pytest tests/contexttrace/test_calendar_agent_diagnosis.py

Build A Repair Plan

repair turns diagnosis into an evidence-backed implementation plan. With a local corpus, it distinguishes retrieval miss, reranking failure, chunking issue, corpus gap, answer overreach, and stale or conflicting evidence:

contexttrace repair trace.json \
  --corpus docs/ \
  --out repair_plan.md \
  --json-out repair_plan.json

The plan records the failed claim, retrieved and corpus evidence, prioritized root-cause-specific changes, and commands to verify the fix. Add only the recaptured passing trace to the generated must-pass regression command.

Local Verification Modes

Mode Use When
lexical Fast default checks with no optional dependencies.
semantic Local paraphrase and role-aware contradiction checks.
local_ml Offline hash-embedding similarity, optionally backed by a local SentenceTransformers model.
nli Local claim+span entailment or contradiction with a local Transformers or ONNX NLI model.
judge Higher-accuracy local LLM judging through Ollama, LM Studio, vLLM, or a local OpenAI-compatible server. The judge sees selected evidence spans, not the full answer prose.

Run the stronger local non-LLM verifier:

contexttrace verify trace.json --mode local_ml --report
contexttrace verify-benchmark --mode local_ml --case-set all

Optional neural local-ML support never downloads models automatically:

pip install "contexttrace[local-ml]"
set CONTEXTTRACE_LOCAL_ML_MODEL_PATH=C:/models/bge-small-en-v1.5

Run local NLI when you want mechanical claim-versus-span entailment:

pip install "contexttrace[nli]"
set CONTEXTTRACE_NLI_MODEL_PATH=C:/models/deberta-v3-nli
contexttrace verify trace.json --mode nli --report
contexttrace nli-calibrate --case-set all --report

Run a local judge with Ollama:

set CONTEXTTRACE_JUDGE_PROVIDER=ollama
set CONTEXTTRACE_JUDGE_MODEL=llama3.1

contexttrace verify trace.json --mode judge --report
contexttrace judge-calibrate --case-set all --report

Remote judges are blocked while local_only: true is active. To use a remote judge, explicitly disable local-only mode and configure the provider/API key.

Diagnose And Regression-Test

# Find whether support existed elsewhere in the corpus.
contexttrace audit trace.json --corpus docs/ --report

# Compare a baseline and current answer after a prompt, model, or retriever change.
contexttrace compare baseline.json current.json --report

# Turn saved failures into replayable endpoint tests.
contexttrace suite create traces/failure.json --out contexttrace-suite.json
contexttrace suite run contexttrace-suite.json --endpoint http://localhost:8000/query --report

Common root causes include retrieval_miss, reranking_failure, chunking_issue, corpus_gap, answer_overreach, stale_source, citation_mismatch, and should_have_abstained.

support_status, truth_status, and source_status stay separate so a claim can be grounded by a source while the source itself remains stale, wrong, or unassessed.

Source metadata can include source_authority, source_timestamp, source_version, canonical, or canonical_source. ContextTrace uses those local fields to flag grounded_but_stale, grounded_but_conflicted, grounded_by_low_authority_source, or supported_by_canonical_source.

Capture Existing Systems

Capture one live endpoint response:

contexttrace capture endpoint \
  --endpoint http://localhost:8000/query \
  --query "What is the refund policy?" \
  --answer-path $.answer \
  --contexts-path $.contexts \
  --citations-path $.citations \
  --out traces/refund_trace.json \
  --verify \
  --report

Or capture artifacts from Python:

from contexttrace import capture_rag_trace, write_rag_trace

trace = capture_rag_trace(
    query=question,
    answer=answer,
    contexts=retrieved_docs,
    metadata={"system": "support-rag"},
)
write_rag_trace(trace, "trace.json")

SDK Example

from contexttrace import ContextTrace

ct = ContextTrace(project="support-rag")

with ct.trace(query="What is the refund policy?") as trace:
    chunks = retriever.search("What is the refund policy?")
    trace.log_retrieval(chunks)
    trace.log_context(chunks[:5])

    answer = llm.generate("What is the refund policy?", chunks[:5])
    trace.log_answer(answer, usage={"total_tokens": 1200})
    trace.log_citations([
        {"claim": "Refunds are available within 30 days.", "source_chunk_id": "chunk_12"}
    ])

    result = trace.evaluate()
    print(result["failure"]["failure_type"])

Integrations

pip install "contexttrace[langchain]"
pip install "contexttrace[llamaindex]"
pip install "contexttrace[fastapi]"
pip install "contexttrace[langgraph]"
pip install "contexttrace[otel]"
pip install "contexttrace[all]"

Includes LangChain, LlamaIndex, FastAPI, LangGraph, and OpenTelemetry hooks.

Privacy

ContextTrace makes no network calls unless you point it at an endpoint or configure a judge provider. Local controls include:

  • local_only: true
  • log_chunk_text: false
  • log_answer_text: false
  • storage_path
  • judge_cache_enabled: true
  • judge_cache_path: .contexttrace/judge_cache.json

Limits

ContextTrace is a diagnostic tool, not a correctness proof. It verifies grounding against provided evidence; it does not certify real-world truth. Claim extraction is rule-based, contradiction detection is conservative, and high-stakes outputs still need human review.

Links

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

contexttrace-1.1.0.tar.gz (244.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

contexttrace-1.1.0-py3-none-any.whl (280.2 kB view details)

Uploaded Python 3

File details

Details for the file contexttrace-1.1.0.tar.gz.

File metadata

  • Download URL: contexttrace-1.1.0.tar.gz
  • Upload date:
  • Size: 244.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for contexttrace-1.1.0.tar.gz
Algorithm Hash digest
SHA256 d105c578db9c32a609792a1d913664b698ba4b7cf25f3d0c34f51f8e389617f6
MD5 7b29eed4d2877735266294a6a10c34b4
BLAKE2b-256 02bdda36001671b92dd1d20efbca507b6c2fcc980cd95fb51f473596ad10a6f8

See more details on using hashes here.

Provenance

The following attestation bundles were made for contexttrace-1.1.0.tar.gz:

Publisher: release.yml on samarth1412/Context-Trace

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file contexttrace-1.1.0-py3-none-any.whl.

File metadata

  • Download URL: contexttrace-1.1.0-py3-none-any.whl
  • Upload date:
  • Size: 280.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for contexttrace-1.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 c1d29a74b96f99896e2923d0dce454bb8ebdd298f5193eac22dd9ae99e17f85d
MD5 b7452c7d28d5942f369262709b44a964
BLAKE2b-256 570ae421b0e38ec37461420d0c8d0318982c3e651f33fd94fe17d386137cedd7

See more details on using hashes here.

Provenance

The following attestation bundles were made for contexttrace-1.1.0-py3-none-any.whl:

Publisher: release.yml on samarth1412/Context-Trace

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page