Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

ContextTrail

test

Turns the Codex and Claude Code sessions already on your machine into an evidence-linked history of a project: what was tried, what failed, what was decided, and what was actually verified.

ContextTrail reads the transcripts the two CLIs keep locally, plus the project's Git history, and asks the CLI you already have installed (sandboxed, read-only) to reconstruct the flow. Every event and every arrow carries a quote from the source, and the code checks each quote against the record before anything is stored. You get a terminal view, a local browser view, find/show commands, and two agent skills so Codex and Claude Code can answer "why did we drop X?" from the saved graph instead of from memory.

Development alpha 0.1.0a5. Linux and macOS, Python 3.11+. What has been measured is further down; nothing on this page is an estimate presented as a result.

The terminal view on the synthetic demo: the event flow on top, the selected event with its links and quoted evidence below

Try it in a minute, without an AI account

git clone https://github.com/kisbi3/ContextTrail.git && cd ContextTrail && ./install.sh
contexttrail demo --path /tmp/contexttrail-demo        # synthetic logs + a deterministic mock runner

The demo writes a small Codex log and a small Claude Code log for one folder, runs the whole pipeline with a mock in place of the model, and opens the result. It shows what the tool builds and how it reads; it says nothing about real model accuracy.

What goes in, what comes out

The demo's input is six records in two logs, the kind of thing the CLIs write all day:

Source Role Record
Codex session user Let's store it as a JSON file for now.
Codex session assistant SQLite is also worth considering.
Codex session tool result Patch applied: JSON storage code.
Codex session tool result concurrent write test: FAILED — JSONDecodeError
Claude Code session user Concurrent writes break it, so let's switch to SQLite.
Claude Code session assistant Switched to SQLite and the tests pass too.

Out of that, the pipeline (extract events with quotes → integrate them into the existing graph → check every claim in code) publishes six events and seven relations. The terminal shows them as a flow:

 ┌─ [01] decision ─────────────────┐
 │ JSON file storage adopted       │
 │ · adopted                       │
 └─────────────────────────────────┘
    │ follows
    ▼
 ┌─ [02] proposal ─────────────────┐
 │ SQLite proposed as an           │
 │ alternative                     │
 │ · proposed                      │
 └─────────────────────────────────┘

 ┌─ [03] change ───────────────────┐      ◀ [01] follows
 │ JSON storage implemented        │
 │ ✗ applied                       │
 │   · check failed: Concurrent    │
 │     write test failed           │
 └─────────────────────────────────┘
    │ verifies
    ▼
 ┌─ [04] result ───────────────────┐
 │ Concurrent write test failed    │
 │ ✗ observed failure              │
 └─────────────────────────────────┘
    │ motivates
    ▼
 ┌─ [05] fix ──────────────────────┐      ◀ [01] revises   ◀ [02] follows
 │ Switch to SQLite decided        │
 │ · adopted                       │
 └─────────────────────────────────┘
    │ follows
    ▼
 ┌─ [06] result ───────────────────┐
 │ SQLite change reported done     │
 │ ! reported done·unverified      │
 └─────────────────────────────────┘

Three things in that picture are the point of the tool:

  • A change never gets a success badge of its own. Event 03 is applied, and it is marked ✗ only because a run that verifies it failed. Verification is an arrow to a result, never a word in a summary.
  • Words are not results. Event 06, "Switched to SQLite and the tests pass too", is reported done·unverified (!) because no test output in the records backs it. The flag goes away only when a run result is linked.
  • Every arrow has a basis. motivates from 04 to 05 is explicit: the person said "Concurrent writes break it". Links the model merely infers are drawn dotted and labelled, or left out.

Any event can be opened with its quotes, in the terminal or by an agent:

$ contexttrail show ev_53ee4e0e
ContextTrail event · current graph v2 · saved result · no AI calls
ev_53ee4e0e  JSON storage implemented
  status: applied · check failed: Concurrent write test failed
  change · tool · 2026-09-22 19:02
  reference: contexttrail:ev_53ee4e0e@v2

What happened
  Patch applied: JSON storage code.

Linked events
  ← follows  ev_bedb4f79  JSON file storage adopted [adopted]
  → verifies  ev_eb72c0aa  Concurrent write test failed [observed failure]

--- source evidence 1 (Codex · tool result · 2026-09-22 19:02 · session demo-cod · line 1) ---
> Patch applied: JSON storage code.
--- end ---

The contexttrail:ev_…@v2 reference can be pasted into a Codex or Claude Code chat, and the installed contexttrail-context skill reads that event with its evidence.

The same thing on a real project

On a real project the records are thousands of lines of conversation, tool calls, diffs and test output, and the model does the reading; the checks are the same. This is the tail of what Claude Sonnet reconstructed from three work units of this repository's own Codex logs (the afternoon macOS sandbox support was added), in 8 calls, about 1.5 minutes and no repair round:

 ┌─ [19] change ─────────────────┐               ┌─ [20] result ─────────────────┐
 │ cli_runner.py gets macOS      ├────verifies──▶│ test_runner: 1 failed, 11     │
 │ Seatbelt sandbox              │               │ passed; codex doctor fails    │
 │ ✗ applied                     │               │ ✗ observed failure            │
 │   · check failed:             │               └───────────────────────────────┘
 │     test_runner: 1 failed, 11 │
 │     passed; codex doctor…     │
 └───────────────────────────────┘

 ┌─ [21] fix ────────────────────┐               ┌─ [22] result ─────────────────┐
 │ Resolve home path and add     ├────verifies──▶│ codex doctor passes on        │
 │ /etc,/var,/tmp literals       │               │ Seatbelt                      │
 │ ✓ applied                     │               │ ✓ observed success            │
 │   · verified: codex doctor    │               └───────────────────────────────┘
 │     passes on Seatbelt        │
 └───────────────────────────────┘

Opening the fix shows why it is marked verified: the patch itself, the error it answered, and the run that checked it are all quoted from the log.

$ contexttrail show ev_cf348e54
ev_cf348e54  Resolve home path and add /etc,/var,/tmp literals
  status: applied · verified: codex doctor passes on Seatbelt
  fix · assistant · 2026-09-23 16:52

What happened
  Two follow-up patches to cli_runner.py: resolved home/TMPDIR paths after CODEX_HOME read
  failures, then allowed literal traversal of /etc, /var, /tmp after codex failed to read
  /etc/codex/requirements.toml.

Linked events
  → verifies  ev_233a98c0  codex doctor passes on Seatbelt [observed success]

--- source evidence 1 (Codex · tool call · 2026-09-23 16:52 · line 2) ---
> - return home
> + return home.resolve()
--- source evidence 2 (Codex · tool call · 2026-09-23 16:52 · line 2) ---
> + literals = sorted(ancestors | {str(credential), "/etc", "/var", "/tmp"})
--- source evidence 3 (Codex · tool result · 2026-09-23 16:52 · line 5) ---
> Failed to read requirements file /etc/codex/requirements.toml: Operation not permitted (os error 1)
--- end ---

Earlier in the same flow a change is applied · unverified because its tests were never run in the log, a request · no answer marks a question the assistant never answered, and an open item records that a pytest run was aborted with no result. Those are the states the tool is for.

What it does and does not do

  • Reads only. It never modifies the project, the transcripts, or Git refs, index or config. Its own state lives under .git/contexttrail/ (or .contexttrail/ without Git).
  • Calls a model only when you ask. Viewing, exporting, the browser page and node selection make zero calls. Analysis runs on analyze, on R in the terminal view, or when an agent runs the contexttrail-update skill with a unit count you gave it; before sending it shows what it will send and asks.
  • Runs your CLI in a sandbox. The Codex or Claude CLI runs under bubblewrap (Linux) or a deny-default sandbox-exec profile (macOS) with its credential file bind-mounted read-only. If the sandbox is unavailable it refuses to run; there is no unsandboxed fallback, and it never reads or copies credential contents.
  • Sends only this project's records. A record is attributed to the project by the working directory the CLI recorded, never by keyword guessing. Out-of-scope conversations never reach the model.
  • Checks before it stores. Every quote must exist in the source at the cited lines. Slips with one possible reading (a quote at the wrong line, an extra escape) are corrected in code and audited; everything else goes back to the model for one repair round, and what still fails is dropped and named in the graph's limitations.
  • Costs tokens from your own account. The pipeline is tuned so a cheap model does the work: on the measured fixture, Claude Sonnet finishes a work unit in about a minute. The plan shown before a run estimates calls, tokens and minutes.
  • Is incremental. Finished work units are never re-sent. Re-running with no new records makes zero calls.

Version: 0.1.0a5 · Linux / SSH primary, macOS measured · names are provisional.

Languages: the terminal, TUI, browser view and CLI help are in English or Korean, chosen from --language, CONTEXTTRAIL_LANGUAGE, the project's saved output language, or the locale. Event titles and summaries are written in the language of your own messages (--language overrides). Everything the model reads is English. A Korean README is not written yet.

More: model tiering, project filters, and test usage. Before opening the UI on a real project, run contexttrail scan . to review input selection.


1. Getting Started

Requires Python 3.11+, Git, and a UTF-8 terminal. Python's curses module and venv support are needed. On Linux or macOS, clone the repository and run once from the repo root:

cd /path/to/ContextTrail
./install.sh
contexttrail --version

install.sh installs the package into ~/.local/share/contexttrail/venv, links ~/.local/bin/contexttrail and (where possible) the project command, and registers Codex/Claude Code agent commands. Existing user command files are preserved; files created by ContextTrail are updated on reinstall. Package repository access may be required. If ~/.local/bin is not in PATH, follow the instructions printed during installation. Running pip install or pip install git+... directly does not register agent commands — run contexttrail install-commands afterward.

After installation, running contexttrail with no arguments in a project directory opens its saved flow view — equivalent to contexttrail view .. AI is called only when you press R in the TUI. To open a different directory, pass it as the first argument: contexttrail /path/to/project.

cd /path/to/project
contexttrail

In a non-Git directory, state is stored in .contexttrail/. In a Git repository, it goes under .git/contexttrail/.

Code was tested on Python 3.13.5. Python 3.11 and 3.12 compatibility is pending verification. The command is contexttrail, with ct as a short alias; python -m contexttrail also works.

Running the demo without AI or credentials

contexttrail demo --path /tmp/contexttrail-demo

Specify a new or empty directory. This generates synthetic Codex/Claude logs and a demo project without calling any external model — it uses a test-only Mock Runner that handles a fixed set of cases. This Mock cannot be selected as the analyzer for a real project.

Example flow:

[01] Adopt JSON storage / adopted
├── follow → [02] Propose SQLite alternative / proposed
│   └── follow → [05] Decide to switch to SQLite / adopted
│       └── follow → [06] Report SQLite migration done / report·unverified
└── follow → [03] Implement JSON storage / report·unverified
    └── follow → [04] Concurrent-write test fails / observed failure
        └── motivates → ↗ [05] (merges/returns)

Viewing saved results and re-running without changes:

contexttrail view /tmp/contexttrail-demo/sample-project
contexttrail analyze /tmp/contexttrail-demo/sample-project --no-tui
# noop if no new records — 0 runner calls

2. Applying to a Real Project

Setting Up a Runner and Diagnosing

Install and log in to the Codex or Claude Code CLI yourself before use. ContextTrail does not handle token issuance, copying, or login. On Linux, real AI execution requires bubblewrap (bwrap) and available user namespaces. On macOS, sandbox-exec and file-based CLI credentials are required. If isolation tools or required CLI options are missing, execution is blocked. The macOS path uses a deny-default Seatbelt profile with a temporary HOME; it has passed Codex doctor --smoke and a small live-segment analysis. Blocking of malicious tools and config is not yet validated.

contexttrail doctor --runner codex
contexttrail doctor --runner claude

The default doctor run checks CLI version, help output, and isolation feasibility — no model calls. The presence of a credential file does not guarantee successful authentication. The following commands make explicit real-account calls, sending a small synthetic input (not your project data):

contexttrail doctor --runner codex --smoke --yes
contexttrail doctor --runner claude --smoke --yes

This smoke test is a simple structured JSON round-trip. It does not substitute for full integration validation of mixed-log analysis, all ReadRequests, or permission blocking. Per-CLI flags and unvalidated items are in Implementation Status; isolation constraints are in Security.

Current alpha authentication constraints:

  • Only file-based credentials are supported: Codex's auth.json and Claude's .credentials.json are bind-mounted read-only into the CLI's isolated environment. The program does not read or copy their contents.
  • Keyring-only auth, per-org user settings, custom API providers, proxies, and certificate inheritance are not supported.
  • If credential renewal requires a write, the run may fail. Log in or refresh via the original CLI, then retry. Permissions are not escalated to work around failures.

Confirming Input Scope → Analyzing

# Check relevant record count, linked worktrees, and limitations — no AI calls
contexttrail scan /path/to/project

# Read logs from both sources and analyze with one chosen runner
contexttrail analyze /path/to/project --runner codex
# or
contexttrail analyze /path/to/project --runner claude

Replace /path/to/project with the actual path on your server. The first analysis prompts for transmission consent. Conversation turns, tool outputs, and code fragments may be sent to the model service of the chosen CLI. This is not an automatic secret redactor — verify your project's transmission policy before proceeding with sensitive material.

The chosen runner, settings, and consent are stored in the per-scope local DB. Analysis proceeds oldest-first through work units. Each run first shows a summary such as "1,480 units pending · 15 this run (≤30 AI calls) · ~2.39M input tokens · ~120 min (estimated)" along with token/time projections for 5, 15, and 30 units, then prompts for the number of units to process (Enter = shown count, n = cancel, --yes = no prompt). Use --units N to specify in advance. Once units are set, the per-unit AI call cap is 6; --max-calls overrides that. Both values apply only to the current run and are not saved. Estimates are calibrated from actual input tokens and timing if ≥3 units with usage records exist. Use contexttrail scan to preview the plan without sending anything (plan_text, plan_choices).

Output language. Event titles and summaries are written in one language per project, determined by the dominant language of user messages during the first analysis run. Script-distinct languages (Korean, Japanese, etc.) are identified directly; for Latin-script languages (English, Spanish, etc.) the system locale is used as a tiebreaker. Source text in other languages is still summarized in the chosen language. Code names, filenames, commands, and quotations remain verbatim. Use --language Korean to change; --language auto re-detects. Already-analyzed events are not re-translated unless re-analyzed.

Speed. While integrating one work unit, extraction for the next unit begins in parallel (with one extraction worker). Token cost is identical; only wall-clock time is reduced. For read-only tool calls (file reads, listings, searches), only the first 400 and last 200 characters are sent; the model may request the rest via a ReadRequest.

--session <session-id> prioritizes that session and its sub-agents over older records. --session current refers to the Codex/Claude Code conversation that invoked this command. Events added this way are marked "out-of-order analysis" — relationships to earlier unanalyzed records may be missing. Use contexttrail analyze /path/to/project for subsequent runs. Runner does not switch silently; completed segments are not re-analyzed just because settings changed.

Analysis model and reasoning level. ContextTrail always specifies the model by name because the isolation environment does not carry the user's CLI config file (~/.codex/config.toml, etc.). Defaults: Codex gpt-6-sol, Claude sonnet; override with --model. Per-stage reasoning level defaults to medium for all stages (lowered from high for integrate/re-review on 2026-09-27). Adjust with --extract-effort, --integrate-effort, --escalation-effort (low·medium·high·xhigh·max). Every call logs the requested model and reasoning level in the local call ledger (contexttrail ops --details) and evaluation reports. Claude records the responding model; Codex does not report it in its response.

Change re-review. If integration results contain unlinked execution results, unlinked fixes, or modifications to existing events/relationships, the same runner reviews the proposed changes one more time. This adds one call to that work unit. Use --no-review to skip for the current run only.

For non-interactive use, consent can be given with explicit --yes. The following command goes beyond a simple example and allows real transmission — confirm scope first:

contexttrail analyze /path/to/project --runner codex --yes --no-tui

Default log paths: CODEX_HOME or ~/.codex (sessions/, archived_sessions/) and CLAUDE_CONFIG_DIR or ~/.claude (projects/). Override when paths differ per server:

contexttrail scan /path/to/project \
  --codex-home /path/to/codex-home \
  --claude-home /path/to/claude-home

Only the project currently being worked on is analyzed. If logs are on a local PC and the project is on a Linux server, the server does not automatically collect logs from the PC. Cross-machine/repo auto-merge is not a feature of this version.

Extraction input contains the source text for the work unit plus limited surrounding context. Up to 4 conversation/Git records from the same worktree within 15 minutes are placed at the front of the surrounding context and additional read list. Temporal proximity alone does not imply causation. Long Git diffs are split into citable fragments of ≤32,000 characters, the same as logs. Commits whose Git command output exceeds the 4 MB safety limit are deferred while the next commit continues — check contexttrail scan for limitations.

Work units are first grouped by session and worktree. Within a session, summary compression, gaps ≥3 hours, and date changes with sufficient gaps are used as boundaries. When the input budget is reached, breaks are made at user-turn boundaries where possible; tool calls/results and source fragments are kept together. Work unit boundaries do not correspond to event node boundaries in the final graph. Already-analyzed segments are not re-chunked in incremental analysis.

To preview a small fixed evaluation fixture before model calls, run contexttrail eval --fixture /tmp/case.json --preview --output /tmp/case-preview. The output input-preview.html shows WorkUnit boundaries and source text, the common system instruction, the extraction instruction, new_records/context_only, existing events, evidence, manifest, the response JSON Schema, the full task JSON, and the generation path of each part. The first unit's input is the exact initial request; subsequent units' existing-graph sections are estimates that depend on earlier model responses. Preview includes private source text; no AI calls are made.

After contexttrail eval, review.html is generated automatically. It shows event/relationship candidates, citation strings vs. actual source lines, and validation errors for each model call side by side. Eval runs also save failed structured responses to call-review/ with private user permissions. You can regenerate HTML for earlier evals with contexttrail review /path/to/eval-output, but source model response text not saved at the time cannot be recovered. This local review view makes no AI calls or external transmissions.


3. Using the Terminal UI

contexttrail view, contexttrail analyze, and contexttrail graph share the same screen. When user messages exist in the graph, the left panel shows a request list; the center event flow panel renders only the span of the selected request — from that message to the next message in the same session. Use [/] to navigate requests, and A to toggle between full-flow and per-request views. Connections to out-of-span events are annotated inline as ← [02] motivates. Events before any recorded request are grouped under "Events without request record."

The event flow panel draws events as boxes with labeled arrows. The left column lists events in record order; execution results that verified a change but do not continue elsewhere are attached to the right of that change as ──verified──▶. Results that lead to the next story beat (e.g. a failure motivating a fix) remain in the column, so failure → motivates → fix reads downward. Adjacent boxes connect with downward arrows; distant boxes connect via a margin line. The margin line notes the source box number and relationship as [02] motivates ▶ before the destination, so the origin is visible without scrolling back. Unrendered connections are annotated on both boxes as → [06] motivates / ← [05] motivates. Events with no relationships show "No connected events." Every user message becomes a request box; events in the same session span with no other relationships are connected from that request as follow-up — this reflects conversational structure, not causal claims. Groups of events sharing no session or relationship (e.g. independent features from separate sessions) are rendered under ══ Flow 1/2 ══ headings. Titles and status are never truncated — they wrap. The selected box is drawn with a bold border; only its connected lines, annotations, and boxes are highlighted. When a panel is too narrow for boxes, it falls back to a branch list.

Status markers: ✓ confirmed (verified change, observed result, answered request), ! needs confirmation (unverified change, completion report, result without a verification target, unanswered request), ✗ failed.

The right selected event panel shows the description, the result that verified this change (or the change this result confirmed), connected events, and source evidence. Results and connected events are numbered 1–9; press the digit to navigate and Backspace to return. Evidence is shown with record type, source, timestamp, and line position. Single-line tool-call arguments with escaped newlines (\n-escaped patches etc.) are reflowed visually on screen only — stored evidence is verbatim. At ≥180 columns, all three panels are side-by-side; at ≥150 columns, the flow and detail panels are side-by-side; below that, they stack vertically.

Key / Option Action
↑↓ / j k Previous/next event. In per-request view, wraps to the next request at the end of a span
←→ / h l Jump to connected box in adjacent panel (change ↔ right-side result). Horizontal scroll in the narrow branch-list view
PgUp·PgDn / Space Scroll one screen
g G / Home·End First/last event
[ ] Previous/next request
1–9 / Backspace Navigate to numbered connection in the selected-event panel / return to previous event
! Jump to next event needing attention (!·✗); wraps from end to start
/ → n N Search titles, summaries, and source evidence; next/previous result. Matching box titles are underlined
Tab / Enter Toggle focus between the event flow and selected-event panels
J K Scroll the selected-event panel without changing focus
e Toggle full evidence expansion ↔ 8 lines per piece
z Zoom current panel to full screen ↔ restore
A Toggle full-flow / per-request-flow view
Esc Close search or zoom; return from selected-event panel to flow panel
? Show full keyboard help
Mouse Click to select a box, request, or numbered connection (clicking a line or [02] motivates ▶ annotation jumps to the connected box); scroll wheel navigates events/requests and scrolls detail
y Copy a reference (contexttrail:ev_6226b954@v12) to clipboard. Paste into a Claude Code or Codex conversation — the agent reads that event with its evidence. Uses pbcopy locally on macOS, OSC 52 over SSH; if neither works, select the text from the status line
R Explicit incremental analysis. Shows pending volume and choices, prompts for unit count, then runs while keeping the previous graph visible (view·analyze)
B Print the browser URL and SSH forwarding instructions for the current version / selected event (view·analyze, optional feature)
Q Quit. During analysis, confirms cancellation and cleans up child processes
--ascii Replace box-drawing characters, arrows, and status markers with ASCII (v ! x). Non-ASCII titles are preserved
--no-tui Print the annotated flow and status without interactive UI
--no-color Disable status colors. Markers (✓ ! ✗) are preserved. Also respects the NO_COLOR env var
--no-mouse Disable mouse capture and restore terminal text selection

Keys also work when a Korean IME is active (e.g. ㅂ → Q, ㅓ → J). If the IME is mid-composition when the terminal delays sending, a second keypress may be needed.

The UI does not depend on terminal image protocols or specific emulators. Mouse is used when available, but every action is accessible via keyboard. Real branches, merges, and back-edges are rendered; already-shown nodes appear as references. Long graphs scroll; deeply nested branches are collapsed to references.

Compatibility targets: standard SSH, Windows Terminal, IDE terminals, tmux, screen. Tested: xterm-256color, screen-256color, tmux-256color, linux, vt100 TERM values, and resize on a Linux PTY. This does not imply complete screen/keyboard compatibility verification on each product. TERM, CJK width, and font issues must be verified on real servers separately.

In TERM=dumb or when output is piped, the UI falls back to plain text. The contexttrail graph text output follows the flow with per-event description, verification, connections, and evidence — the same content as the right panel. ContextTrail agent commands installed into Codex/Claude Code read this output. If curses initialization fails on an unknown terminfo, a clear error is shown; use --no-tui to view saved results.

The analysis tokens counter at the bottom of the screen is the sum of calls where the runner reported both input and output tokens in this project's local LLM call ledger. If some calls lack usage data, + (confirmed calls/total calls) is appended. It does not include tokens used in the original Codex/Claude conversations or estimated costs.

Developer Execution Tracing

LangSmith tracing (--langsmith) is a developer tool for inspecting analysis structure and is not a production feature. It does not appear in command help; see LLM Ops for usage.

Viewing Analysis Runs in Studio (Developer)

contexttrail analyze and contexttrail eval invoke the same LangGraph execution graph as Studio's contexttrail_analysis. The LangGraph runtime is included in the base install; the optional studio install adds a development server. You can inspect extraction input preparation, model calls, candidate evidence validation, integration input preparation, graph change summaries, and SQLite publishing in live execution nodes. extract_input.request and integrate_input.request in node state are the exact tasks to be sent; candidate_audit and graph_change_audit contain source citations and reflected events/relationships. Additional read and repair calls appear in sub-traces. Claude can be read as log input but only Codex is used as the model runner. Project paths and eval fixtures for live mode are fixed via server environment variables, not Studio input.

python -m pip install -e '.[dev,studio]'
# Use only the LangSmith key from your private .env, with a separate tracing project
dotenv -f /path/to/private/.env run -- env -u OPENAI_API_KEY \
  LANGSMITH_TRACING=true LANGSMITH_PROJECT=contexttrail-dev \
  langgraph dev --no-browser --host 127.0.0.1 --port 2025

Replace /path/to/private/.env with your own private file path. In the Studio UI URL printed by the server, select contexttrail_analysis and run with empty input {} to practice with synthetic logs and the Mock Runner. To see a real Codex analysis, stop the server and restart with the project and fixture paths fixed as environment variables:

dotenv -f /path/to/private/.env run -- env -u OPENAI_API_KEY \
  CONTEXTTRAIL_STUDIO_SCOPE=/path/to/project \
  CONTEXTTRAIL_STUDIO_EVAL_FIXTURE=/path/to/reviewed-small-fixture.json \
  LANGSMITH_TRACING=true LANGSMITH_PROJECT=contexttrail-dev \
  langgraph dev --no-browser --host 127.0.0.1 --port 2025

Studio input {"mode":"eval","confirm_live":true,"max_units":1,"max_calls":10} analyzes the reviewed small fixture with real Codex and saves to a separate temporary state. {"mode":"live","confirm_live":true,"max_units":1,"max_calls":10} scans the full project and publishes to the existing ContextTrail state DB. Only one work unit is processed by default; if units remain, the run ends with partial. Actual model input/output is visible in the Codex structured response sub-trace in Studio. Records from contexttrail analyze run separately in the terminal do not appear retroactively in Studio UI (see LLM Ops for developer tracing). The Studio dev server has no user authentication — bind only to 127.0.0.1 and shut down after use. See Studio Architecture for details.


4. Browser Detail View

Press B in the TUI, or run in a separate terminal:

contexttrail serve /path/to/project --port 8765

The server binds to 127.0.0.1 only by default. Set up SSH port forwarding from your PC (replace user@server with your actual connection details):

ssh -L 127.0.0.1:8765:127.0.0.1:8765 user@server

Open the full URL including the token printed in the server terminal in your PC browser. If a port conflict causes a different port to be selected, use the actual port shown. The URL contains a private access token for your records — do not share it externally.

Select a graph node to inspect event/relationship/evidence levels, preserved source excerpts, and diff identification at the time of the event. F5, page load, and node selection do not call AI. Analysis must be explicitly requested via the "Analyze Changes" button and confirmation dialog. The token becomes invalid when the server stops.

No external CDNs or web fonts are used. The browser graph is a custom SVG layout generated from a safe internal Mermaid subset. This is not a full Mermaid engine bundle. Layout and actual browser rendering for complex graphs are pending further validation.


5. Saving and Exporting

To open the saved event flow directly in the terminal:

contexttrail graph /path/to/project
contexttrail graph /path/to/eval-output

In a terminal, this opens an interactive view with the event flow and selected event's description and evidence side by side. Use arrow keys to select events, Tab to focus the detail panel, and Q to close. When a screen cannot be opened (e.g. when piped or run from a coding agent), the same command prints event summaries, relationships, and evidence excerpts to stdout. No AI calls or external transmissions are made.

The same command works on contexttrail eval output directories. Stdout may contain private source citations — review before sharing.

To find or inspect individual events:

contexttrail find "install script"        # events with all words in title, description, or source evidence — newest first
contexttrail find                         # recent events and open items
contexttrail show ev_6226b954             # status, connections, and source evidence for one event
contexttrail show contexttrail:ev_6226b954@v12   # copied reference; notifies if event changed since v12

Output is brief and machine-friendly by default; --json is also supported. No AI calls are made — reads only from stored results.

Installing Codex / Claude Code Commands

./install.sh automatically registers two commands as personal skills in the current user's Codex and Claude Code installations. If you installed via pip directly, run contexttrail install-commands once. Files created by ContextTrail are updated on reinstall; separately authored files are not overwritten. Use contexttrail install-commands --force only if you want to replace those too. If the agent does not recognize the new commands, start a new session.

Task Codex Claude Code
Update graph $contexttrail-update [N units | current session] /contexttrail-update [N units | current session]
Load context from saved graph $contexttrail-context [reference | search term] /contexttrail-context [reference | search term]

The legacy slash-style variants /prompts:contexttrail-update and /prompts:contexttrail-context for Codex CLI/IDE are also installed for compatibility; Codex recommends the skill form. Both commands invoke the contexttrail CLI from the shell — no MCP server required.

  • Update runs only when the user calls it by name. Claude Code uses disable-model-invocation; Codex uses allow_implicit_invocation: false in agents/openai.yaml to prevent the agent from calling it on its own. It first shows pending units and per-choice token/time projections, then prompts for unit count. Providing a count (e.g. /contexttrail-update 5) skips the prompt; current session prioritizes the active conversation's session. Analysis uses the Codex Runner only (analyze --no-tui --brief --units N, printing a few-line summary instead of the full flow). Since it can take a long time, the agent runs it in the background where possible and reports progress as unit done k/N lines. Interrupted runs resume from the last completed unit. Running inside Codex may be blocked by Codex's own isolation — the agent will guide you to run the command in a separate terminal.
  • Context loading reads only stored results; no AI calls. The agent may invoke it proactively when the user asks about past decisions, attempts, or verifications. The agent uses contexttrail find "<query>" (or recent events and open items if no query) and contexttrail show <event> (description, connections, cited source evidence). Both support --json. Output and skill instructions note that source evidence is a quotation from past records and should be treated as reference — not instruction.
  • Attaching events as evidence: Press y in the TUI or click "Copy agent reference" in the browser to copy a reference like contexttrail:ev_6226b954@v12. Paste it into a Claude Code or Codex conversation — the agent reads that event and its evidence with show. If the event was re-analyzed and changed or removed since the copy, show will say so.
  • Using context loading in Claude Code sends cited Codex/Claude records to the model in that conversation (Anthropic).

Export

contexttrail export /path/to/project --format md --output ./project-flow.md
contexttrail export /path/to/project --format mmd --output ./project-flow.mmd
contexttrail export /path/to/project --format json --output ./project-flow.json

All formats are generated from the same stored graph — no AI calls. Markdown includes events, relationships, source evidence, and limitations. JSON includes the graph and cited evidence. MMD is Mermaid code. Use --force to overwrite an existing file. Export to DB, raw logs, or Git internal files is blocked.

Storage locations: Git repos use <git-common-dir>/contexttrail/<scope-key>/; plain directories use <folder>/.contexttrail/. The same repo root/worktree set shares state; different specified sub-paths get separate scopes. Full source text is not permanently copied — only metadata, digests, cited excerpts, and stage results are stored.

Permissions: state directory 0700, DB and new exports 0600. This does not imply disk encryption or automatic secret redaction. SQLite locking on network filesystems (NFS, etc.) is not validated.


6. What has been measured

Both real CLIs have been measured end to end on fixtures and on this repository's own logs. Analysis quality is still a draft to check, not a record. Below is every number actually measured.

Measured Not yet done
Real model calls Codex CLI 12 analysis calls (8 on this repository's own logs, 4 on the Linux external project, both gpt-6-sol) plus 2 doctor --smoke round-trips; 20 calls across the 6 scored fixture runs (Linux · table). Claude Code CLI: 42 eval runs on macOS with claude-sonnet-5-5 (2 installer smoke runs, 40 scored runs on the repairfix-v2 fixture), 3–4 calls each (record §9–14) Claude Runner on Linux: 0 calls (that host has no Claude CLI). No Claude run on an external project yet
Real analysis scale Linux: 5,282 in-scope records → 8 events / 9 edges, 4 runner calls, read-only check passed. macOS self-analysis: 4,199 records → 17 events / 15 edges (Linux · macOS) —
Cost per work unit Claude Sonnet, repairfix-v2 (26–34 records per unit): 1.5–2.4 minutes per run of 2 units, 130k–240k cached input tokens and 12k–21k output tokens, after the integration step was changed to publish a code-built draft and ask the model only for its changes (§12–14). The first Claude baseline was 2.7 minutes and about twice the tokens for 1 unit Codex cost after the same change: not measured
Archive parse audit 156 files (Claude 40, Codex 116), 31,748 records selected, limitations: 0 as of the a4 parser (audit · report) Stale for this parser. The record-type split in sources/local.py now reports types that audit did not. Re-run needed
Semantic quality Codex: 79/111 expectations met (71%) across 6 runs of 2 fixtures. Claude Sonnet on repairfix-v2: 11.8/18 mean over the latest 5 runs (11, 11, 13, 13, 11), two work units, integration against an existing graph, 0 bookkeeping rejections (§14) 2 human-checked WorkUnits total (14 records on macOS, 1 unit on Linux). Not a rate. The three expectations that fail in every Claude run are two-hop verifies judgements (a test run covering the code it calls)
Reproducibility Characterized for one configuration. Five same-configuration Claude Sonnet runs of repairfix-v2 scored 11, 11, 13, 13, 11; the per-item breakdown shows the spread comes from events being split or merged differently between runs, not from relations (§11, §14). Codex installer scored 8/18 and 16/18 on consecutive runs, confounded by different integration effort No same-configuration repeat with Codex
Platform validation Linux/Ubuntu 24.04 + bubblewrap: 1 end-to-end run on an external project (report) · macOS: sandbox-exec + Codex smoke + self-analysis + the 42 Claude eval runs above Claude Runner under bubblewrap: 0 calls (no Claude CLI on the Linux host)
Tool-denial tests Canary escape probe at startup on both macOS and Linux (SECURITY) ~, .ssh, project tree, and real credential write-blocking unverified on both
Tests 437 in the suite, all passing (2026-10-06; a4 results · state); CI runs them on Linux + macOS × Python 3.11–3.13 0 real model calls in CI, by design

Evidence behind these numbers is published, not summarized: design & evaluation history · performance plan and Claude measurements · pre-live audit · real-CLI evaluation · a3 validation · what is left.

Redaction in those files: real project names, native session UUIDs, source snapshot IDs, and machine-specific absolute paths are replaced with placeholders. What is deliberately kept is the aggregate evidence — record and session counts, work-unit counts, token totals, durations, and the archive SHA-256 digests, because a digest is what proves the audit did not modify the archive. The LICENSE copyright name is unchanged, and so is the GitHub account in the badge above.

The quality row is the honest one: extraction is good enough to be useful on a project you remember well, and there is no measurement showing it is stable across repeated runs. Treat a reconstructed flow as a draft to check, not a record.

A zero-real-model walkthrough is available after installation: python scripts/prelive_walkthrough.py --output /tmp/pf-prelive-walkthrough

7. Explicit Limitations of This Alpha

  1. Analysis quality is measured only on two small fixtures and this repository's own logs. Both CLIs have run end to end (Codex on Linux and macOS, Claude on macOS), but a reconstruction of a large, unfamiliar project has not been checked by a person. The known failure modes are events split or merged differently between runs, and two-hop verifies links (a test run that covers the code it calls) that cheap models do not draw.
  2. Not all historical records and files are read without limit. The default Git commit scope is the 50 most recent per worktree, plus tracked staged/unstaged diffs. Untracked content, all merge parents, and an arbitrary-revision comparison UI are out of scope. Sub-agents whose parent linkage cannot be verified are excluded.
  3. Long source text is split while preserving semantic lineage. The parser splits into stable fragments of ≤32,000 characters and preserves fragment_of/index/count. Binary/Base64 payloads (e.g. images) are excluded from text model input; only media type, size, and SHA-256 marker are kept. The total input budget per work unit is still limited — very large context combinations may be deferred.
  4. Historical context retrieval is also limited. Defaults: 24 relevant events, 200 indexed events + 250 source records + 100 pinned files, up to 2 additional evidence reads, and at most 1 format/reference repair. After repair, invalid candidates are discarded; discarded candidates remain in the graph's limitation list by title. If selected context exceeds the budget, the work unit is deferred/failed. Large-project tuning remains to be done.
  5. Evidence presence checks do not guarantee semantic correctness. Citation scope, actor source, and execution evidence presence are checked in code, but whether the source text actually supports a given interpretation or causal claim requires AI and user evaluation.
  6. Not all MVP acceptance criteria are complete. Safe mixed-source analysis with both real runners, semantic regression evaluation, and real terminal/browser validation are outstanding. Explicit historical re-analysis commands and user graph editing are also not yet provided.

Budget tuning example (review increased data transmission and model limit effects together):

contexttrail analyze /path/to/project --history-limit 100 --record-chars 64000 --unit-chars 80000

The constraint record_chars ≤ unit_chars < task_chars must hold. task_chars and detailed context values are adjusted in the Python AnalysisConfig. Code or config changes alone do not trigger automatic re-analysis of completed records.


8. Development and Regression Testing

python3 -m venv .venv && .venv/bin/python -m pip install -e '.[dev]'
PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 .venv/bin/python -m pytest -q

Or simply scripts/test.sh -q, which picks the right interpreter. On most current Linux and macOS systems there is no bare python, so prefer python3 or an explicit virtualenv.

This reproducible command prevents unnecessary external pytest plugins from auto-activating. Tests create synthetic logs and Git fixtures in temporary directories; HTTP tests use a loopback port; TUI tests use a PTY. No real AI accounts are called.

CI runs this suite on Linux and macOS across Python 3.11, 3.12, and 3.13 — see .github/workflows/test.yml. Linux is included because it is the platform README claims as primary; a green badge there is a real signal, not a formality.

The mouse assertion in tests/test_cli_ui.py checks the portable invariant — with --no-mouse no mouse-reporting enable sequence may appear, and any enable that does appear must have a matching disable before exit — rather than one terminfo-specific escape sequence, so it does not depend on the local terminfo database. A separate test in the same file renders across xterm-256color, screen-256color, tmux-256color, linux, and vt100.

Structure:

src/contexttrail/
  sources/           local log adapters
  runners/           Codex·Claude CLI / bubblewrap isolation
  analysis.py        incremental planning, ReadRequest, extract/integrate
  schema.py          JSON contract, evidence validation, GraphDelta application
  git_context.py     scope/worktree, pinned Git evidence
  store.py           SQLite, lock, ledger, atomic publish
  render.py          Mermaid subset, character graph, SVG, export
  ui.py              curses TUI
  webview.py         loopback detail server
  prompts/           SPEC-based analysis instructions
  assets/            local HTML·JS·CSS
  demo.py            Mock Runner for synthetic fixtures only

See Documentation Guide for the public document list. Design basis: PRD and SPEC. Also read Implementation Status, Test Report, and Security & Execution Constraints.

Contributions are welcome: see CONTRIBUTING.md for the invariants every change must keep and how model-visible changes are evaluated, CODE_OF_CONDUCT.md, and SECURITY.md for private vulnerability reporting. Changes by date are in CHANGELOG.md.

Distributed under the MIT License (LICENSE). External packages installed alongside this tool and their licenses are listed in THIRD_PARTY_NOTICES.md.

Metadata

Release files for contexttrail 0.1.0a5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for contexttrail 0.1.0a5
File Size Uploaded
contexttrail-0.1.0a5.tar.gz 351.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for contexttrail 0.1.0a5
File Interpreter ABI Platform
contexttrail-0.1.0a5-py3-none-any.whl Python 3 none any Details

Total release size: 592.1 kB

Release files / contexttrail-0.1.0a5.tar.gz

Download URL contexttrail-0.1.0a5.tar.gz
Size 351.5 kB
Tags Source
SHA-256 checksum
How to use checksums
8b96afbf8f3dde3dd395df72c89d96f7570cc1e05b43eedb5687c621378b8e9d
BLAKE2b-256 checksum
How to use checksums
682a0ebdc4487c596fea85e9ac5406ba63b78bb3995525084a16e8128963a8fd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / contexttrail-0.1.0a5-py3-none-any.whl

Download URL contexttrail-0.1.0a5-py3-none-any.whl
Size 240.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
19b3a72973b1cdfeecaf0a561192dfc01d6fa5156c6678aa3c9168fd246d8787
BLAKE2b-256 checksum
How to use checksums
c3c4e61541b6e7314cefc91da3ec460194667b79facedeaaa721ebebcce4c577
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0a5 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page