Skip to main content

TailTrail logo

TailTrail

Plan first. Change with evidence. Finish without drift.

TailTrail is a local, approval-first workflow for AI-assisted software delivery. It helps an agent understand an existing project, propose a bounded change, collect real validation evidence, and keep multi-file work tied to the approved requirement.

It works with Codex, GitHub Copilot, Claude, Cursor, ChatGPT, and Gemini.

The self-contained tailtrail wheel and sdist support CPython 3.12 and 3.13, have no runtime dependencies, verify their packaged resources before command dispatch, and do not need a source checkout. See INSTALL.md.

Get a plan in two minutes

  1. Install TailTrail into the project you want to work in. Use the one installation guide—it has Windows, macOS/Linux, update, and host-specific instructions.
  2. Open a new chat in your AI host.
  3. Ask TailTrail to plan the task:
tailtrail start "add payment retry handling"

TailTrail returns a Planning Lock and run ID. It does not implement the task, run tests, or change Git until you approve the plan.

Leave TailTrail without losing that run at any time:

tailtrail stop

Ordinary prompts then return to the host agent. Resume later with the exact saved identity: tailtrail resume --run-id <run-id>. Stop is not rejection or cancellation, and resume never approves or advances the workflow.

Choose your host

Host Fast path
Codex Codex quickstart
GitHub Copilot Copilot quickstart
Claude Claude quickstart

TailTrail in plain language

Coding agents can generate code quickly, but larger tasks often fail in less obvious ways: a caller is missed, one requirement is forgotten, tests prove only the easiest path, or repeated corrections move the implementation away from the original request.

TailTrail adds a local control and evidence layer around the agent. It keeps the approved intent visible, selects only the controls the task needs, checks the result from several angles, and produces one completion report. It does not replace the coding agent—the agent still reads and writes the code—but it makes the delivery easier to inspect, correct, resume, and trust.

This is useful when you want to:

  • keep a small fix small and avoid unnecessary refactoring;
  • deliver a multi-file feature without missing callers, contracts, or tests;
  • ask questions about a plan before approving it;
  • recover from repeated failures without losing unrelated work;
  • see what was actually validated instead of accepting a generic “tests pass”;
  • run the same governed workflow from Codex, Copilot, Claude, CLI, or MCP.

The daily flow

flowchart LR
    A["Describe the task"] --> B["TailTrail Start\nPlanning Lock"]
    B --> C{"Approve?"}
    C -->|"Revise"| B
    C -->|"Approve"| D["Scoped implementation\nand evidence"]
    D --> E["Completion Report\nrequirements, tests, drift"]

For a normal code change, the entire conversation can stay this simple:

tailtrail start "fix the zero quantity validation defect"
tailtrail discuss --question "Why was this scope selected?"
tailtrail approve
tailtrail continue
tailtrail flow status
tailtrail close

These six verbs use one orchestration façade. TailTrail resolves a run only when it is unambiguous, approves only the exact plan or next frozen stage, and keeps advanced workflow commands available for diagnostics. Bare tailtrail status still means installer status; use tailtrail flow status for the auto-resolved task or tailtrail status --run-id <run-id> for an explicit one.

Start plans select detail automatically: AIDLC Off is Quick; AIDLC Lite is Expert without dedicated Architecture/Behaviour planning detail; Standard, Full, hands-free, and Intent Bridge plans are comprehensive. Users do not need a presentation flag. Add --verbose in any mode to force the complete canonical projection; requirements, scope, controls, approval, and workflow authority remain identical. Completion Reports are always comprehensive.

Reports use one canonical presentation contract across CLI, MCP, Codex, Copilot, and Claude. Narrow terminals wrap without dropping sections, verbose mode keeps all applicable sections, and a collapsed host surface must tell the user to open the complete report instead of presenting a partial substitute. Maintainers can verify the deterministic plan/debug/closure matrix with:

tailtrail presentation conformance

Maintainers can also check whether registered commands, MCP tools, files, documentation ownership, and core module dependencies remain consistent:

tailtrail maturity maintainability validate

Use hands-free: or end-to-end: only when you deliberately want TailTrail to break a larger delivery into approved slices.

What TailTrail packs

The Evaluation Harness includes a real-evaluation portfolio protocol: 18 task classes across five repository fixtures, blinded A/B grading, repeated runs, immutable positive/neutral/negative observations, and an honest claim gate. Run tailtrail eval real-portfolio report --root . to inspect coverage. Until all required observations exist, it reports no-performance-claim.

Enterprise operations include an offline conformance suite for compatibility, transactional installation/update/rollback, policy, linked CI, retention/export, migration/recovery, threat controls, and support boundaries. Run tailtrail enterprise-readiness --root . conformance; local success stays separate from hosted Windows/macOS/Linux release qualification.

TailTrail does not run every feature for every task. Navigator selects the smallest useful workflow during planning; other controls remain armed and activate only when their trigger—such as drift, a failed correction, UI work, or release risk—actually occurs.

Harnesses and assurance loops

Implemented Harness What it checks or controls Why it matters
Requirement Completion Harness V1–V4 Maps stable requirement IDs to likely and actual code paths, preservation rules, evidence, checkpoints, convergence, closure, and recovery. Prevents “code changed” from being mistaken for “the requirement is complete.”
Architecture Fitness Harness Checks callers, layers, contracts, dependency direction, expected files, and unexpected architectural change. Catches service-only fixes, missed callers, wrong-layer logic, and architecture drift.
Behaviour Harness Compares approved user/API scenarios with observed behaviour evidence. Proves the user-facing flow, including failures and side effects, rather than accepting unit tests alone.
Maintainability Harness Looks for duplicate logic, unnecessary abstractions, test-chasing, excessive churn, and unjustified scope. Keeps agent-generated changes understandable and aligned with existing project patterns.
Context Continuity Harness V1–V3 Carries forward the active requirement, prior decisions, failed attempts, drift, and “do not repeat” reminders; its watcher/advisory layer can remind the main agent when intervention signals appear. Reduces repeated mistakes and keeps a long-running agent focused without reloading the entire history.
Program Delivery Harness and deterministic orchestrator Breaks hands-free or end-to-end programmes into dependency-ordered features, slices, checkpoints, and approval gates. Makes large deliveries resumable and prevents one uncontrolled implementation pass.
Evidence-Aware Testing Selects a testing profile, minimum evidence tier, requirement links, receipts, CI inputs, flaky-test posture, and evidence metrics. Matches proof to the task instead of running arbitrary tests or trusting a generic pass statement.
Higher-Tier Testing and Release Confidence Covers integration, contract, behaviour/E2E, migration, environment, deployment, rollback, release policy, and calibration evidence. Extends confidence beyond unit tests when the change crosses system or release boundaries.
Token Harness and budgeting Routes and reduces safe context, preserves exact evidence, records estimates, and accepts measured telemetry only when linked. Keeps context manageable without treating estimates as exact model usage or dropping critical source and policy.
Evaluation Harness Runs deterministic scenarios, datasets, normalization, baseline comparisons, workflow outcomes, token evidence, and delivery evaluation. Measures whether TailTrail improves completion and drift control instead of relying on product claims.
Meta-Harness Reviews TailTrail’s own workflow fit, confidence, guardrails, context, metrics, learning, and proposal readiness. Detects when TailTrail itself selected too much, too little, or the wrong control.
Benchmark and Efficacy Harness Runs repeatable benchmark fixtures and analyzes captured efficacy results. Provides measured local evidence for product evaluation while keeping live-model claims separate.
Harness convergence, templates, and finalization Selects project-specific templates, compares checkpoints, bounds correction cycles, finalizes selected Harnesses, and creates one Completion Report. Gives all Harnesses a shared lifecycle instead of producing disconnected assessments.

Other major product features

Feature area What is included How it helps
Navigator and TailTrail Start Automatic task classification, feature selection, skipped/deferred explanations, focused validation, token estimate, and a Planning Lock. Gives the user a TailTrail decision before implementation starts.
Target and code intelligence Enterprise Target Workspace Resolver, input-role registry, Navigator-owned graph creation/reuse/incremental refresh/rebuild, caller/test discovery, immutable hash-bound run mappings, semantic evidence labels, and cross-repository reference mapping. Keeps the agent in the correct editable repository, preserves useful context across tasks, and validates ownership against current local evidence.
Canonical requirements and anchors Versioned requirement IDs, requirement-to-impact matrix, immutable approved intent, phase/slice anchors, actual state, and amendment history. Provides the stable reference used for implementation, drift detection, evaluation, and selective recovery.
AIDLC integration Off, Lite, verified official Standard, and verified official Full modes; Question Orchestrator grounding and requirement traceability; official questions, recommendations, stage approvals, session resume/redo/jump/recovery, and evidence/closure adapters. Matches lifecycle depth to task complexity, improves question relevance, and never presents a local questionnaire as official AIDLC.
Intent Bridge Detects, imports, versions, maps, amends, and converges existing structured requirement sources. Lets source-owned specifications remain authoritative while TailTrail manages delivery evidence and drift.
Interactive Plan Mode Evidence-backed explain/discuss, bounded investigation, question clarification/challenge, plan revision, AIDLC mode switch, and Expert Plan Customization. Lets users challenge any file, requirement, feature, risk, validation, token, or approval decision without rejecting the whole plan.
Durable Workflow Runtime Canonical ownership, workflow identity, state machine, task locks, freshness, declarative capabilities, mode-aware execution authority, approvals, adapters, retries, pause/resume, CI continuation, replay, retention, and closure. Lite/Off can reuse one approved-plan grant for safe local work; official AI-DLC and Intent Bridge retain their material gates. Moves long-running work from chat memory into deterministic local state while avoiding per-command approval fatigue without hiding sensitive authority.
Advanced runtime boundaries Approval-gated contracts for multi-agent graphs, source-writing MCP operations, cloud/Kubernetes runners, model-based diagnosis, live evaluation, and externally measured claims. Lets advanced execution be integrated explicitly without pretending a local contract is a running external service.
Failure, drift, and bounded correction Failure intake receipts, classification, sanitized fingerprints, requirement/drift mapping, correction packets, repeated-cycle detection, and recovery/replan routing. Acknowledges pasted failures, avoids infinite retry loops, and fixes the relevant requirement under the same run.
Safe Git recovery and Mode B diagnosis Git readiness, task recovery boundaries, checkpoints, selective reconciliation, conflict classification, and an evidence-based Recovery Diagnostician. Preserves completed and unrelated work while reverting or repairing only the active task’s owned delta.
UI Consistency Guardrail Read-only discovery of components, tokens, layouts, responsive rules, accessibility patterns, and existing visual tests. Makes UI agents reuse the repository’s design system instead of introducing a parallel one.
Closure, dashboard, and guarded learning Execution-evidence recorder, selected-Harness finalizer, requirement/control status tables, token posture, acceptance choices, deterministic evaluation, candidate-only learning, and a local workflow dashboard. Shows what is complete, missing, unresolved, measured, or unavailable before learning from the result.
MCP and host adapters Inspection-first MCP tools, approval-gated controlled operations, and composed Codex, GitHub Copilot, Claude, Cursor, ChatGPT, and Gemini guidance with conformance receipts. Makes the same workflow callable and debuggable across hosts without giving every tool unrestricted write authority.
Guardrails and quality intelligence Dependency Gate, security/review lenses, Sonar and vulnerability evidence, CI summaries, policy checks, and versioned JSON/SARIF repository enforcement. Preserves safety and makes heavy or state-changing checks explicit, reviewable, and evidence-labelled.
Enterprise workflow controls Enterprise target policy, identity binding, tenant/repository/actor context, leases and fencing, audit events, backup/restore validation, migration/rollback, observability, and conformance. Adds deterministic local enterprise contracts while keeping remote-service claims evidence-gated.
Packaging, installation, and supply chain Transactional install/update/repair/uninstall, ownership manifests, rollback journals, self-contained package validation, platform qualification, host quickstarts, and release manifests. Makes TailTrail installable and recoverable without requiring users to understand its internal schemas on day one.
Reporting and product learning Outcome telemetry, quality loops, value reports, sanitized learning candidates, refresh/review, and public-claim boundaries. Improves future routing from trusted evidence while current source, tests, policy, and user direction remain authoritative.

A simple mental model

Navigator decides what matters
    -> AIDLC clarifies requirements when needed
    -> the agent implements the approved scope
    -> Harnesses check intent, architecture, behaviour, and evidence
    -> bounded loops correct gaps or recover safely
    -> one Completion Report shows what is complete and what is not

Use the right page

Trust boundary

TailTrail is local and evidence-aware. It does not replace source inspection, tests, CI, scanners, code review, security review, or release approval. It never claims those checks passed unless it has their actual receipt.

For contributors, the repository CI validates Python compatibility, adapter contracts, registry consistency, installer smoke behavior, selected guardrail classes, and dependency decisions. See IMPROVEMENT-PLAN.md for the delivery roadmap and context/guardrail-layers.md for the layered enforcement model.

Metadata

Release files for tailtrail 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tailtrail 1.0.0
File Size Uploaded
tailtrail-1.0.0.tar.gz 2.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for tailtrail 1.0.0
File Interpreter ABI Platform
tailtrail-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 4.6 MB

Release files / tailtrail-1.0.0.tar.gz

Download URL tailtrail-1.0.0.tar.gz
Size 2.0 MB
Tags Source
SHA-256 checksum
How to use checksums
aa40a0fe242c47732f9cc4c3d43428826541c7e3c2d2a89c41a6af4e90162f53
BLAKE2b-256 checksum
How to use checksums
82d37435f40189e88571c58b7637c149e9ba25ccf77b1a7db2fbf2b1ed963ce2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release files / tailtrail-1.0.0-py3-none-any.whl

Download URL tailtrail-1.0.0-py3-none-any.whl
Size 2.6 MB
Tags Python 3
SHA-256 checksum
How to use checksums
aac02311be71d506357ac5659129e3c7be1a105badc7fb23e410882636af3d6c
BLAKE2b-256 checksum
How to use checksums
f463343a3b5e6be1898f4b0c1d95b46171ab62a9517e28bc346dd739a40faa65
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release history Release notifications | RSS feed

1.1.0

2 release files

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page