Skip to main content

transcripto

your coding agents keep a transcript of every session. it is the most valuable dataset you own, and you cannot scroll back far enough to read it. transcripto indexes it, keeps only the turns you actually typed, and grades them.

one command, no account, no signup, no cloud. it reads files that are already on your disk and never opens a socket.

uvx transcripto coach

check which build you got. trace, --harness cursor, and the "run index first" errors below all arrive in 0.1.2. On anything older, trace is not a command, cursor is rejected, and the index-gated commands fail with a raw sqlite traceback instead of an instruction. One line settles which one you are holding:

uvx transcripto --version          # 0.1.2 or newer = this README is accurate

Read 2026-08-31: PyPI was serving 0.1.1 while this README described 0.1.2, so uvx transcripto gave the older build. If --version is not even a recognised flag, you have 0.1.1 — it was added in 0.1.2 precisely because there was no way to tell. The repo always matches this README and needs nothing installed:

git clone https://github.com/Morkeeth/transcripto && cd transcripto
python3 transcripto.py coach

three harnesses, one instrument

transcripto coach                     # Claude Code, ~/.claude/projects
transcripto coach --harness codex     # Codex, ~/.codex
transcripto coach --harness cursor    # Cursor, ~/.cursor/projects/*/agent-transcripts

Authorship is not the same gate in all three, and the tool says so rather than pooling them. Claude Code stamps promptSource: typed, which is the measured-reliable signal: about 95% of raw type: user records are not the operator at all. Cursor has no such field. Its one honest equivalent is the <user_query> wrapper it puts around a submitted prompt, which injected and tool-result records do not carry. That is a weaker signal and it is labelled weaker.

trace — what actually happened after you asked

ask shows what you typed. find shows what a file went through. Neither answers the question that matters after the fact: you asked for X, did anything durable happen?

transcripto trace "the gate"

It walks each of your matching prompts forward inside its own session and lists the writes and edits that followed, stopping at your next prompt so one turn cannot claim the next turn's work. Green dot = something durable landed. Red = nothing was touched.

Honest limit: a write following a prompt in the same session is CO-OCCURRENCE, not proof the write was caused by that prompt or that it was correct. Same proxy coach uses, labelled the same way.

what you get back

this is a real run on one machine, pasted unedited, 2026-08-28:


  YOUR PROMPT HABITS, GRADED   (offline, your machine only)

  harness: claude
  corpus : 2720 transcript(s), 381,804 records
  kept   : 3678 prompts you actually typed (0.96% of records)
  episodes: 1946 ranked, 910 survived (47%)
  tiers  : commit 373 | write/edit 537 | reverted 2 | nothing durable 1034

  SURVIVAL IS A PROXY: survival = a durable Write/Edit or an un-reverted git commit in-episode. A PROXY, not proof the work was correct or shipped.

  SURVIVES MOST  do more of these:
     64%  (167/260)  detailed (>40 words)
     63%  (66/104)  states-a-check-or-done-condition
     58%  (414/708)  intent:CHANGE
     58%  (21/36)  no-object (pronoun/vague)
     57%  (99/175)  cites-a-file-or-path

  SURVIVES LEAST  these tend to loop:
     33%  (4/12)  intent:REVERT
     39%  (411/1043)  intent:none
     40%  (4/10)  intent:TEST
     42%  (298/715)  terse (<8 words)
     45%  (77/173)  intent:DESCRIBE

  + your best landed prompt, with its witness:
      "mTERMINAL 8 — Mountain of Helicon · ~/CODE/mountain-of-helicon Read ~/CODE/mountain-of-helicon. Two…"
      COMMIT-WITNESSED: git commit  ·  corrections: 0

  - your worst looped prompt, with its witness:
      "ok, and lets see they might solve it in the future so i can go back to my beloeved routine :) Befor…"
      NO-DURABLE-RECORD: read-only Bash only, no file change  ·  corrections: 15  ·  assistant turns: 221

those are my numbers on that date, and they move every session i run, so treat them as a snapshot rather than a constant. yours will be different, which is the whole point. the last two lines are the ones that sting: it hands you back your own best and worst prompt, verbatim, with the receipt for why it scored each one.

on that machine, on that date, prompts that wrote down what done looks like survived 63% of the time (66 of 104). prompts with no stated intent survived 39% (411 of 1043). i had spent a year blaming the model.

one day later, 2026-08-29, the same command on the same machine read 63% (67 of 107) and 40% (424 of 1072) over 2,874 transcripts. the percentages held and the denominators moved, which is what a snapshot is supposed to do.

the proxy caveat, which travels with every number

an episode "survived" if a Write or Edit landed, or a git commit ran and nothing reverted it inside the same transcript.

that is a durable keystroke, not a durable outcome. a commit is not proof the code was right. a revert in a later session is invisible to it. a prompt whose payoff was a decision rather than an edit reads as dead. it is a coaching signal, not a verdict. if it ever prints something that flatters you, distrust it.

the caveat is printed in the output itself, every run, on purpose.

why your own gate matters here

at fleet scale roughly 95% of the type: user records in a transcript are not you. they are tool results, injected skill bodies, sub-agent prompts, and messages from other terminals, all wearing your role. transcripto gates on promptSource (typed/queued, no meta, no sidechain) so it grades what you typed.

you can watch the gate do work: in the run above, 3678 of 381,804 records survived it. that is 0.96%.

the same gate is what makes cost produce a number a spend tracker cannot:

cost per human decision  last 30 days  2026-07-28 → 2026-08-27

  API-equivalent spend      $8,892.49
  your decisions            2934   turns you actually typed (promptSource typed/queued)
  ────────────────────────────────────────────────────
  cost per human decision   $3.03

  46.4k agent messages · 16 per decision · 11.5B tokens · 149 sessions
  57.7k raw `type: user` records in the same window. dividing by those instead
  would read $0.15, 19.7x too cheap.

it prints both, so the gate's effect is something you can check rather than something i am asserting. these are API-equivalent dollars at list rates, because a transcript has no cost field, only token counts. on a subscription you did not pay this.

honest limits

read these before you quote a number at anyone.

  • survival is a proxy, described above. durable keystroke, not durable outcome.
  • one operator's corpus. every figure in this README comes from one machine. it is an existence proof that the measurement runs, not a finding about how people prompt. run it on yours and you get yours.
  • three harnesses today: Claude Code, Codex, Cursor. nothing else is supported. aider and the rest are not read. and the three are not equal: Claude Code has a measured-reliable authorship field, Cursor has only the <user_query> wrapper, which is weaker and is labelled weaker wherever it is used.
  • the habit labels are heuristics. "states-a-check-or-done-condition" is a pattern match over your text, not comprehension. it will misfile some prompts.
  • correlation, not instruction. detailed prompts surviving more often does not prove that padding a prompt causes survival.

privacy

it runs locally and never touches the network. there is no socket, no urllib, no requests, no subprocess, no telemetry, no analytics, and no account. that is a claim, so here is the grep that settles it against the single file it ships as:

$ grep -nE '^[[:space:]]*(import|from) ' transcripto.py
8:import sys, os, json, glob, re, sqlite3, argparse
9:from datetime import datetime, timezone
260:    import time
1051:            import datetime
1061:    from the separator), so the result is checked on disk and dropped if it is

five lines, four of which are imports and all four are stdlib. time and datetime sit inside functions, which is why the pattern allows for indentation — anchor it at ^import and you would miss two, so do not take my word for the anchor either. line 1061 is the pattern catching a docstring that happens to begin with the word from; it is prose, not an import, and it is left in rather than tuned out, because a grep you tuned until it agreed with you proves nothing.

what the list does NOT contain is the actual claim: no socket, no urllib, no requests, no http.client, no subprocess. that one is checkable too, and the right answer is no output at all:

$ grep -nE '\b(socket|urllib|requests|http\.client|subprocess)\b' transcripto.py
$

your transcripts stay in ~/.claude, ~/.codex and ~/.cursor. the index it builds stays in ~/.trace.

the rest of it

coach and cost read your transcript files directly and need nothing set up. the other six read a local index, so run this once first:

transcripto index      # a few minutes on a large corpus, incremental after that

on a 2,874-file corpus that was 164 seconds, measured 2026-08-29. if you skip it, the six say so and exit 2.

transcripto index      build / refresh (incremental)
transcripto watch      live, new sessions get picked up as your agents work
transcripto ask        YOUR OWN messages about a topic, newest first + a rollup
transcripto search     full-text across everything (you + agents + tool logs)
transcripto find       every session that wrote / edited / read a file
transcripto trace      what durably happened after each prompt you typed (0.1.2+)
transcripto sessions   recent sessions + the first prompt YOU typed in each
transcripto stats      what you actually work on
transcripto cost       what ONE of your decisions costs
transcripto coach      which of YOUR prompt habits survive (a proxy)

ask is the one that kills "wait, did i lose something?". it answers "what was i thinking about X across ALL my sessions", in your own words only.

$ transcripto find USER-JOURNEY.md          # run 2026-08-29
USER-JOURNEY.md  4 touches across sessions (3 were writes/edits)

2026-08-20  WROTE  ~/CODE/mountain-of-helicon-main/USER-JOURNEY.md            abd9e871
2026-08-21  WROTE  ~/…/Obsidian LIFE/00 Dashboard/suite-user-journey.md       0f845ede
2026-08-27  read   ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md                 cddfde29
2026-08-27  WROTE  ~/CODE/hack-fleet-ata/docs/USER-JOURNEY.md                 cddfde29

the file you lost, found across every session you ever ran, with the session id that touched it. find needs transcripto index first.

Codex

transcripto coach --harness codex

reads ~/.codex (sessions + archived_sessions), normalises it into the same rows, and applies the identical survival proxy. it also ingests history.jsonl purely as a control on the gate: it reports how many of its input lines also show up as typed rollout turns, so you can see the gate agreeing with a second source.

install

uvx transcripto coach

no install, nothing to set up. or put it on your PATH:

pipx install transcripto

or run the single file with no packaging at all:

git clone https://github.com/Morkeeth/transcripto
cd transcripto
python3 transcripto.py coach

no dependencies, stdlib only, one file. the packaging adds nothing at runtime, it just gives the file a name on your PATH.

tests

./test_coach.sh        15 assertions
./test_codex.sh        14 assertions
./test_cost.sh         12 assertions
./test_small_n.sh       7 assertions
./test_label_bands.sh  13 assertions

61 assertions, all green, re-run 2026-08-31.

offline, no keys, on fixtures that inherit the real transcript shape including all four ways a non-human record disguises itself as type: user.

the load-bearing one in test_coach.sh is REVERTED IS NOT SURVIVED: a commit that got reset --hard in the same session left no durable record. flip that one line and the suite goes red, which is the point. a generous proxy is a broken one.

test_label_bands.sh is the other one, and it exists because 0.1.1 shipped the defect it pins. SURVIVES MOST took the top five habits and SURVIVES LEAST took the bottom five, which overlap whenever you have fewer than ten rankable habits — so a new user, who necessarily has few, read the same habit at the same percentage under both "do more of these" and "these tend to loop". on a 3-habit corpus 0.1.1 reprinted all three, all at 65% (22/34). the suite is red on the published 0.1.1 file and green on this one, and its wide band asserts the fix leaves a large corpus byte-identical.

why though

your agent history is proof. every "yeah it's done" has a real trace sitting behind it. transcripto is the index that makes it checkable.

it is the fuel layer. on top of it you check what your agents claim against what the trace shows, which is mountain of helicon. the pitch was never "search your history". it is prove your agent did what it said, from your own local traces.

local, MIT, no telemetry. star it if it finds you something you'd lost ™

Metadata

Release files for transcripto 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for transcripto 0.1.2
File Size Uploaded
transcripto-0.1.2.tar.gz 32.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for transcripto 0.1.2
File Interpreter ABI Platform
transcripto-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 64.5 kB

Release files / transcripto-0.1.2.tar.gz

Download URL transcripto-0.1.2.tar.gz
Size 32.0 kB
Tags Source
SHA-256 checksum
How to use checksums
5aea0babd7f93ab78dd52535a38159805c1b2fbccb74b308237304a43b6fa6f4
BLAKE2b-256 checksum
How to use checksums
183d8700531362f0388402acd1df13f4060b0aa3f152d24eae95ca103c22c425
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.26 {"installer":{"name":"uv","version":"0.9.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / transcripto-0.1.2-py3-none-any.whl

Download URL transcripto-0.1.2-py3-none-any.whl
Size 32.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4617d8c7d7fb1251f31b5dde57964b58e6bbfcb2dedbb07dfb223021adbc3799
BLAKE2b-256 checksum
How to use checksums
828811b231606e21786ae0c6fdd33e91b3405dfc94a1974793b9cb1d4bf1f147
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.9.26 {"installer":{"name":"uv","version":"0.9.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.2.1

2 release files

0.2.0

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page