Skip to main content

where are we

Stop paying your agent to grep.

PyPI CI License

One tree walk writes entry points, routes, data model, step signatures, and every duplicate or dead test into AGENTS.md — or JSON for your own harness. Your agent starts working at turn one, not turn 41.

What you save

Measured on a real 184-feature behave suite and one production agent run:

Without the map With the map
Orientation before real work ~40 turns of ls/find/grep 1 turn: one --ask
Repo context re-sent each turn the map inlined, ≈ 64k tokens a pointer, ≈ 212 tokens (300× less)
That context over one run 27.4M tokens — a quarter of the whole run a few KB total
Budget it drained a 5-hour allowance gone in 74 min the budget goes to work
Cost to build it one tree walk: 10s, offline, 0 tokens

The map (121 KB) stays on disk to grep; a 616-byte pointer is what an agent carries. You pay a ten-second tree walk once and stop paying for rediscovery on every turn. The same walk pays for itself a second way in CI: the call graph it already wrote is what says which scenarios a commit reaches, so what to re-run after a change is a selection the map hands you rather than the whole suite every time.

What it does

One command turns a repository into a map an agent reads before it works: layers, entry points, routes, data model, contracts, tests, and where every name is defined. It reads the tree offline in seconds, and the same tree gives the same map every time.

Without it a session spends its first turns rediscovering the repo:

# turn 1   ls; find . -name "*steps*"
# turn 7   grep -rn "def click_pay" .
# turn 19  cat conftest.py; cat tox.ini; cat Makefile
# turn 34  grep -rn "BASE_URL" .
# turn 41  first line of actual work

With the map, turn 1 is the work. Measured on a real suite: forty-odd orientation turns become one --ask, and the agent reads an 849-byte pointer instead of grepping a repository it has not seen.

$ where-are-we --repo . --agent-file AGENTS.md --max-lines 200

framework map: 66 step modules, 1359 steps, 179 features, 1782 scenarios -> ./framework_map.md

AGENTS.md gets a pointer - 849 bytes, not the map:

## The framework map

`framework_map.md` (123 KB) is a generated map of this suite and the product it
tests. It is on disk on purpose: read from it, do not carry it. Ask it before
grepping the repository — it already knows.

    where-are-we --ask "the words you need"

That prints only the rows that mention those words, whole, and says how much of each section it left out. `--sections` lists what
is in it.

It has these sections:

- Where things are
- What a step may call
- Steps that overlap (14 pairs) — check whether one already does what you need
- What past runs measured (slowest first)
- 

And the map answers questions instead of being read:

$ where-are-we --ask "refund settled invoice"

## What past runs measured (slowest first)
- `billing/`
  - `refund.feature:88` Refund a settled invoice — ~252s, failed 3×
  - `credit_note.feature:12` Refund a settled invoice by credit note — ~40s
… 61 rows in this section do not mention these words

## Steps that overlap (14 pairs)
- 0.88: "the invoice is settled" (`billing_steps.py`) ≈ "an invoice has settled" (`api_steps.py`)
… 2 more matching rows; 11 rows in this section do not mention these words

An answer is whole rows, never a row cut in the middle, and it fits the limit it was given (12,000 characters for the CLI and the MCP) rather than filling it. The one exception is a continuation: where a more: handle lands on a single row longer than the whole budget, that row comes back cut and marked … (row cut to fit; N of M characters) with the handle on the row after it, because the alternative is a chain that dead-ends and strands every row behind it. Rows under one directory are printed under it once. Each section ends by saying what it left out — how many matching rows did not fit, and how many rows did not mention the words at all — so the reader knows whether to ask again with more words or to open framework_map.md.

As a Claude Code plugin

/plugin marketplace add ngavrish/where-are-we
/plugin install where-are-we@where-are-we

The plugin builds the map at session start, puts its pointer into the session's context, serves all eighteen of the map's tools over MCP (affected, ask, at, callees, callers, context, dead, defines, find, hot, impact, more, path, range, rank, reaches, sections, unreached) and ships nine skills (ask, orient, rank, readmes, spec-map, where-defined, and three for the graph answers: what-to-rerun for affected, reaches and unreached, lines-to-edit for range and path, what-to-look-at for hot and dead). It needs where-are-we on PATH (pipx install where-are-we). Details in plugin/README.md.

Everything it does, and what it is measured to save

Each row names what the tool does and what that is measured to save. Where a number exists it is cited; where none exists the row says so and names the measurement that would settle it. Numbers come from one real run unless said otherwise: a 184-feature behave suite and one production agent run (1,409 turns, $59.65), read from that run's event log. Estimates are marked as estimates.

The map itself

Feature What it does Measured impact If not measured, how to
One tree walk → framework_map.md, framework_map_brief.md, framework_map.json Indexes layers, entry points, routes, data model, public surface, call graph, steps, scenarios, fixtures, CI, duplicates, dead code, every declared name with its line. Declarations are indexed by dedicated regex for Python, TypeScript/JavaScript, Rust, Kotlin, C# and Ruby (a tree-sitter parse tree instead, for the languages it has a grammar for, where pip install "where-are-we[precise]" is present), and by a generic pattern for anything else Build: ~10 s on the 184-feature suite, offline, 0 tokens (README). 75 sections on this repository, 28 on the demo suite
Deterministic output Same tree, same map, byte for byte tests/golden/check.py (CI step golden): two builds of three fixture trees are byte-identical on every CI run; 156 ask cases pinned n/a
schema: where-are-we/1, stable JSON contract (SCHEMA.md) Sections may be added within a major; existing shapes keep Not measurable; a promise Consumers: any agent runner that reads the JSON (find_text, _definitions_for)
Fingerprint (<commit>:<newest mtime in nanoseconds>), content_root and --force A build is skipped when the tree has not moved. The mtime keeps the precision the filesystem reports, so an edit inside the same second as the build before it is still seen, and it covers every file the map indexes rather than a fixed list of extensions, so an edit to a .go, .rs, .kt, .cs, .rb, .java, .yaml, .tf or .proto file is seen too. A skip needs the content root to agree as well as the fingerprint: the root is a sha256 over the sorted (path, hash) pairs of every indexed file, so a rewrite that kept its byte count and had its timestamp put back moves the root and the build runs. --watch asks the same two questions on every tick. Two caveats: the root is compared only when the map already in --out has one, so the first build after upgrading an existing --out from 1.4.x still decides on the fingerprint alone; and the hash of a file is retaken only when its mtime, size or inode change time moved, so on a filesystem where st_ctime is a creation time rather than an inode change time (Windows) the same-size restored-mtime rewrite is missed until something else about the file moves. --force distrusts the hashes as well as the parses, so one command always recomputes the root from the bytes Not measured CI steps idempotent (two builds of the same tree, the second prints unchanged since it was built), an edit in the same second as the build is still seen, a same-size rewrite with a restored mtime is still seen (one-shot and under --watch), and the fingerprint watches exactly the files the map indexes (five languages edited in turn, each one rebuilt)
Incremental rebuild: _cached() around every ast.parse, tree-sitter parse and declaration scan, keyed by path, kind and the sha256 of the file's bytes, with (mtime, size, ctime) as the pre-filter that says whether the hash needs taking again (--force reads nothing from it and rewrites it; WAWE_NO_CACHE=1 neither reads nor writes it) A rebuild only re-parses the files that actually changed since the last build in the same --out, and a change is a change of content: a same-size rewrite with the timestamp put back is re-parsed without --force, which is what the mtime-and-size key could not tell from no change at all. A tree nobody touched still costs one stat per file and no reads, and a tree whose timestamps all moved while its content did not (a restored CI cache, a cp -r, a tar x, a container rebuild) re-parses nothing where it used to re-parse everything Measured 2026-09-07 on this repository at 1.5.0, exported with git archive to /tmp/wawe-self/repo so the figures are reproducible, 260 indexed files: cold build best 2.94 s and median 3.31 s over six runs; warm build, one file touched so the tree has moved, best 1.69 s and median 1.83 s; a build over a tree nothing touched at all is the skip, about 0.2 s. WAWE_DEBUG_PARSES=1 prints parsed 260 files / hashed 260 files cold and parsed 0 files / hashed 1 files warm. parsed N files counts files: the count of computations, about three and a half times larger because one file is asked for its declarations, its symbols, its call graph, its step phrases and its redaction diff, is PARSE_COUNT on the mapper facade. Against 1.4.1 on the same tree, eight alternating pairs so that a drifting machine cannot favour either side: cold, the median paired difference is +0.34 s, 2.78 s before against 3.09 s after at their best. Both sides of a pair are slower than the isolated cold run above, because alternating runs contend for the same disk; the paired difference is the number to read, and it is the content hashing and the one redaction pass, paid once and cached under each file's hash so that no warm build pays either again. Map output identical with and without the cache (tests/golden/check.py, plus a WAWE_NO_CACHE=1 build diffed byte for byte against a cached one, fingerprint excluded)
## This map is incomplete What a bound cut (file count, spec depth) is named at the top of the map Not measured; a correctness feature: an answer of "absent" is never given past a bound Count answers that say "indexed: …" per run
Every artefact written to a temporary and renamed into place Nothing ever reads a half-written map, and a build killed mid-write leaves the previous one whole. Covers framework_map.{json,md,html}, the brief, tags, spec_map.*, semantic_index.{npy,json}, .wawe-cache.json, .pointer-head, and the two files written outside --out: the agent file and .framework-map.json. The start-of-build sweep of a dead writer's temporaries covers what lives in --out, which is everything but those last two: a build knows where its output directory is and does not know where a previous run was told to put an agent file Measured 2026-09-05 on a 3,000-file repository: a reader polling sizes through a build saw 26 zero-byte hits on framework_map.json before and none in 4.2 million reads after; six readers calling defines and ask through a rebuild loop saw 4 unparseable JSON reads and 4 empty defines answers before and none after; SIGTERM at the first touch of the map files left a zero-byte JSON beside a stale .md before, and the previous three files byte for byte after CI step every artefact is replaced, never truncated in place
Every file read into memory is bounded _slurp reads at most 400 KB for a scan and _slurp_source at most 2 MB for a parser, once per (path, limit) and cached, so a build holds the limit rather than the size of what it walks. A file a parser only saw part of is named in the map. The one unbounded read is the content hash, which streams a file in 1 MB blocks and keeps none of it: a bounded read cannot notice a change past the bound, so this is about memory rather than about bytes off the disk, and a large asset is re-read whenever its stat block moves Measured 2026-09-06: a 198 MB .py cost 785 MB of peak RSS and 3.97 s before, 39 MB and 0.19 s after; a 196 MB .ts cost 590 MB and now 29 MB; a .py and a .ts together cost 1171 MB and now 42 MB CI steps a very large file costs the read limit, not its size and a large module is parsed, not silently dropped
An optional source that misbehaves does not end the build --runs-api, git log and git blame are history the map is better with and fine without, so a tracker answering something that is not HTTP, or an author name the locale cannot decode, costs that section and not the run Measured 2026-09-06: a socket answering GARBAGE NOT HTTP exited 1 with a BadStatusLine traceback and wrote no map before, and exits 0 with a map after; LC_ALL=C with the author Renée Müller exited 1 with UnicodeDecodeError before, and after writes blame_owners {'a.py': ['Renée Müller (1)']}, the name exact CI step an optional source that misbehaves does not end the build
A symlink or a pipe in the tree is not read A file whose link resolves outside the repository is skipped, so nothing outside it is copied into a map that gets committed and pasted into prompts, and anything that is not a regular file is skipped, so a FIFO does not block the build forever. A link that stays inside the repository is still followed, and the map says what it left out Measured 2026-09-05: a passwd.py symlinked to /etc/passwd put the whole file into the map's line index before and is absent after; a FIFO named x.py hung the build past a 60 second bound with nothing written before, and after it finishes and indexes its sibling CI steps a symlink out of the repository is not read into the map and a pipe does not hang the build, and dead temporaries are swept
.wawe.toml, .wawe-ignore, WAWE_MAX_FILES, WAWE_JUNIT_DIRS (os.pathsep separated; default: the repository's own reports, test-results, junit, build/test-results, plus /runs, never /tmp) A project states its invocation, exclusions and where its JUnit history lives once. .wawe.toml's [synonyms] table adds a project's own words to --ask's built-in groups Not measured
--product, sibling guessing only for a suite, --product none The application under test is indexed beside its suite; a plain code repository does not index its neighbours Measured 2026-09-03 on this repository: before the fix the map held 231 files of three unrelated sibling repositories (3.1 MB JSON, defines answering with their paths); after, indexed: suite 75, 818 KB
--also Fold other repositories into one map Not measured
--diff What changed since the map already in --out, naming the files whose content hash moved, were added or are gone before the map keys that moved with them. It reads the parse cache and does not write it, so that everything it prints is measured against the map on disk and the same command over the same tree answers the same way twice the pointer names what moved since the last session; CI step "pointer says what changed since the last session" proves it
--watch SECONDS Rebuild whenever the tree moves, by the fingerprint and the content root together, the same two questions a one-shot build asks: a full rebuild each time, writing every artefact a one-shot build writes, and an iteration that raises is printed and the loop carries on Not measured CI step --watch rebuilds whole, writes every file, and survives a failure: twenty files added while watching all reach the map, a deleted name leaves it, framework_map.md and --html are written, and replacing the output directory with a plain file prints rebuild failed, still watching without ending the watcher
--html The brief as a page Not measured
--ctags<out>/tags Every declaration the map holds, in universal-ctags format: name, file, the line as the EX command, kind:, line: and end: where a parser knew it. A row names its file relative to the directory the tags file is in, which is where every ctags reader resolves it from (vim's tagrelative is on by default): with the usual --out . at the repository root that is src/a.py, and with --out .wawe it is ../src/a.py, and both open. Sorted as bytes, so the binary search the header promises works. vim, emacs, helix, kakoune and readtags open it with no server running, on a checkout mounted read only, in a language whose server is not installed. A build without the flag leaves whatever tags is already in --out alone, exactly as a build without --html leaves framework_map.html: the file may be one real ctags wrote, and this tool does not delete files it was not asked to write Measured 2026-09-07 on this repository at 1.5.0, built with --out . at the repository root: 672 rows over 617 names, 54,375 bytes, written in the same build that writes the map. An --out somewhere else writes the same rows with the path that reaches each file from there CI step --ctags writes a sorted, correct tags file: the pseudo tags are present, LC_ALL=C sort -c passes, every row is a declaration site the map holds and every site has a row, charge names app/billing.py:4 and that line declares it, two builds are byte identical, and readtags reads it where the runner can install universal-ctags
--cost [THRESHOLD], --cost --json What each section costs to carry: rows, bytes and tokens, heaviest first, with a total and a threshold that hides the small ones. A section is a ## heading, and the sections are the ones a reader carries: every one in framework_map.md, then every one in framework_map_brief.md beside it whose heading the map does not already have, which is the rule every read composes the two files by. Each is measured on its own file, so its byte count is what wc -c would give for those lines, and the report ends with a measured: line per file whose header plus sections is that file's size. Tokens are bytes / 4 and the output says estimate, because no extra this project has carries a tokenizer that can be reached without downloading a model Measured 2026-09-07 on this repository at 1.5.0: 71 sections over the two files, 34,414 bytes, 8,580 tokens, 409 rows. wc -c agrees: framework_map.md is 1,330 bytes (140 of header, 1,190 in 3 sections) and framework_map_brief.md is 33,446 (222 of header, 33,224 in 68 sections). The heaviest section is ## Defined here at 5,955 bytes and 52 sections are under 500 CI step the map says what each of its sections costs, which splits both files independently and asserts every section against that split, and fails on a build that stops measuring framework_map.md
--export FILE The map as one self-contained file for a channel with no filesystem (a PR comment, a paste): the incompleteness notice, the indexed: counts, every section of framework_map.md and of the brief beside it with what each costs, then the brief itself. --export - writes it to stdout; an empty path is refused Measured 2026-09-07 on this repository at 1.5.0: 37,779 bytes against 1,330 of framework_map.md, 33,446 of the brief and 2,702,888 of framework_map.json CI step the map says what each of its sections costs: the file parses back into the section list --sections prints, in the same order, with the byte counts --cost reports
--init.framework-map.json manifest A starter manifest the map reads stated facts from Not measured

What an agent carries vs what it asks

Feature What it does Measured impact If not measured, how to
The pointer (--pointer, --agent-file) ~600–850 bytes in the prompt naming the map, its sections and how to ask; the map stays on disk Map inlined: ≈ 64k tokens re-sent every turn, 27.4M tokens over one run, a quarter of that run; pointer: ≈ 212 tokens (300× less); a 5-hour allowance gone in 74 min vs the budget going to work (README, measured on one production run)
Orientation replaced by one --ask The first turns of a session stop being ls/find/grep ~40 orientation turns → 1 on the measured suite (README)
--ask / MCP ask: whole rows, ranked sections, honest tail Only rows that mention the words, never a cut row in the answer itself (a more: continuation may cut and mark one row longer than the whole budget, so the chain advances instead of stranding it), limit a strict ceiling, "… N more matching rows (more:rows:...); M rows do not mention these words" - the tail names what was left out and carries the handle more fetches it with Before 0.12: an answer could exceed its limit 68× (3 KB head at limit 50) and cut a row mid-word; after: ≤ limit on every golden case (150), 5/5 fixture checks. .wawe/.wawe-ask.log records every answer; wawe-measure --ask-log .wawe prints median/p95/max tokens
more (MCP), --more HANDLE What an answer left out, by the handle it printed: a section's unshown rows, the rows in it that never matched, definitions past the block's cap, the sections that did not fit, find's hits past its limit. Stateless: the handle names a section, the words and a position, and more recomputes the ranking from the map on disk, so a rebuilt map answers "no such handle in this map" rather than a slice of some other list Measured 2026-09-07 on the suite golden fixture built at /tmp/wawe-cost, question "invoice checkout": chasing the handles from ask(..., 1500) to exhaustion returns all 198 rows the unbudgeted answer holds, in 17 calls; from ask(..., 12000), 198 in 1 call. At 350 characters, 197 of 198 - the missing row is 446 characters long and cannot fit in a 350-character answer, and more says so
--ask synonyms and stemming "login" also searches "signin", "auth"; "invoices" also searches "invoice"; a synonym or a stem scores at half the weight of the literal word, so it never outranks an exact hit; the first line says (also matched: signin, auth) when an expansion found something the literal words did not; .wawe.toml's [synonyms] table adds a project's own words to the built-in groups Not measured
Rows under one directory printed once - \features/checkout/`` then the files Not measured Bytes of an answer before/after on a 40-row directory
## Defined here (defines, _definitions_for) A name → file:line, every declared name in every walked file Not measured as turns saved; the README's claim is one question instead of grep -rn Count Grep calls per session before/after (the run's call events)
defines (MCP), --defines NAME, the map's spans key Every home of a name, not the first one walked: charge: a.py:10-24 (function), b.py:88-91 (function), sorted by file then start, with the kind of each. The end line is exact where a parser knew it (ast for Python, a tree-sitter node with the precise extra) and ? where the pattern table read the language, which sees where a declaration starts and not where it stops. definitions keeps its one-home shape, so nothing that reads it moves Measured 2026-09-07 on this repository's own map at 1.5.0: 617 names over 672 declaration sites, 35 of the names declared in more than one place, and before this those 35 answered with one file chosen by walk order. 104 of the 672 sites carry no end, which is every site the pattern table found. The key is 123,353 bytes of the 2,702,888-byte map on disk, 4.6 percent Count the names your own map declares twice: `jq '[.spans[]
at (MCP), --at FILE:LINE The whole definition enclosing a line, the way a stack trace names it: the innermost one, whole lines, from the map's own line index rather than the file. Cut to the answer budget with a more:at: handle for the rest. A line nothing encloses is answered no definition encloses FILE:LINE with the nearest declarations, never with the wrong function Not measured as turns saved; the claim is one lookup instead of a Read at a guessed offset and a second one when the guess cut the function in half Count Read calls that follow a stack trace before/after
rank (MCP), --rank [FILE,...], the map's rank key What the repository is built around, best first, by PageRank over its own file graph: nodes are files, an edge runs from a file that uses a name to the file that declares it weighted sqrt(uses), and aider's four multipliers apply (x10 a long snake/kebab/camel name, x10 a name the question asked about, x0.1 a leading underscore, x0.1 a name more than five files declare). Given files it is personalised on them, which answers "what should I read given that I am editing these" rather than "what mentions this word". 100 power iterations at damping 0.85, scores rounded to 9 digits before sorting, ties by path: the same order on every machine and both supported Pythons Measured 2026-09-07 on this repository's own map at 1.5.0: the top four are declare.py's find_text at 0.061419466, build.py's build, ask.py's inner build, which is where the answer budget is spent, and eval.py's add. --rank src/where_are_we/ask.py does not change that order and raises find_text from 0.061419466 to 0.080008437, a third, which is the personalisation showing up as a score rather than as a swap. The scores are of a map that indexes this file, so editing it moves their last digits and the order is what the CI step pins. The key is 30,765 bytes of the 2,702,888-byte map on disk, 1.1 percent Compare the top ten to the files you would name yourself: where-are-we --rank --limit 10
--files a.py,b.py, files on MCP ask The files you are working in: inside every section the rows naming one of them are printed first and the rest follow with the section's usual tail and handle. A path or a directory prefix, relative to the repository root, and --files - reads a newline list on stdin, which is what git diff --name-only hands over. A row that names only a basename, which is how both call graphs write a key, is resolved through the files the map indexed. The tail's handle carries the scope, so more reorders the section before it continues and a scoped answer reaches every row an unscoped one does Not measured as turns saved; the claim is that the half of the answer about your own directory is at the top of it rather than under the cut Diff --ask X against --ask X --files <your dir> at a small limit
Cross-file call graph (call_graph_files) Function to callees declared in another file, Python by AST, TypeScript, JavaScript and Go by pattern. An edge is charge (a.ts) where one indexed file declares the callee, or where the caller's own import line says which file it means, and charge (a.ts|c.ts)? where several declare it and nothing says which: the edge names every candidate, sorted, and the question mark says the map is choosing. Three calls are no edge at all: one to a name the calling file declares itself, one made through a module this tree does not declare, which is what ast.walk(...) and from os import path then path.join(...) are, and one on a receiver the function built out of a builtin, which is what out = set() then out.add(key) is. A parameter's default counts as what the function bound, None excepted; *args and **kwargs are a tuple and a dict; a name a nested function never binds is the one written around it; and d.setdefault(k, set()) is a set whatever d is. Two more are read further: a receiver that is a file of this tree carrying a name it does not declare is followed through its own import lines to the file with the def, and self.name(...) goes to the class the method is in or to the one base that declares it, where the file says where that base came from. callers, callees and impact match on the name alone and print the mark as they find it Measured 2026-09-07 on this repository's own map at 1.5.0: 60 keys, 129 edges under them, 1 of which carries the mark, and it is a call on d[k], which is the shape nothing written in the file settles. 1.4.1 mapping the same tree writes the same 60 keys, the same 129 edges and the same 1 mark: the receiver rules that took that count from 20 to 1 are 1.4.1's, and this release does not move it Count Grep calls spent chasing a callee across files before/after
Every edge as a row that says how it was resolved (xrefs) The graph normalised: one row per edge, {subject, edge, object, file, candidates, line, resolution}, sorted by subject, object and line and capped by nothing. Every path in every row is absolute, so the table joins to itself: the file a calls row settled on is the subject of the declares rows for that file. edge is calls (a cross-file call), declares (a spans site, whose candidates is empty because a declaration is not a choice between files) or imports (one import line, with the file it is in and the line it is on, for the package pairs import_graph summarises). A declares row is the site as spans holds it, joined onto nothing: every site is recorded with the absolute path of the root it was walked under, so a page object of an --also root names that root's file and the tags file written from the same key names it too. resolution names the rule that placed the edge and is one of eight values: sole_declarer one indexed file declares the name, import_line the caller's own import line named the file, receiver_import the module the receiver was bound to named it, re_export a facade's own import lines were followed to the file with the def, base_class a self.name(...) call reached the one base that declares it, declaration the row is the declaration site, import_statement the row is the import line, and ambiguous several files declare the name and nothing said which, which is the ? the rendering carries. call_graph_files is derived from the calls rows by the renderer that writes it, and holds less than they do: cut to 8 callees a key and 60 keys, and keyed by basename, so where two files of one basename declare one function name they share a key and the file the walk read last owns it while both keep their rows here. There is no local value and no overrides edge: a call inside the file that declares the callee never becomes an edge, and no rule computes an override Measured 2026-09-07 on this repository's own map at 1.5.0: 893 rows, 221 of them calls (191 sole_declarer, 11 re_export, 7 receiver_import, 5 import_line, 7 ambiguous) and 672 declarations; call_graph_stats counts 214 edges and 7 marked, which is the same 221 rows counted the other way. The key is 273,961 bytes of the 2,702,888-byte map on disk, 10.1 percent, of which the declares rows are 182,558 of the 256,974-byte table. On the suite golden fixture, built at a fixed /tmp/wawe-cost so the number is reproducible, 75,821 bytes of 182,322, 41.6 percent: the share is larger because that map is small and 224 of its 265 rows are declarations. The whole release against 1.4.1 on the same tree is 2,277,523 to 2,702,888 bytes, 19 percent, for spans, rank, content_root and this; on the suite fixture the same two releases are 54,317 to 182,322 bytes, which is more than a tripling, because a small map is mostly the four keys this release adds and a large one is mostly its lines Read xrefs instead of parsing the pipes and the ? back out of a rendered edge
wawe-eval --map OUT --graph, call_graph_stats Per language group: how first-party a tree's calls are (callee names looked at, names some indexed file declares, names several declare) and what the graph came to (cross-file edges written, and how many carry a ?). --json carries it. With the precise extra it parses the TypeScript, JavaScript and Go again with tree-sitter and prints the delta Measured 2026-09-07 at 1.5.0, fixtures built at /tmp/wawe-cost. suite fixture: python 85 sites, 40 first-party, 40 edges, rate 0.4706. code fixture: python 7 sites, 0 first-party, 0 edges. poly fixture: ts_js 3/3 with 1 edge and go 5/5 with 0, rate 1.0 each. This repository: python 2818 sites, 737 first-party, rate 0.2615, 123 ambiguous names (share 0.0436), 214 edges and 7 of them marked; ts_js 2/2; go 3 sites, 2 first-party, rate 0.6667. The tree-sitter comparison line needs the precise extra, which is not installed in the environment these figures were taken in; without it --graph prints regex vs tree-sitter: not compared, and with it, measured 2026-09-05, tree-sitter saw 1 and 2 sites for the poly fixture's same code The row above is the measurement
callers (MCP), --callers, "Called by" in ask Who calls a name, the other direction of the call graph: every file:func that mentions it, exact and case-sensitive Not measured as turns saved; the claim is one lookup instead of grepping every file for a call site Count Grep calls spent finding call sites before/after
callees (MCP), --callees What a name calls, the other direction of callers: every callee with the file it is defined in, from the same two graphs, cross-file only Not measured as turns saved; the claim is one lookup instead of reading the function to find out what it reaches Count Read calls spent opening a function to list its calls before/after
impact (MCP), --impact NAME [--impact-depth N] The blast radius of a name: every file:func that reaches it within N hops (1 to 6, 3 by default), grouped by hop and sorted inside each. A visited set walks a cycle once, and the keys defining the name itself are the change rather than its radius. The first line of every reply states the rules the answer was built under, unconditionally: hops are followed by name, so where several files define one name their callers are unioned; only cross-file calls are in the graph; and the map keeps at most 60 cross-file and 120 step graph keys, so on a large repository the radius is a floor. Under each hop, where the map holds xrefs, a how: line names in the same order the rule that placed each of those edges, and step_graph for a hop through the behave step graph, which records a bare name and no rule. Where the key data shows the clash a note: names up to five of the keys and how many more. Capped at 200 file:func entries in the whole reply, note keys included, with a line naming how many are left and at which depth; no handle to fetch them, because the tool already takes a depth and narrowing is the reader's move. A depth outside 1 to 6 is refused: exit 2 on the command line, JSON-RPC -32602 on the tool Not measured as turns saved; the claim is one lookup instead of running callers outward by hand, hop after hop Count callers calls per session before/after
context (MCP), --context NAME Everything the map holds about one name, in one answer: where it is declared with every span, the rows of the map that mention it, its callers, its callees, and its impact one hop out. The five calls an agent used to make on landing on a name, composed from the same functions over the same map. The budget is allocated in two passes: every block is given the smaller of what it needs and its floor share - 15% declared, 35% map rows, 15% callers, 15% callees, 20% impact - and what nobody claimed is then handed on in that same order to the blocks still short. So when the five answers together fit the budget every one of them is printed whole, and the same name at the same budget is always the same answer. Whole rows; a block that could not print itself ends in a tail carrying a more:ctx: handle the more tool resolves Measured 2026-09-08 at 1.6.0 on the suite golden fixture, all 224 declared names at 12000 characters: 182 answers hold every row --defines, --ask, --callers, --callees and --impact --impact-depth 1 return, and the other 42, whose five answers do not fit 12000 at all, have every cut row behind a handle. wawe-eval --tool context over 100 names reports recall with handles 1.0 at 1500 and 12000, and 0.6984 and 0.9943 for the first answer alone at those two budgets Count the tool calls a session spends on one name before/after (the run's call events)
affected (MCP), --affected FILE[,FILE], --changed [REF], --affected-format, --affected-depth N Which tests a change reaches, from the graph the map already holds: the scenarios whose steps call into those files, the feature files they are in, the routes and page objects reached, and the files no xrefs row names, which is the answer saying what it does not cover. The walk starts at every name the changed files declare and follows the calls rows upward, callee to caller, to a depth of 6 by default and 12 at most, with a visited set so a cycle is walked once. A step function is the join of two spans sites on one line, the function and the phrase its decorator binds; a scenario is reached when one of its own step lines holds that phrase, normalised and cut to 40 characters, which is the same substring rule feature_links is built with and which a step phrase written as a regular expression can miss. A changed feature file is a test rather than something a test calls, so every scenario in it is affected, at no hops, and the row says so. A route is named when the file it is served from is reached: routes_served records the basename of that file and no handler, so the granularity is the file. --changed asks git what moved in the repository the map was built from, so a pipeline passes nothing. --affected-format behave prints one --name per affected scenario, anchored on its name, and nothing else: --name is the only behave option that unions, -i is a plain store that also decides which files are collected, and a tag is never emitted because behave applies --tags per scenario while the map records a feature file's tags as every @word anywhere in it, an email address in a step line included. Past 200 scenarios each feature file's names alternate inside one --name, so a large selection is eight arguments rather than two thousand. A selection is a machine artefact rather than prose, so it is exempt from the answer budget: printed whole, with no tail and no handle, however large. --affected-out FILE writes it there and prints the blocks a person reads to stdout, so a pipeline parses a file it asked for rather than an answer it has to cut a head off. That file is empty, zero bytes, when the change reaches nothing and when nothing changed at all, so an empty file means run nothing and never the previous run's selection. Test for it: xargs given empty input runs behave with no arguments, which is the whole suite. The MCP tool is bounded like every other tool this server declares: a selection that fits its 12000 character reply comes back whole, and one that does not comes back as its count, its first 20 selectors and the --affected-out command that writes all of it. Two scenarios of one name in one file are one argument and behave runs both, which over-selects rather than under: the block above names each of them with its own line. pytest prints node ids. Whole rows, a floor share of the budget per block, and a more:aff: handle on every block that was cut; a block there is no room for at all is not dropped silently but printed as … 2 rows in pages; raise the budget with the handle that fetches them Measured 2026-09-07 on a three-scenario behave suite (CI step a change names the scenarios that reach it): a change to the checkout page object names 2 of 3 scenarios at 1 hop, the product function under it names the same two at 2 hops and none at depth 1, and the other page object names the third. behave itself is run on the printed selection in that step: behave --dry-run with it reports 2 untested and 1 skipped, and the other way round for the other page object. Three more suites in the same step run the selection across feature files: 24 of 24 over 8 files, 3 of 4 where one file is wholly affected and another partly, and 240 of 240 through the 8 combined patterns the large form prints. A fourth, 80 feature files of 30 scenarios, checks the exemption: 2400 of 2400 selected, 80 patterns and 56 KB of them, whole on stdout and in the file, with behave running every one. The expected set is recomputed in the step by the opposite traversal and asserted equal. On the suite golden fixture the answer is 0 of 5 scenarios, and the first line says why: 3 of those 5 hold no step any module in that fixture binds. Turns saved not measured Count the scenarios a session re-runs after a one-file change, before and after
reaches (MCP), --reaches NAME Which tests reach one product function or class: the other direction of affected, walked from that one name up the same calls rows with no depth cap, because the question is whether anything reaches it at all. The answer names where the name is declared with each span and, for a class, the members walked; the scenarios that reach it grouped by feature file; the pytest cases that reach it; and the routes reached. Under the first scenario of each feature file it prints the chain of calls, steps/shop_steps.py:step_pay -> pages/checkout.py:pay -> app/billing.py:charge, as a line of its own, so the reason is on the page rather than in a second question. A class is answered by what is declared inside it, found two ways: the dotted name spans writes beside a method (CheckoutPage.pay), and containment in the class's range. Neither is available when the parser did not find where the class stops and the language has no dotted spelling, and then the first line says the counts are a floor rather than answering nothing in silence. Where the name is declared under a product root of its own the first line says that too: the call graph is extracted from the files under --repo and from no other root, so no edge crosses into it. Whole rows, a floor share per block, and a more:rch: handle on every block that was cut Measured 2026-09-07 on the CI step the graph says what the tests reach and what they do not: the product function under the checkout page object is reached by the two scenarios whose steps call it and by neither of the third; a function covered only by a pytest case is reached by that case; the class is reached through its three methods; the same class with an unparseable file answers with the floor sentence; an unknown name answers no declaration of 'x' in the map Count the greps a session runs to decide whether a function is covered
unreached (MCP), --unreached [--limit N] What no test reaches: every product function and class with no call path up to a behave step function, a pytest case pytest_tests names, or a declaration in the files another runner holds its cases in (js_tests, api_tests, perf_suites, other_suites, more_suites, whose cases the map records by title rather than by function, so every declaration in them counts). Grouped by file and ranked by the map's own rank, best first. One walk down from the entry points rather than one walk up per definition. Product is every file that declares something and that the map does not name as suite, read off the keys the map already writes, so a product under a root of its own and a product beside its suite are answered the same way. The precision of that split is the precision of the map's own suite heuristics: a file the mapper miscalls a page object is not product here, is not in the count and is not in the list, so the answer prints the files it set aside under a head of their own. A definition named main, __main__ or __init__, and any file on a test path, are left out. That is a narrower list than dead's, and the two are two rules on purpose: dead also drops every other dunder, every name starting test and every step function, which is right for a question about what nothing calls, where a name the runtime or the runner calls is not an answer. Here every name a rule drops leaves the denominator with it, so the same widening would report better coverage of a smaller product. Each answer's ## How this was counted block prints its own list. The first line states the graph's own resolution rate from call_graph_stats, because a graph that placed half its callee names calls half the product unreached whatever the suite covers; where the product is under a root of its own it also says that no calls row crosses into it, which is a limit of the map rather than a fact about the tests. A map with no step function and no test case says there is nothing to reach from and prints no list Measured 2026-09-07 in the same CI step: on a suite whose steps reach charge, whose pytest case reaches audit and where nothing reaches refund, refund is the one definition listed, under a first line reading The graph resolved 8 of 16 callee names (50 percent); on the code golden fixture, which has no suite, the answer is that there is nothing to reach from; on a PRODUCT_SRC tree the whole product is listed and the first line says why Count the definitions a session reads to find what the suite does not cover
path (MCP), --path A,B [--path-depth N] The shortest call chain from A to B over the map's own xrefs calls rows, printed one hop per line with the rule that placed each edge and the line the call site is on. Breadth first, the rows out of a node sorted by callee, call-site line and resolution and every frontier sorted, so the chain printed is the same chain on every run; a visited set carries across hops, so a graph with a cycle in it terminates and the first chain found is the shortest. An ambiguous edge is followed into every candidate rather than guessed at, and its hop names every file that declares the callee and says which of them the chain took. Each end is a name, or FILE:NAME where several files declare it, matched by the repo-relative path, a suffix of it or the basename. Only cross-file calls are in the graph, which the first line says out loud, because it is the reason a chain a reader can see in the source can be missing here; the other reason is the depth, 1 to 12 and 6 by default. Where there is no chain the answer names the frontier the walk stopped at, so "no path" says how far it got. Whole rows, the answer budget, a more:pth: handle under a block that was cut Measured 2026-09-07 in the CI step path shows one chain, with how each hop was resolved: on the 1.3.0 three-hop fixture a reaches d in three hops with sole_declarer on each, a cycle of three terminates and prints two hops, and an ambiguous edge prints both files that declare the callee. Turns saved not measured; the claim is one call instead of walking callers or callees outward by hand Count callers/callees calls per session before and after
range (MCP), --range NAME Every home of a name as file:start-end kind, the text of the shortest of them, and the three line numbers an editor anchors on: the first line of the definition, the last, and the one after the last, each printed with the text of that line so an Edit can match on the text rather than trust a number. The shortest site is the one whose text is printed, because a name declared twice is usually a small real one and a large one that shadows it; a site whose end nothing measured cannot be the shortest, and is listed with ? and the reason it is unknown, which says whether the read limit is ruled out. One caveat, and the first line repeats it whenever it applies: a line holding [redacted] had a value that looked like a secret written over on the way into the map, so it is not the line on disk, the anchor row says so, and an edit anchored there will not match. The file is not read back to recover it: nothing in that module opens a source file, and a map is often read far from the tree it was built from. Reads only, and writes nothing. Whole rows, the answer budget, a more:rng: handle on the text block, which is the one that grows, and the rest of the rules in ## How this was counted Measured 2026-09-07 in the CI step range hands an editor the lines it needs: a name with two homes prints both; sed -n on the printed end and the line after it shows the next line is dedented or blank on every site of the fixture; a .go declaration the pattern table read prints ? with the reason Count Read calls with a guessed offset per session before and after
dead (MCP), --dead [--limit N] A list of questions, not a list of dead code, and the answer's first line says that before it says anything else. On a library most rows are calls this map could not place rather than definitions nothing calls: only cross-file calls are in the graph, so a call inside the file that declares the callee leaves no row, and neither does a call through an imported module (extract.Ctx(...), tests.coverage_by_file). It is sharp on a test suite, where a page object is called from step modules, and blunt on a library: 387 of 556 on this repository at 1.6.0, and 359 of 517 at 1.5.0, of which a spot check of that list found one genuinely dead. Both counts move with the tree: they are what this repository looked like at those two releases, not a property of the rule. The first line is the counts and the one sentence that decides what they mean, under two hundred characters so it survives being re-read on every turn; every other rule is a row of ## How this was counted, the block the budget may cut, which is the shape unreached prints under too. What it is: the definitions no xrefs calls row lands on, grouped by file, one file per row so a block cut to a budget loses whole files rather than a header with its names below it. One def is one row whatever it is spelled as: spans holds a class as LoginPage and as class LoginPage, and a method as LoginPage.sign_in and sign_in, and a call on any spelling keeps the site off the list. Only counted in the file kinds this map's call graph actually reaches, read off the table rather than hard coded, because a def quoted inside a Markdown fence is documentation and not dead code. ## How this was counted states the whole exclusion list (a dunder, main, a name starting with test and every pytest_tests case, a step function, a definition in a file a route is served from or that entry_points names). routes_served records a basename and no path, so where two files share one the map cannot say which serves the route and both are excluded; a row of that block counts the basenames it guessed on. A map whose graph holds no calls row at all says No call graph in this map rather than reporting a clean nothing Measured 2026-09-07 in the CI step dead and hot read the graph and the history: on a fixture a function called from nowhere is listed, one referenced only from a route file is not, and a step function is not. The map's own "Page-object methods nothing calls" section is a different rule over a different table and is unchanged: the golden expected/ is byte for byte what it was Count the definitions a review reads that nothing reaches
hot (MCP), --hot [--limit N] The definitions with the most of the codebase behind them that also change the most: the map's own rank score times the commits its most-changed-files section counted, top N with both numbers printed, so a reader can see which of the two put a row where it is. That count is the map's git_commits key, added in this release for exactly this: git_history keeps at most five commit lines a file, so counting those lines weighed every busy file the same and the multiplier reordered nothing. Two bounds. The first line carries the one that decides what the ranking is, the way the impact row states its own: the most-changed section is the forty busiest files, so a file outside it counts 1 however often it changed and the ranking is rank's own order for those. The second is a row of ## How this was counted: a map built before git_commits can only count commit lines, and there the block says so and every count at the cap prints 5+. A merge commit names no file in the log that section is built from and is not counted. rank is the map's top 200, so this ranks within those, and only the file kinds the map's call graph reaches, which is the rule dead applies for the same reason: a def quoted inside a Markdown fence is documentation, and the two most-committed files in a repository are usually its README and its changelog, so without the filter a reader asking where to look first was told to read documentation third. On this repository 15 of the 200 are left out that way and the block says how many. Ties by path, then line, then name, so two builds of one tree print the same order. Same shape as the others: a first line under two hundred characters, and the rules in ## How this was counted Measured 2026-09-08 at 1.6.0 in the CI step dead and hot read the graph and the history: on a fixture of two files with equal rank, the one with ten commits is printed above the one with one, and the product beside each row is the score times that count; on this repository's own map the top three rows are src/where_are_we/ask.py:2650 context, src/where_are_we/ask.py:867 build and src/where_are_we/_mapper/build.py:307 build, and no row is .md. The products themselves move with every commit, because both halves do, which is why the CI step asserts the kinds and the order rather than the numbers Count the files a review opens before it finds the one that matters
find (MCP) Where a phrase or string lives, with the line Not measured Same
sections (MCP), --sections The headings, now map + brief (75 vs 3 before 0.12.1) Measured 2026-09-03: a code repository's --sections went from 3 empty suite headings to 75
wawe-eval Generates questions from the map, asks each at no budget to get the rows the map holds for it, and reports what a budgeted answer shows of those, three ways: macro (per question), pooled (all rows), and over the five rows that ranked highest, plus a count of the rows too long to print at that budget at all. With more in the build it also reports what the answer plus its handles reaches Measured 2026-09-08 at 1.6.0 on the suite fixture built at /tmp/wawe-cost, 100 questions, seed 0: first-answer recall 0.4292 / 0.8392 / 0.9996 at budgets of 350 / 1500 / 12000 characters; pooled 0.0886 / 0.2884 / 0.997; top-5 0.6317 / 1.0 / 1.0; 34 / 0 / 0 rows longer than the budget; mean answer 303.0 / 973.3 / 2220.0 characters. Recall with handles, over the rows that fit, is asserted 1.0 by the CI step wawe-eval: the budget loses no row the map holds, which exits 1 on any such row a handle fails to return, from the release that adds more onwards --agent compares the map tools against grep and read on the same questions; not run in CI
`--for author coder, --only, --skip, --max-lines` A brief tailored to who reads it; capped per section Not measured in tokens; the per-section cap keeps every head (3×50 rows at 30 lines → every head present, before: the last sections dropped)

The other map

Feature What it does Measured impact If not measured, how to
--specs, --spec-cmd, --spec-source, --spec-depth, --spec-limitspec_map.md/json A ticket and its links two hops out, from any command that returns JSON, or from --spec-source github|linear with no command to write Not measured Turns spent fetching tracker pages before/after
ask over both maps One question, answers from code and spec Not measured

Semantic answers (optional extra)

Feature What it does Measured impact If not measured, how to
pip install "where-are-we[semantic]" — embedding index, "Related by meaning" tail Keyword hits plus nearest paragraphs by meaning Not measured for answer quality An A/B of questions with and without the tail
--corpus NAME=PATH External corpora (rules, runbooks) in the same index Not measured
WAWE_EMBED_CACHE Embeddings cached across builds in one sqlite file Full five-corpus build: 6 min → about 30 s, five and a half minutes were recomputing unchanged vectors (CHANGELOG 0.11.0)
WAWE_EMBED_MODEL, WAWE_RERANK_MODEL, --no-semantic Model choice; skip the index Not measured

Docs the repository is missing

Feature What it does Measured impact If not measured, how to
`where-are-we --docs plan write` Drafts a README per directory that lacks one, from the map's facts, with a TODO: for the purpose only a human knows Not measured
wawe-readmes The same drafts as one command (both call readmes.describe; --docs is the map-aware wrapper): --repo, --write, --help; lists by default, writes only with --write Fixed in 0.12.3: until then the entry point parsed no arguments and wrote into $AGENT_REPO unasked

Ways in

See it on a repository you know: FastAPI 0.115.0 mapped.

Feature What it does Measured impact If not measured, how to
CLI where-are-we Everything above
MCP server (--mcp): ask, find, defines, sections The map as tools, stdio On the measured production run: 194 map calls vs 68 repository searches (call events of that run)
LSP server (--lsp): textDocument/definition, workspace/symbol The map as an editor's language server, Content-Length framed stdio Not measured -
Library: build, brief, digest, init_manifest, main Python API Not measured
GitHub Action (ngavrish/where-are-we@v1): inputs repo, product, out, agent-file, comment; outputs brief, summary Map on CI, optional PR comment Not measured
pre-commit hook Rebuild on commit so a map is never stale Not measured
`--install-hook git claude cursor codex
Claude Code plugin (/plugin marketplace add ngavrish/where-are-we) SessionStart builds .wawe/ and hands the session the pointer; the eighteen tools over MCP (affected, ask, at, callees, callers, context, dead, defines, find, hot, impact, more, path, range, rank, reaches, sections, unreached); skills orient, ask, rank, where-defined, spec-map, readmes, what-to-rerun, lines-to-edit, what-to-look-at; WAWE_STRICT=1 refuses repository searches. Installed from the marketplace the tools are named mcp__plugin_where-are-we_where-are-we__{affected,ask,at,callees,callers,context,dead,defines,find,hot,impact,more,path,range,rank,reaches,sections,unreached}; under --plugin-dir the prefix differs, so prompts name the server where-are-we, not the prefix Verified 2026-09-03 in a fresh repository: hook built the map, tools answered, pointer reached the context. Turns saved not measured Sessions with vs without the plugin: Grep/Glob/Bash grep counts
Packages: PyPI wheel + sdist, deb (apt repo with key), rpm, Homebrew tap, GitHub release with SBOM (SPDX) and sigstore signatures Install anywhere

Honesty features (not savings, guarantees)

Feature Guarantee
no match for … indexed: product N files, suite M files An absence names what was searched; it is not a claim the thing does not exist
## This map is incomplete Every bound that cut is written at the top
Whole rows, a ceiling, a tail An answer never pretends to be complete: what was left out is counted
The map says what it maps "a map of this repository" vs "of this suite and the product it tests", from the map's own counts (0.12.2)
A map file is never torn Every artefact is written to a temporary and renamed into place; a reader sees the previous map or the new one, never a half-written one, and a build killed mid-way leaves the previous map intact (CI step every artefact is replaced, never truncated in place)
The servers stay up MCP and LSP answer malformed params, arguments or limit with a JSON-RPC error and keep serving; both exit quietly when stdout closes (CI steps mcp malformed params..., lsp malformed params..., mcp and lsp exit 0 quietly when stdout closes early)
--html escapes repository content A docstring or a file name holding markup renders as text on the page (CI step --html escapes repository content instead of interpolating it)
--install-hook is one unit Every target is checked before any is written; a refusal installs nothing and names its cause, a rerun finishes the job (CI step install-hook git refuses a symlinked hook file)
A guard can tell a read from a write --effects ships the class of every flag this tool's parser knows, --effects -- <command line> classifies one line without running it, and --dry-run names every path a write would touch and touches none of them. The table cannot drift from the parser or from the effects.json in the wheel: a CI step compares all three (CI steps every flag the parser knows is in the effects table, and nothing else is, a command line says what it would write before it runs, --dry-run names every path it would write and writes none of them)
A source in UTF-16 is read A file with a byte order mark is decoded, indexed and answerable; a binary that merely starts with one is not (CI step a UTF-16 source file is decoded, indexed and answerable)
An edge the map is guessing at says so A cross-file callee several files declare is written charge (a.ts|c.ts)?: every candidate, sorted, and a question mark, which is the map declining to pass a choice off as a lookup. A call to a name the calling file declares is no edge at all, and an import that names one file settles it without the mark. The mark is about the name: a call through a variable to a name one file declares is written plain, because what the map does not know there is the receiver rather than the name. Every impact reply states the rule (CI step an ambiguous edge names every file that declares the callee, and is matched without the mark). An unmarked edge covers two different degrees of certainty, so the map's xrefs key says which rule settled every edge, marked or not, and impact prints it under each hop (CI step every edge says how it was resolved)
The graph says what it looked at and what it wrote call_graph_stats counts, per language, the callee names the walk looked at, how many of them some indexed file declares and how many several declare, which is a measure of the tree rather than of the graph, and beside them the cross-file edges written and the ones written with a ?, which is the graph. Over the whole walk, not the 60 keys that survive the cap. wawe-eval --map OUT --graph prints them, and with tree-sitter installed prints what a real parse makes of the same code beside them (CI step wawe-eval --graph: the map says how much of its call tree it resolved)

How to measure it on your own sessions

wawe-measure reads Claude Code's own transcripts (~/.claude/projects/<project>/<session>.jsonl) and counts, per session, how many turns were spent looking around versus doing something else:

pip install where-are-we
wawe-measure --since 2026-09-01           # table, one row per session, a median row
wawe-measure --since 2026-09-01 --json    # the same rows as JSON
wawe-measure --sessions /path/to/jsonls   # a directory of transcripts instead of ~/.claude/projects

Definitions:

  • A turn is one assistant message.
  • A search is a Grep or Glob tool call, or a Bash call whose command starts with (after an optional cd ... &&) grep, rg, find, ls, ag, ack, fd or tree.
  • A map call is a tool call whose name contains where-are-we (the MCP tools) or a Bash call whose command contains where-are-we --.
  • orientation_turns is how many turns went by before the agent did something other than look around (an edit, a write, a test run): the count of turns before the first turn with a non-search, non-read, non-map tool call.

Measured 2026-09-05, wawe-measure --since 2026-09-01 against 30 sessions on this machine (one developer, several projects, not a controlled run):

sessions median searches median orientation_turns
with a map call 5 5 3
without a map call 25 1 2

Five sessions used the map at all in this window, and those five ran longer and searched more, not less: this is one developer's mixed transcripts, not a before/after comparison, and settling the "orientation replaced by one --ask" claim above needs matched sessions on the same task, one with the map and one without.

Does the budget lose answers

Every answer ask gives is cut to a byte budget. wawe-eval measures what that cut costs. It generates questions from the map itself (every declared name, every step phrase's distinctive word, the longest word of every section heading), asks each one at no budget at all to get the rows the map holds for it, and then asks it again at each budget.

First-answer recall is the share of the rows of the full answer that the budgeted answer shows before any more, averaged over the questions.

wawe-eval --map .wawe --questions 100 --budgets 350,1500,12000

Two averages are published, because they answer different questions and neither is "the" recall:

  • first-answer recall is the macro average: each question's recall is computed, then those are averaged, so a question with two rows counts as much as a question with three hundred. This is the number for "what does a typical question lose".
  • pooled recall is the micro average: all rows shown over all rows there were, so the biggest questions dominate. This is the number for "what share of everything the map could have said was said".
  • top-5 recall is rank aware: ask ranks what it shows, so this is the share of the first five rows of the full answer that survived the cut, averaged per question. It is the number for "was the answer at the top still there".
  • rows over budget is not a recall at all. It counts reference rows longer than the whole budget, which nothing can print at that budget: not the first answer, and not more, which reports such a row rather than skipping it. On the suite fixture one 446 character row does this, and it turns up in 34 of the 100 questions at 350 bytes and in none at 1500. These rows count against first-answer recall, because a reader who asked for them did not get them; they are left out of the recall the exit code asserts, because that one is about rows a handle could have returned.

Measured 2026-09-08 at 1.6.0 on the three golden fixtures with more in the build, 100 questions per fixture (fewer where the map has fewer), seed 0, built under the fixed root /tmp/wawe-eval the CI step uses:

fixture questions budget first-answer recall pooled recall top-5 recall recall with handles rows over budget mean bytes
suite 100 of 285 350 0.429 0.089 0.632 0.981 34 303
suite 100 of 285 1500 0.839 0.288 1.000 1.000 0 973
suite 100 of 285 12000 0.9996 0.997 1.000 1.000 0 2220
code 12 of 23 350 0.756 0.585 0.785 1.000 0 263
code 12 of 23 1500 1.000 1.000 1.000 1.000 0 413
poly 10 of 21 350 0.733 0.656 0.753 0.983 0 277
poly 10 of 21 1500 1.000 1.000 1.000 1.000 0 367

Read the suite row at 350 bytes together: a 303 byte answer holds under half of what a typical question could have said and a tenth of every row across all of them, and about two thirds of the five rows that ranked highest. That is the shape of the cut. It is not a claim that nothing was lost.

The claim that nothing is lost belongs to more, the tool that fetches what a tail line says was left out. recall_with_handles counts the rows that fit the budget after every more: handle in the answer has been followed, and every handle in the replies after that. At 1500 bytes, the MCP server's floor, and above, it is 1.000 on every fixture: the budget cuts the first answer, the handles give all of it back. At 350 it is 0.98: an answer that small cannot always hold a row and the handle that points at the rest, so wawe-eval asserts 1.0 only from --assert-from (default 1500) and prints the smaller budgets. The CI step wawe-eval: the budget loses no row the map holds is that exit code.

How first-party a tree's calls are, and how big the graph came out

The cross-file call graph is built by name: a callee is placed in the files the walk saw declare it. call_graph_stats, a top-level key of framework_map.json, counts what that came to, per language group.

Three of its counters describe the tree that was read. sites is the callee names the walk looked at, one per function per distinct name; resolved is how many of those names some indexed file declares; ambiguous is how many several files declare. resolution_rate is resolved / sites, so it is the share of what a function calls that is first-party: Python builtins, string and dict methods and every standard library name sit in the denominator and can never leave it, and a call to a name the caller's own file declares sits in the numerator without ever becoming an edge. It is a fact about the code, not a score the graph can raise by resolving more.

Two more describe the graph that was written. edges is the cross-file edges the walk wrote, and marked how many of them carry a ? because several files declare the callee and nothing in the caller says which is meant. Those are the numbers that move when the graph gets better or worse.

All five are taken over the whole walk, before call_graph_files is cut to its 60 keys, so they measure the walk rather than the cap.

wawe-eval --map .wawe --graph          # a table
wawe-eval --map .wawe --graph --json   # the same numbers as JSON

Measured 2026-09-07 at 1.5.0, the three golden fixtures built under /tmp/graphfix and this repository, exported with git archive to /tmp/wawe-self/repo, mapped with --product none --no-semantic:

map language sites resolved ambiguous edges marked resolution rate ambiguous share
suite fixture python 85 40 0 40 0 0.4706 0.0
code fixture python 7 0 0 0 0 0.0 0.0
poly fixture ts_js 3 3 0 1 0 1.0 0.0
poly fixture go 5 5 0 0 0 1.0 0.0
where-are-we python 2818 737 123 214 7 0.2615 0.0436
where-are-we ts_js 2 2 0 0 0 1.0 0.0
where-are-we go 3 2 0 0 0 0.6667 0.0

The code fixture's 0.0 is the honest reading of a repository whose only calls are into argparse and a decorator: nothing it calls is declared in it, so the graph has no edges and the rate is 0. Read the two halves together, though. The poly fixture scores 1.0 in both languages and holds one edge, because a rate of 1.0 says every name it calls is declared nearby, not that a graph was drawn; edges is the half that says a graph was drawn.

This repository's 0.2615 is what a Python tree looks like when most of what a function calls is a builtin, a method on an object or a standard library name. Of its 2818 sites, 727 are Python builtins and 847 are str, list, dict and set method names no file here declares, so 1,574 of the 2,818, more than half the denominator, is unreachable by any name graph; and 516 of the 737 first-party names never become a cross-file edge, most of them a call to a name the calling file declares itself, which the cross-file graph does not hold either. The graph under it is 214 edges, 7 of them marked.

Of the 60 keys the map keeps, 1 edge carries the mark. It is pool[kind] .add(word) in eval.py, a call on a subscript: nothing written in that file says what the container holds, so the map names both files with a def add and says it is choosing. 1.4.1 mapping this same tree writes the same 60 keys, the same 129 edges and the same 1 mark, so that count is 1.4.1's receiver rules holding rather than anything this release changed; the release before them wrote 20 marks over the same 60 keys. Some of those edges did not vanish, they were corrected: the edge hooks.py:_ensure_map used to carry for build named two candidate files, and is now build (build.py), read off mapper.py's own from ._mapper.build import build. The rest are gone from the graph rather than corrected: out.add(key) in a function that wrote out = set() two lines up is set.add, and no def add in this tree is what runs.

With pip install "where-are-we[precise]" the same numbers are computed again for the TypeScript, JavaScript and Go from a tree-sitter parse and printed beside the pattern pass:

regex vs tree-sitter, go: sites 5 vs 2 (-3), resolved 5 vs 2 (-3), rate 1.0 vs 1.0 (+0.0)
regex vs tree-sitter, ts_js: sites 3 vs 1 (-2), resolved 3 vs 1 (-2), rate 1.0 vs 1.0 (+0.0)

That is the poly fixture, and the gap is the pattern pass admitting what it over-counted: its body scan starts at the signature line, so a function's own name reads as a call, and if ( reads as one too. Without the extra the line says so and the run carries on.

--agent is the other half: the same questions asked through the Claude API twice, once with the map tools and once with grep and read over the repository, scored against answers taken from framework_map.json. It needs ANTHROPIC_API_KEY and pip install "where-are-we[eval-agent]", refuses cleanly without them, costs money, and is not run in CI. How to read the table and how to run the A/B: docs/examples/eval.md.

Command line

A CLI is the tool. pip install where-are-we gives two commands:

  • where-are-we — build the map and answer from it.
  • wawe-readmes — offer a repo the docs it is missing.
where-are-we --repo . --agent-file AGENTS.md   # build the map, drop a pointer
where-are-we --ask "refund settled invoice"    # answer from an existing map
where-are-we --install-hook git                # rebuild on checkout/merge/commit

Every flag is under All options below; the map also answers over MCP (--mcp) and as a library.

Where is it defined

$ where-are-we --ask "MAX_PERSISTED_FORECAST_RESULTS"

## Defined here

- `MAX_PERSISTED_FORECAST_RESULTS` — src/constants/forecastStorage.ts:31

Every name in every file the walk reaches, with its line — functions, classes, constants, types, step phrases, scenario names. A question about a name is a question about where it is, and an answer without the line sends the reader to grep for it anyway.

A name may have more than one home, and --defines names all of them with the line each one ends on:

$ where-are-we --defines charge

charge: billing/core.py:10-24 (function), legacy/pay.py:88-91 (function)

An end of ? is a declaration this map could not measure: a language the pattern table read, which has seen the line a declaration starts on and nothing that says where it stops, or a file over a megabyte, of which the parser was handed the first megabyte and so never saw the end. A guessed end would be worse than none for anyone editing by anchor. Installing the precise extra turns ? into a line for TypeScript, JavaScript, Go, Rust, Kotlin, C# and Ruby, and changes nothing about which names are declared where.

Upgrading to 1.5.0 does not add spans to the map you already have: a build skips a tree that has not moved. Run it once with --force after upgrading, or wait for the next commit. The other direction of the same index is --at, which takes the file:line a stack trace hands you and prints the definition around it, so the next step is not a Read at a guessed offset:

$ where-are-we --at billing/core.py:15

billing/core.py:10-24 charge (function)
def charge(amount):
    ...

When a name is not there, the answer says what was indexed rather than declaring the absence real. A map that overstates its reach turns "I did not look" into "it is not there".

Both of those answer "where is this name". The question before it is which names are worth knowing at all, and no amount of word matching answers that, because importance is a property of the graph rather than of the text. --rank runs PageRank over the graph the map already holds: files are nodes, an edge runs from a file that uses a name to the file that declares it, and the weight is the square root of how often. Aider's repo map is where the idea and the four multipliers come from.

$ where-are-we --rank --limit 4

0.061419466 find_text /repo/src/where_are_we/_mapper/declare.py:135
0.060390690 build /repo/src/where_are_we/_mapper/build.py:283
0.047834365 build /repo/src/where_are_we/ask.py:835
0.033650027 add /repo/src/where_are_we/eval.py:136

That is this repository's own map, unedited: the declaration scanner, the build, ask's inner budget loop and the counter wawe-eval runs over its questions. The scores are of a map that indexes this README, so editing this file moves their last digits; the order is what the CI step pins, across both supported Pythons.

Name the files you are working in and the walk is personalised on them, which turns "what matters here" into the question actually worth asking: what should I read given that I am editing this. On ask.py, find_text goes from 0.0614 to 0.0800, a third, and keeps first place, which is right: ask.py is what calls it. Personalisation shows up here as a score rather than as a swap, because the unpersonalised ranking already has it first.

$ where-are-we --rank src/where_are_we/ask.py --limit 2

0.080008437 find_text /repo/src/where_are_we/_mapper/declare.py:135
0.059723023 build /repo/src/where_are_we/_mapper/build.py:283

The same files bias --ask without changing what it found: --ask charge --files billing/ prints the billing/ rows first inside every section and the rest after them, with the section's usual tail. The handle on that tail carries the scope as one more field, so more reorders the section the same way before it continues: a scoped answer reaches every row an unscoped one does, which is the guarantee the tail is there to make. --files - reads the list on stdin, so git diff --name-only | where-are-we --ask charge --files - asks the map about the branch you are on.

Paths are compared relative to the repository root, and a prefix stops at a directory separator, so --files bill is not billing.py. A path nothing indexed matches is named on stderr rather than answered as if it were the whole repository.

Landing on a name usually raises all five questions at once, and --context answers them together:

$ where-are-we --context charge

Context for `charge`: declared, map rows, callers, callees, impact to depth 1. 12000 characters, floor shares 15/35/15/15/20 percent.

## Declared in
- charge: billing/core.py:10-24 (function)
- charge: legacy/pay.py:88-91 (function)

## What the map says
## Defined here
- `charge` — billing/core.py:10
...

## Callers
- checkout.py:pay

## Callees
- settle (billing/ledger.py)

## Impact
Impact of `charge` to depth 1. How to read it: ...
depth 1: checkout.py:pay

The shares are floors, and the budget is allocated in two passes. The first asks every block what printing all of itself would cost. The second gives each block the smaller of that and its share - 15 percent of the budget for the declarations, 35 for the map's own rows, 15 for the callers, 15 for the callees, 20 for the impact - and then hands what nobody claimed to the blocks still short, in that same order. So when the five answers together fit the budget, every one of them is printed whole; and the same name at the same budget always composes the same answer, since the needs come from the map and the order is fixed.

A block that could not print all of itself ends in … N more lines (more:ctx:...), and --more on that handle returns the rest of that block. A block that could not even be given the room to name itself and carry that handle is left out rather than printed as a count nobody can follow, and the blocks are served in order, so below about a thousand characters the first of them are printed and the rest are not there. That is why the MCP server never asks for less than 1500: under it some rows are out of reach.

The other map: the specifications

A codebase is not the only thing an agent gropes around in. The other is the tracker — the ticket, its parents, what it links to, what mentions it — and it gropes there the same way and for the same reason: no map, so it asks, and asks again.

Measured on one run of a real pipeline: sixteen tickets fetched over and over. One agent pulled fourteen neighbours to understand the task; the next agent pulled the same fourteen again, because a session cannot see another session's memory. Three were fetched three times inside a single session, since finding an answer already in a conversation costs more than asking for it fresh. Every answer then sat in the context for ever, and every later turn paid to re-read it.

$ where-are-we --specs APF-1934 --spec-cmd 'python3 fetch.py {key}'

  APF-1934 (1 so far)
  APF-1860 (2 so far)
  APF-2752 (3 so far)
spec map: 3 ticket(s) -> ./spec_map.md

This tool knows nothing about any tracker, which is the same contract as the rest of it: you hand it a command that turns a ticket key into JSON, it walks the links two hops out, and it writes spec_map.json and spec_map.md. Jira, Linear, GitHub Issues, a text file — it never finds out.

Two trackers it does not need a command for: --spec-source github builds the gh issue view {key} --repo owner/name --json ... call itself, reading owner/name off the repository's origin remote; --spec-source linear builds the GraphQL call over curl and needs LINEAR_API_KEY set. Either way --spec-cmd is filled in rather than typed.

--ask answers from both maps, because a question about a piece of work is as likely to be about what was asked for as about where the code is.

What a map leaves out, it says

Both walks are bounded, because a repository and a tracker are both graphs and a graph will hand over everything if asked. What the bound cut is named in the map itself, at the top:

## This map is incomplete

- the file walk stopped at 40000 files under /work — raise WAWE_MAX_FILES or add
  to .wawe-ignore; what is below that count is mapped and the rest is not

A limit that stops quietly produces a map that looks complete and is not, and the reader has no way to tell — which is worse than a small map, because a small map that says so can be asked to grow. An absence in a silent map reads as a fact about the codebase.

flag bounds
--spec-depth hops from the starting ticket (2)
--spec-limit tickets fetched at most (60)
WAWE_MAX_FILES files read from the repository (40000)
.wawe-ignore paths never read at all

What it has that the others do not

Every other tool in this space does one of three things. Some hand the model a pile of text and hope: repomix, gitingest, code2prompt. Some hand it a search index whose answers move when the embedding model does: claude-context, chunkhound, continue's @codebase, greptile. Some hand it a live language server that has to be installed, started and kept warm per language: serena, mcp-language-server, opencode. This one treats the budget as a contract and publishes what the contract costs. An answer is whole rows or no row, never a row cut in the middle, and only a continuation asked for a row longer than its whole budget returns one cut, marked with how much was taken off. It is at or under its stated ceiling. Every section ends by counting what it left out and carrying a handle that fetches it. And wawe-eval measures, on every CI run, what the cut loses: first-answer recall, pooled recall, top-5 recall, rows longer than the budget, and recall after every handle is followed, with that last number asserted at 1.0 from 1500 characters upward.

Nobody else states that. Repomix counts tokens and never says what --compress dropped. Aider binary-searches its map into --map-tokens and says nothing about what fell out. Every vector-backed tool returns top-k and is silent about rank eleven. Graphify is the only other project that publishes reproducible retrieval numbers and tags where an edge came from, and it still needs a model for anything that is not code.

The same honesty runs through the rest. An absence names what was searched (indexed: suite 260 files) instead of asserting the thing does not exist. ## This map is incomplete names every bound that cut. A cross-file edge the map is guessing at is written charge (a.ts|c.ts)? with every candidate listed, rather than passing a choice off as a lookup. call_graph_stats publishes how much of the call tree resolved, 0.2615 on this repository at 1.5.0, instead of quietly reporting only the edges it managed to draw.

And the test suite is a first class thing here, which it is nowhere else: step phrases with their overlaps and duplicates, feature to step traceability, scenario history and slow steps read out of JUnit XML, quarantine tags, fragile locators, test-id ownership, and the product under test indexed beside the suite that drives it. It builds in one tree walk, offline, for zero tokens, byte identically, out of the standard library, and what reaches the prompt is an 849-byte pointer rather than the map.

Why install it

  • The first forty turns stop repeating. The answers never change between sessions and need no model to produce, so produce them once and commit them.
  • A step that exists stops being written twice. Overlapping phrases, dead phrases and uncalled page-object methods are listed by name.
  • It reads a repo it has never seen. Detection is by shape, not directory name: a page object is a class that owns selectors, wherever it lives.
  • It costs a tree walk. 3s on a 6k-file repo, 2.5min on a 36k-file one, cold. Deterministic — same tree, same map, no API bill.

Install

pip install where-are-we
pip install "where-are-we[semantic]"   # + local embeddings for a semantic --ask
brew tap ngavrish/tap && brew install where-are-we
curl -fsSL https://ngavrish.github.io/where-are-we/install.sh | sh

macOS, Debian 13, Ubuntu 24.04, Fedora, RHEL (any of them with Python 3.12 or newer). Or ghcr.io/ngavrish/where-are-we. The [semantic] extra adds fastembed (ONNX on CPU, no service, no database); without it --ask still answers by keyword.

Needs Python 3.12 or newer (stdlib only, no dependencies to install alongside it); 3.10 and 3.11 are no longer supported.

Output

File Contents
framework_map_brief.md the digest for a prompt
framework_map.md every step phrase, every scenario with its line number
framework_map.json the same as data, under a versioned contract

--agent-file writes a pointer into AGENTS.md, CLAUDE.md or .cursorrules between markers. The rest of the file survives.

Why a pointer and not the map

A prompt is re-sent in full on every turn — that is what a conversation is — so anything put in one is paid for on every turn of the session, read or not.

Measured on a real run: the brief inlined whole was 253 KB, the agent carrying it took 424 turns, and the map alone came to 27.4 million tokens re-sent — a quarter of everything that run consumed, and the reason a five-hour allowance emptied in seventy-four minutes. Trimming it to an index still cost 6k a turn for a document most turns never opened.

in the prompt per turn
the brief, inlined 253 KB ≈ 64k tokens
an index of its sections 27 KB ≈ 6k tokens
a pointer 849 B ≈ 212 tokens

The sections are still named in the pointer, because an agent that cannot see that a section exists goes back to grepping the repository — which is the thing this was built to end. Naming them costs two hundred tokens; carrying them costs sixty-four thousand, every turn.

command what it prints
--pointer what belongs in a prompt: the path, the sections, how to ask
--ask "words" only the rows that mention those words, their stems and their synonyms, whole, ranked by section; says what it left out and what a synonym matched
--ask "words" (with [semantic]) the keyword hits plus a "Related by meaning" tail from a local embedding index
--corpus NAME=PATH fold an external corpus (a rules dir, a runbook) into the same semantic answers
--no-semantic skip the embedding index even when fastembed is installed
--defines NAME every place that name is declared, each as file:start-end (kind), in path order
--at FILE:LINE the whole definition that encloses that line, with a handle for the rest of it
--context NAME the five answers about a name in one: declared, map rows, callers, callees, impact one hop out
--affected FILE[,FILE] which tests a change to those files reaches: the scenarios, their feature files, the routes and page objects, and the files the graph has no row for
--changed [REF] the same, for the files git diff --name-only REF names (HEAD by default)
--affected-format behave|pytest that selection as the runner's own list: one --name per affected scenario, or pytest node ids. Printed whole, with no ceiling and no handle
--affected-out FILE write that selection, and nothing else, to FILE, and print the answer a person reads to stdout. Empty file means run nothing
--affected-depth N how many call hops --affected follows upward (1 to 12, 6 by default)
--reaches NAME which scenarios, pytest cases and routes reach that function or class, grouped by feature file, with the call chain for the first scenario of each
--unreached the product definitions no test reaches, grouped by file and ranked, under a first line saying how much of the call graph resolved
--path A,B the shortest call chain from A to B over the map's xrefs calls rows, one hop per line with the rule that placed each edge and the line the call is on; each end is a name or FILE:NAME
--path-depth N how many call hops --path follows forward (1 to 12, 6 by default)
--range NAME every home of NAME as file:start-end kind, the text of the shortest one, and the three lines an editor anchors on: the first, the last, and the one after the last. A line the map redacted is printed with a warning beside it: it is not the line on disk, so an edit anchored there will not match
--dead questions, not dead code: the definitions no xrefs calls row lands on, grouped by file. A call through an imported module or inside the declaring file leaves no row, so on a library most rows are unplaceable; the first line says so and ## How this was counted names the exclusions
--hot the map's own rank score times the commits its most-changed-files section counted, both numbers shown; that section is the forty busiest files, so anything outside it counts 1
--rank [FILE,...] the definitions this repository is built around, best first; given files, what to read while editing them
--limit N how many rows --rank prints and how many definitions --unreached ranks (default 200 for both; below 1 is refused)
--limit N how many rows --rank, --dead and --hot print (default 200, 40 files, 40; below 1 is refused)
--ask "words" --files a.py,b.py the same answer with the rows about those files first in every section; --files - reads the list on stdin
--mcp serve the map over MCP on stdin/stdout instead of answering once
--sections the section headings
--cost [N] what each section of framework_map.md and of the brief beside it costs: rows, bytes and estimated tokens, heaviest first, hiding what is under N bytes
--export FILE the map and its brief as one file to paste: the notice, the counts, the priced section list, then the brief. - for stdout

A pipeline asks the map what to re-run and then respects an empty answer:

where-are-we --out .wawe --changed "$base" \
  --affected-format behave --affected-out sel.txt
[ -s sel.txt ] && xargs behave < sel.txt || echo "nothing to run"

The [ -s ] is the whole point of the file being empty rather than absent or prose: xargs handed empty input runs behave with no arguments, which is the entire suite, so a commit that reaches no scenario would run everything rather than nothing.

WAWE_EMBED_CACHE=<file> caches the semantic index's embeddings in one sqlite file keyed by model and text, so a rebuild does not recompute vectors it already has. Unset keeps the old behaviour.

What it reads

  • Code (UTF-8, or UTF-16/UTF-32 with a byte order mark) — languages, entry points, make targets, npm scripts, container commands, HTTP routes and status codes, data model, module public surface, call and package graphs, cycles, unimported files, complexity hotspots, duplicate blocks.
  • Runtime — queues, topics, gRPC, cron, Kubernetes probes and resources, Terraform, Pulumi, Ansible, cache keys, permissions, metrics, spans, log fields, error types, retries, timeouts, breakers, rate limits, transactions, idempotency, outbound services, installed versions from lock files.
  • Contracts — OpenAPI, GraphQL, migrations, mocks, feature flags and their branch points, locale keys, pinned images, secret paths (never values).
  • Decay — deprecations, coverage, docs pointing at deleted files, git history and who touches what.
  • Tests — layers, entry points, callable step signatures, hooks, locators, timeouts, fixtures, tag meanings, overlapping and unused step phrases, dead page-object methods, slow scenarios from past junit.
Supported stacks

Test runners — behave, pytest, jest, vitest, playwright, cypress, robot, JUnit, TestNG, Cucumber (JVM/JS/Ruby), rspec, go test, xUnit, NUnit, SpecFlow, PHPUnit, Behat, Rust, XCTest, ExUnit, Flutter, Spock, clojure.test, hspec, busted, Foundry, karate, gauge, k6, gatling, JMeter, Locust, Espresso, Detox.

Languages — Python, TypeScript, JavaScript, Go, Java, Kotlin, Scala, Ruby, Rust, C#, PHP, Swift, C, C++, Elixir, Erlang, Dart, Groovy, Clojure, Haskell, Lua, Perl, R, Julia, Objective-C, F#, Solidity, Shell, SQL.

Web — Flask, FastAPI, Django, Express, Nest, Go net/http, chi, Spring, Rails, React, Vue, Svelte, Angular, Storybook.

Infrastructure — Docker, Compose, Kubernetes, Helm, Terraform, CloudFormation, Pulumi, Bicep, Ansible, Chef, Puppet, GitHub Actions, GitLab CI, Jenkins, CircleCI, Azure Pipelines, Buildkite, Drone.

Data — PostgreSQL and friends, MongoDB, Elasticsearch, DynamoDB, Cassandra, ClickHouse, Kafka, RabbitMQ, SQS, NATS, Pulsar, MQTT, dbt, Airflow, Spark, notebooks.

In a pipeline

- uses: ngavrish/where-are-we@v1
  with:
    agent-file: AGENTS.md
    comment: "true"
- repo: https://github.com/ngavrish/where-are-we
  rev: v1.0.0
  hooks: [{id: where-are-we}]

What to re-run after a change

The map already knows which scenarios reach the code a commit touched. This is the pipeline that runs those and nothing else, and falls back to the whole suite whenever the answer is partial.

where-are-we --out .wawe --changed "$BASE" \
  --affected-format behave --affected-out sel.txt > affected.txt
cat affected.txt                      # the head says what was selected and why
if grep -qE 'No xrefs row names|No calls row lands on' affected.txt; then
  behave                              # the answer says it is partial
elif [ -s sel.txt ]; then
  xargs behave < sel.txt
else
  echo "nothing to run"
fi

--changed "$BASE" reads git diff --name-only "$BASE", so the pipeline passes no file names of its own; with no argument it is HEAD, which is the working tree against the last commit.

Five things that shape make true, each of which a shorter version gets wrong:

  • The file is the selection and nothing else. One argument pair per line, no head, no prose, no tail, whatever the size: xargs behave < sel.txt is all the parsing a pipeline does. The head and the blocks go to stdout, which is where a person reads them.
  • An empty file means run nothing, and [ -s ] is what says so. A commit that reaches no scenario writes zero bytes, and a --changed run over a tree nothing has touched writes zero bytes too and says so on stdout. Both are exit 0. GNU xargs with empty input runs the command with no arguments, and behave with no arguments is the whole suite, so a pipeline that pipes straight into xargs runs everything at the moment it should run nothing. BSD xargs, which macOS ships, runs nothing there, so the fault appears on the CI runner and not on the laptop the pipeline was written on.
  • The selection is --name and only --name. One --name per affected scenario, anchored on the whole name (--name '^Pay\ with\ a\ new\ card( -- @|$)'), because --name is action="append" in behave and several of them are a union. Past a couple of hundred scenarios one --name per feature file alternates that file's affected scenarios instead, which is the same selection in fewer arguments. Nothing emits -i or --tags.
  • Two scenarios of one name are one selector, and behave runs both. The block head says so on every answer that can be affected by it. That over-selects, which is the safe direction, and it is the one place the selection is not exactly the list printed above it.
  • A file the graph has no row for is not "not affected". It lands in ## Unreachable from the graph, the first line says No xrefs row names N of the files given, and the selection file says nothing about it, correctly, because there is nothing to say.
  • Nor is a file the map declares and no call row lands on. That one is in the map, so the answer looks complete, and it is either a thing nothing calls or a thing reached by a route no rule places (importlib.import_module plus getattr is the shape). The first line says No calls row lands on N of the files given, so this map does not know what calls them: either nothing does, or a caller the resolver could not place, and follows it with the graph's own resolution rate, which is the number that says which of the two is likelier. Both sentences are what the grep in the pipeline above reads: a partial answer is a full run.

--affected-format pytest is the same shape with node ids from pytest_tests, one per line for xargs pytest < sel.txt.

What this cannot see. affected walks xrefs rows and nothing else. A call the resolver could not place is not a row, so the scenario that reaches the change through it is not in the selection: on a test suite, where a page object is called from step modules, that is few, and on a library it is most of the graph. unreached carries the map's own resolution rate in every first line, affected carries it in the one branch where a reader cannot otherwise tell a real zero from an unplaced caller, and wawe-eval --map OUT --graph prints it per language whenever you want it. Run the whole suite on a schedule, and on every release, whatever the selection says.

What writes and what only reads

A pre-execution guard sees an argv and has to decide. This tool ships the answer instead of leaving it to guess: where-are-we --effects prints what every flag does to the disk, and a command's class is the highest class any of its flags carries.

class what it touches flags
read answers from what is already there. It may append one line to <out>/.wawe-ask.log, the map directory's own record of what was asked --ask, --more, --defines, --at, --context, --affected, --changed, --affected-format, --affected-depth, --reaches, --unreached, --path, --path-depth, --range, --callers, --callees, --impact, --impact-depth, --sections, --cost, --pointer, --mcp, --lsp, --repo, --product, --also, --rules, --for, --only, --skip, --max-lines, --corpus, --no-semantic, --quiet, --effects, --json, --dry-run, --help
read answers from what is already there. It may append one line to <out>/.wawe-ask.log, the map directory's own record of what was asked --ask, --more, --defines, --at, --context, --affected, --changed, --affected-format, --affected-depth, --path, --path-depth, --range, --dead, --hot, --callers, --callees, --impact, --impact-depth, --sections, --cost, --pointer, --mcp, --lsp, --repo, --product, --also, --rules, --for, --only, --skip, --max-lines, --corpus, --no-semantic, --quiet, --effects, --json, --dry-run, --help
writes-map-dir the map files and the parse cache under --out --out, --html, --ctags, --force, --watch, --diff
writes-repo the repository being mapped: a manifest, an agent file, the READMEs a directory has none of, and the files --export and --affected-out were told to write, which are at whatever path the caller named --init, --agent-file, --docs, --export, --affected-out
writes-config where a tool other than this one reads: .git/hooks, ~/.claude/settings.json, ~/.codex/config.toml, a Cursor rule, a Gemini setting --install-hook
network off this machine: a tracker fetch, a runs API --specs, --spec-cmd, --spec-source, --spec-depth, --spec-limit, --runs-api

The order is read < writes-map-dir < writes-repo < writes-config < network. --effects --json prints the same table as {"schema": "where-are-we-effects/1", "flags": {...}, "order": [...], "notes": {...}}, where notes says per class what it touches, including the answer log a read can append to. That JSON is installed beside the code as effects.json, so a guard written in something other than Python reads the file instead of the table, and reads the same caveats a human does.

--effects -- <command line> classifies one command line with this tool's own parser, running nothing:

$ where-are-we --effects -- where-are-we --out /tmp/m --ask x
writes-map-dir
--out writes-map-dir
--ask read

A flag is resolved the way argparse resolves it, so --eff is --effects and the class of an abbreviated line is the class of the line that runs.

A command line naming none of --ask, --more, --defines, --at, --context, --affected, --changed, --reaches, --unreached, --path, --rank, --callers, --callees, --context, --affected, --changed, --path, --range, --rank, --callers, --callees, --context, --affected, --changed, --path, --range, --dead, --hot, --rank, --callers, --callees, --impact, --sections, --cost, --export, --pointer, --mcp, --lsp, --init, --install-hook, --specs, --dry-run, --effects or --help builds the map into --out, so where-are-we --repo . is writes-map-dir on the strength of the build alone and says so as a build writes-map-dir line.

--dry-run prints every path the command can write, one per line, and exits without writing any of them:

$ where-are-we --repo . --install-hook git --dry-run
would write /repo/.git/hooks/post-checkout
would write /repo/.git/hooks/post-merge
would write /repo/.git/hooks/post-commit

would write when nothing is there, would replace when a file is. The paths come from the same expressions the writers use, so the preview names what the real run names; whether a listed file is then written depends on what is already in it, since a target that already says what this tool would say is left alone.

It covers every command line, not only the ones that write:

the line the preview
--init the manifest
--install-hook KIND the files that kind installs, and for claude or codex with HOME unset, the same refusal the real install gives, exit 2
--agent-file, or a plain build the map files under --out, the parse cache, framework_map.html with --html, tags with --ctags, and the agent file
--docs write the documents it would create, or nothing to write: every directory already explains itself
--specs spec_map.json and spec_map.md. The tracker command is not run
--export FILE FILE, at whatever path was given. The file is not written
--ask, --cost, --mcp and the other reads nothing to write: --ask only read, and no answer, since an answer is not a preview

The optional semantic index adds semantic_index.json and semantic_index.npy to the same directory as the map.

Environment

Every variable the tool reads. A flag always wins over the variable it defaults from.

name read in what it does default
AGENT_REPO _mapper/walk.py, cli.py, readmes.py the repository to index or answer about, when --repo is not given. main() also writes it back so the walk and the product guess see the resolved path unset: --out's parent when that is a .wawe, then /work if it exists, then the current directory
RUN_DIR cli.py, _mapper/build.py where the map files are written, when --out is not given .
PRODUCT_SRC _mapper/walk.py, cli.py the product under test, colon or comma separated, when --product is not given. none switches the sibling guess off unset: the siblings of a repository that looks like a test suite
RULES_REPO _mapper/build.py, cli.py a directory of agent rule files to fold into the map, when --rules is not given /rules
RUNS_API_READ _mapper/build.py, cli.py base URL of a runs API whose recent verdicts go into the map, when --runs-api is not given unset: no runs section
SPEC_ROOTS cli.py the ticket keys --specs walks from, comma separated unset
SPEC_FETCH_CMD cli.py the command that fetches one ticket as JSON, when --spec-cmd is not given unset: --specs refuses to run without one
SPEC_SOURCE cli.py which built-in tracker command to use (jira, linear, github, cmd), when --spec-source is not given cmd
WAWE_SPEC_DEPTH specs.py how many link hops out from each root ticket the spec map walks 2
WAWE_SPEC_LIMIT specs.py the most tickets one spec map will fetch 60
WAWE_MAX_FILES _mapper/walk.py the most files one walk will visit before it stops and says so in the map 40000
WAWE_NO_CACHE _mapper/build.py, _mapper/walk.py set to anything: parse every file again and leave the parse cache exactly as it was. --force re-parses but rewrites the cache unset: the cache is read and written
WAWE_DEBUG_PARSES _mapper/build.py set to anything: print the parse count and the hash count per build to stderr, to see what an incremental rebuild actually re-read and re-hashed unset: silent
WAWE_JUNIT_DIRS _mapper/build.py extra directories of JUnit XML to read past runs from, separated by the platform's path separator unset: the repository's own reports directories
WAWE_POINTER_MAX _mapper/state.py the byte cap on the pointer, the block a SessionStart hook puts into context 4000
WAWE_VOCAB _mapper/render.py cap on how many vocabulary entries the brief prints, split across the groups 0, meaning no cap
WAWE_ASK_LOG ask.py set to 0 to stop appending a row per answer to <out>/.wawe-ask.log unset: the log is written
WAWE_EMBED_MODEL semantic.py the embedding model the optional semantic index uses BAAI/bge-small-en-v1.5
WAWE_RERANK_MODEL semantic.py the cross encoder that reranks semantic hits Xenova/ms-marco-MiniLM-L-6-v2
WAWE_EMBED_CACHE semantic.py a directory to keep embeddings in between runs unset: no cache
WAWE_STRICT the Claude Code plugin, not src/ set to 1 and the plugin's PreToolUse hook refuses Grep, Glob and Bash searches over the repository, so the map is asked instead unset: searches are allowed
PYTHONIOENCODING the interpreter a codec narrower than the map's text no longer fails: characters it cannot carry are replaced unset: the locale's codec
ANTHROPIC_API_KEY eval.py, read at import, used only by wawe-eval --agent the Claude API key the agent A/B calls with; without it the command refuses and sends nothing unset

Each variable is read in one place, and named there. WAWE_NO_CACHE, WAWE_DEBUG_PARSES and WAWE_POINTER_MAX are read once when _mapper/state.py is imported (NO_CACHE, DEBUG_PARSES, POINTER_MAX), WAWE_MAX_FILES when _mapper/walk.py is (MAX_FILES), WAWE_VOCAB when _mapper/render.py is (VOCAB_CAP), WAWE_ASK_LOG when ask.py is (LOG_ANSWERS), the three WAWE_EMBED/WAWE_RERANK ones when semantic.py is, and the two WAWE_SPEC ones when specs.py is. So a process that sets one of those after importing the package keeps the value it started with.

The rest are read per call. Five of them are the ones a flag writes back into the environment for a later stage to pick up (AGENT_REPO, PRODUCT_SRC, RUN_DIR, RULES_REPO, RUNS_API_READ). Three more are argparse defaults, which main() evaluates when it builds the parser (SPEC_ROOTS, SPEC_FETCH_CMD, SPEC_SOURCE). The last is WAWE_JUNIT_DIRS, which a caller that builds several maps in one process sets per build.

Keeping it honest

where-are-we --init                 # starter .framework-map.json
where-are-we --docs                 # list the docs the repo lacks (--docs write to create)
where-are-we --install-hook git     # post-checkout, post-merge, post-commit
where-are-we --install-hook claude  # before the first turn of a session (agent = claude)
where-are-we --install-hook cursor  # a Cursor rule plus its MCP config
where-are-we --install-hook codex   # an AGENTS.md block plus ~/.codex/config.toml
where-are-we --install-hook gemini  # a GEMINI.md block plus .gemini/settings.json
where-are-we --diff                 # what changed since the last map

Every --install-hook kind installs as one unit: each target (the three git hooks, a rule file and its MCP config, a markdown block and its settings file) is checked before any of them is written. A target that is a symlink, not valid JSON, or not writable stops the whole install with the cause named and nothing changed; fix it and rerun, and the rest lands. Rerunning on a finished install says "already installed" and touches nothing.

Autodetection gets the shape right and the vocabulary wrong, so a repo states its own in .framework-map.json and what it states wins:

{
  "name": "billing-e2e",
  "purpose": "End-to-end tests for the billing portal.",
  "layers": {"steps": "steps/*.py — steps own no selectors, they call page objects"},
  "product_src": ["../billing-web/src"],
  "conventions": ["After a fix, re-run only what failed."]
}

.wawe.toml holds CLI flags as defaults, .wawe-ignore keeps build output out, existing files are never overwritten, and anything shaped like a credential is redacted before it reaches a file. The commit and the newest file in the tree are recorded with the map, and so is a sha256 over what every indexed file holds, so a re-run on an unchanged tree costs a stat walk and a re-run on a tree whose timestamps lie still sees it.

What is redacted

The map holds every indexed line of every indexed file, and the map gets committed and pasted into prompts, so these rules replace a credential with [redacted] before anything is written:

  1. A whole PEM block, -----BEGIN ... PRIVATE KEY----- through -----END ...-----, header and body alike.
  2. An issuer prefix at the start of a word: AKIA... (AWS), ghp_/gho_/ ghs_/ghu_/github_pat_ (GitHub), xox?-... (Slack), sk_live_/ sk_test_/rk_live_/rk_test_ (Stripe), sk-/sk-proj- (OpenAI), pypi- (PyPI), a JWT.
  3. A base64 blob of forty characters or more that carries a + or ends in = padding.
  4. The password inside a URL: postgres://admin:pw@host/db keeps the scheme, the user and the host and loses the password.
  5. The value on a line whose left-hand side names a secret. The last segment of the key has to be secret, password, passwd, token, api_key, private_key, credential, auth or authorization, in an assignment, a dict or JSON key, a YAML key or an export. A quoted value and a bare value after = are replaced wherever they sit on the line, so a .env line inside a shell string counts too. A bare value after : is replaced only on a line shaped like YAML: the key starts the line, is not quoted, and nothing after the value turns the line back into code.

Key names are kept, and so are the quotes around a redacted literal, so a question about where a password is set still gets the file, the line and the syntax. What is not redacted, deliberately:

  • Code on the right-hand side. No value rule admits a bracket, so token = lexer.next_token() and PASSWORD = os.environ["PW"] stay.
  • A name that merely mentions a credential. The secret word has to be the last segment, so token_count, max_token_count, auth_backend, secret_name, api_key_header, private_key_path and credential_kind all stay.
  • A number, a True/False/None, or a bare type name, whatever the key is called: has_token = True and the dataclass field token: str = "" stay.
  • A commit sha, and a path. Rule 3 needs a + or an =, and a forty-character hex sha has neither. A slash is not a gate either, because src/main/java/com/example/service/impl/CustomerServiceImpl is a run of letters and slashes and nothing else.

.wawe.toml's [synonyms] table adds a project's own words to --ask's built-in groups (login/signin/auth, invoice/bill/billing, and eighteen more):

[synonyms]
invoice = ["proforma", "receipt"]

merges into the group that already has invoice, or starts a new group when none does. --ask "invoice" then also searches proforma and receipt.

All options
--repo PATH                  the repository to index
--product PATH,…             source roots of the application under test
--also PATH,…                other repositories to fold into the same map
--out DIR                    where the three files land
--agent-file FILE            also write the brief into AGENTS.md, CLAUDE.md, …
--docs [write]               offer the repository the documentation it lacks
--for author|coder           author gets the whole vocabulary; coder gets the rest
--only "routes,data model"   keep only these sections in the brief
--skip "coverage,history"    drop these
--max-lines N                cap the brief per section; the full map is untouched
--diff                       what changed since the map already in --out:
                             the files that moved, then the keys
--init                       write a starter .framework-map.json
--install-hook KIND          wire it into something that already runs:
                             git|claude|cursor|codex|gemini (agent = claude)
--watch SECONDS              rebuild whenever the tree moves
--html                       also write framework_map.html
--ctags                      also write <out>/tags, in universal-ctags format
--force                      rebuild even when nothing moved, reading
                             nothing from the parse cache
--quiet                      no summary line
--dry-run                    print every path this command would write, and
                             write none of them
--defines NAME               every place NAME is declared, with its span
--at FILE:LINE               the whole definition that encloses that line
--context NAME               declared, map rows, callers, callees and
                             impact one hop out, in one answer
--affected FILE[,FILE]       which tests a change to those files reaches:
                             scenarios, features, routes, page objects
--changed [REF]              the same, for what `git diff --name-only REF`
                             names (HEAD by default)
--affected-format behave|pytest
                             that selection as the runner's own list,
                             printed whole
--affected-out FILE          write that selection to FILE and print the
                             answer a person reads to stdout; an empty
                             file means run nothing
--affected-depth N           how many hops --affected follows (default 6)
--reaches NAME               which scenarios, pytest cases and routes reach
                             that function or class, with the chain of calls
--unreached                  the product definitions no test reaches, ranked,
                             with the resolution rate
--path A,B                   the shortest call chain from A to B, one hop
                             per line with how each edge was resolved
--path-depth N               how many hops --path follows (default 6)
--range NAME                 every home of NAME, the shortest one's text,
                             and the lines an editor anchors on; a line this
                             map redacted is printed with a warning beside
                             it: it is not the line on disk
--dead                       questions, not dead code: the definitions no
                             call row lands on, by file
--hot                        rank score times commits, both numbers shown;
                             churn covers the forty busiest files
--rank [FILE,...]            the definitions the repository is built around,
                             personalised on the files you name
--files FILE[,FILE]          on --ask: those files' rows first in every
                             section; `-` reads the list on stdin
--limit N                    how many rows --rank, --dead and --hot print
                             (default 200, 40 files, 40)
--cost [THRESHOLD]           what the map and its brief cost per section:
                             rows, bytes and tokens, heaviest first
--export FILE                the map as one self-contained file to paste,
                             or `-` for stdout
--effects [--json]           what every flag does to the disk; with
                             `-- <command line>`, the class of that line

As an MCP server

where-are-we --mcp --out /path/to/the/map

Eighteen tools over JSON-RPC on stdin and stdout: affected, ask, at, callees, callers, context, dead, defines, find, hot, impact, more, path, range, rank, reaches, sections, unreached. defines answers where a name is declared, in every file that declares it, with the line each declaration ends on; at takes the file:line a stack trace names and returns the whole definition around it; context answers all five questions about a name at once, so landing on one costs a single round trip; rank answers what the repository is built around, and takes the files you are editing so the answer is about them; find answers where a phrase appears, which is the other half of what a grep was for. ask takes files for the same reason --files exists. The same index answering the same questions; what changes is that the question is an argument and the answer is a tool result, rather than a shell command and its output sitting in the conversation to be re-read on every turn after.

It reads the JSON the mapper wrote, on the same machine, offline.

As a library

from where_are_we import build, brief

m = build("/path/to/repo")
open("AGENTS.md", "w").write(brief(m))

Examples

Real output on a behave suite, a Go service and a React app — docs/examples. Generated by running the tool, not by hand.

Why it exists

Built inside an agentic QA pipeline where seven branches ran at once, each opening with the same forty greps. Three runs died at their deadline with the branches still reading. None of it was specific to that pipeline, agent, or language.

Contributing

Issues and PRs welcome — CONTRIBUTING.md. A change keeps the contract in SCHEMA.md and comes with a case in tests/ built from a real directory.

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

where_are_we-1.6.0.tar.gz (655.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

where_are_we-1.6.0-py3-none-any.whl (350.9 kB view details)

Uploaded Python 3

File details

Details for the file where_are_we-1.6.0.tar.gz.

File metadata

  • Download URL: where_are_we-1.6.0.tar.gz
  • Upload date:
  • Size: 655.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for where_are_we-1.6.0.tar.gz
Algorithm Hash digest
SHA256 4816f5ea0cb9f9bd1d93edbc7dcef22797e84d7d390d87cdf32a0a89293e2369
MD5 0e94249670d9acc484b520b9100124e3
BLAKE2b-256 a9c41e7f6fdd988faee4e40d2e629da586feaded0f588750da49f23ceda87a9f

See more details on using hashes here.

File details

Details for the file where_are_we-1.6.0-py3-none-any.whl.

File metadata

  • Download URL: where_are_we-1.6.0-py3-none-any.whl
  • Upload date:
  • Size: 350.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for where_are_we-1.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 cb0c8eac413041e51a595b56088c14fb7207f43175e05cfa91e3f07935345081
MD5 f12066fc5672153b27b5ecfe626ef6ab
BLAKE2b-256 9a3110062770dcd5bf4a913355d50055f8f836f713bf9c1d697d25a68f69aa5c

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

1.6.0 This release

2 files

1.5.0

2 files

1.4.1

2 files

1.4.0

2 files

1.3.0

2 files

1.2.0

2 files

1.1.3

2 files

1.1.2

2 files

1.1.1

2 files

1.1.0

2 files

1.0.0

2 files

0.12.3

2 files

0.12.2

2 files

0.12.1

2 files

0.12.0

2 files

0.11.2

2 files

0.11.1

2 files

0.11.0

2 files

0.8.1

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.1

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page