Skip to main content

Seamcheck

Finds the bugs that sit between two things — a request and the route meant to serve it, a cache key written in one service and read in another, a job queued that no worker consumes, an element four files are fighting over. Each side is valid on its own, which is why nothing else catches them.

Sponsor PyPI Python License

pip install seamcheck && seamcheck map

It reads your source. It never runs your code, and it makes no network call — no API key, no account, nothing to sign up for.

The report is served from your machine: one link for this computer, one to type on a phone on the same wifi. A phone usually is not on that wifi, so seamcheck config --tunnel always remembers, for this machine, that every run should also print a public HTTPS link. That is the one thing here that leaves the machine — anyone holding the link can read the report while the command runs, and it dies when you stop it. It is off until you ask for it, and --local-only overrules it for a single run.

For agents

This tool is built for you as much as for a person. findings, symbols and diff answer a bounded question at a few KB, instead of the whole graph json prints - 72 MB and about 18 million tokens on a 500k-line project:

seamcheck findings --file app/views.py     # what is wrong here
seamcheck symbols --search submit_push     # name -> id
seamcheck diff --since origin/main         # what this branch changed
seamcheck check --since origin/main        # the gate: 0 clean, 1 findings, 2 no baseline

findings, symbols and diff take --limit/--cursor to page, print one JSON envelope on stdout and nothing else (prose goes to stderr). A failed query (a bad --status, an unresolvable --since) comes back as a code in the envelope's error field (unknown_symbol, no_baseline, bad_argument, ...) rather than a string you have to parse, and the process exit code matches it - see Exit codes. A directory with nothing this tool recognises never reaches the envelope at all: it exits 4 with a plain message on stderr, before any of the three starts. The three of them share one scan cache keyed to the file tree, so a second question against an unchanged tree is nearly free; the whole graph is still there, with --full --yes, when you actually need it.

There is an MCP server with the same functions behind it: seamcheck_unverified (call this first - the queue of findings nobody has judged yet, worst first), seamcheck_check, seamcheck_explain, seamcheck_triage, seamcheck_why_wrong, seamcheck_report, seamcheck_share, seamcheck_services, seamcheck_findings, seamcheck_symbols, seamcheck_diff, seamcheck_snapshot.

The pipeline recipe · Using it from an agent

Why I made it

I was building a game — a fairly large Django app with a lot of hand-written JavaScript — and I kept losing afternoons to the same kind of bug. The Python was fine. The JavaScript was fine. The route one asked for and the route the other served were one character apart, and nothing I already had read both sides of that.

So I wrote something to find them, for myself, on that project. It kept catching things I would not have found on my own, and after a while it seemed like other people might have the same afternoons to lose. So here it is.

It does not catch everything, and I am sure there are things it gets wrong. When it cannot tell, it tries to say uncertain rather than guess. If you find it being confidently wrong somewhere, please open an issue — that is the most useful thing anyone can send me.

What changed, per release: CHANGELOG.md.

What it looks like

Four stacked tiers - the browser, the seam, the server, the store - with one chain lit from a page through its module and its request to the route, the handler and a Redis key

The four tiers, and one chain through them. A page and its module are the browser; the request it makes is the seam; the route and handler are the server; the key they read is the store. Nothing on this picture is inferred from a name — every hop on the right is a line of source the scan read. The second request in the seam is the unresolved one: it goes nowhere, and nothing else in the project would have said so.

The store band of the map: Postgres, Redis and Firebase in separate lanes side by side, each with its oracle badge, and below them the deployables beside each other, split by service and language

Three data stores and three services, one screen. Postgres has a schema to check against; Redis has none, so it can only ever show that two halves of your own code disagree; Firebase has rules. Seven findings are visible before a card is read — a missing row-security policy, a table nothing migrates, a Firestore collection with no rule, a cache key with no expiry, and two renamed background jobs: one in Node, one in Django.

One request followed from the browser to the cache, five hops left to right with arrows on the wires

Click a finding and follow it. Five hops, browser to cache, with the direction on the wires. This one ends on a Redis read in a TypeScript service, of a key a Python service writes — one character apart. No compiler on either side spans that gap.

Four JavaScript and TypeScript files converging on one DOM element

Four files writing one element, in two languages. The element is there, and found four times. Whichever runs last wins, which is how a display bug survives being "fixed" in one of them.

One function's world: the route it answers, the Postgres table and three Redis keys it reaches through its helpers, and the background task it queues - with a line under the canvas counting them

One function, and what it costs. Type three letters and pick submit_push: the map draws everything that function touches - following the calls, because a handler that delegates owns almost none of it itself - plus one hop out to what reaches it. The line under the canvas counts the round-trips per call. This handler is meant to be Redis-only, and it writes Postgres once. That is the whole diagnosis, and at thirty thousand concurrent users it is the difference between a cache read and a connection out of a pool of 45.

A page, then its sections. The map is not one drawing of the whole codebase — the Page picker lists your HTML pages, and Section lists the scripts each one loads: the code that actually runs from that script tag, followed through every import to the selectors, URLs and keys it touches. A widget on a forty-module page is a section you can open alone; Whole page is all of them at once. Whatever no page ever reaches sits in the Not reached from any page buckets, which is a finding in itself.

...and then the function. The third picker is the one you reach for while you are writing code: start typing and every function in the project is offered, prefix first. Picking one leaves the pages behind entirely - a function's symbols are never all on one page, since the page holds its route, the store layer holds its keys, and whatever nothing reaches sits in a bucket - and draws its own world instead: what it touches, what its helpers touch, and one hop out to whatever reaches those. Called by lists everything that calls it, each one a click. Building something new, the holes are the drawing: an unresolved request means the backend is not there yet, an unused route means the frontend is not calling it yet, a key written and never read means nobody consumes it. More on reading the map →

The four words

Every symbol gets exactly one. Nothing is counted twice, and nothing is dropped:

connected Something reaches it, and the evidence is attached.
unresolved Something reaches for it by name and it is not there. Usually a bug.
unused Both ends are visible and nothing connects them. Usually a decision.
uncertain No evidence either way. Never a claim that it is dead.

uncertain is the important one. A route assembled at runtime genuinely cannot be known by reading source, and I would rather it said so than guessed. Every uncertain names the evidence it is missing.

Two numbers, two different denominators, and quoting one as if it were the other is the mistake this page used to make:

  • Coverage — verdicts ÷ symbols. How much of my project can it speak to at all?
  • Precision — true claims ÷ claims. When it says something is broken, is it right?

Precision says nothing about uncertain, because uncertain is not a claim. A backend answering uncertain to everything would score flawless precision and be useless.

Turning uncertain into evidence

Some of it can never be settled by reading source. A selector assembled from a variable, a URL concatenated at call time — no reader resolves those, and that is the floor of what static analysis can know. The browser knows, though.

pip install 'seamcheck[observe]'
seamcheck observe          # visit the pages the graph knows about

It drives your running app with a probe installed ahead of the app's own scripts, and records every selector actually queried and whether it found anything, every URL actually requested, and every class actually applied. That evidence is keyed to the commit, and it converts uncertain rows into answers instead of guesses.

It also settles the one finding no amount of reading can. A multi-writer report says two files write the same element — a risk, not a defect: it becomes a defect when the two disagree. At runtime that has a signature, so the run watches every multi-writer element sit still for twelve seconds with nothing touching the page. A value that moves while the page is idle is the finding that is real. One that never moves says the writers coexist. One the page never rendered is untested, not clean — and is reported that way rather than as a pass.

With one caveat it states rather than hides: a page the run never visited leaves no trace, and looks exactly like a page that is broken. So everything it promotes is labelled as observed, and uncertain going down is always traceable to a specific run over specific pages. The goal was never a smaller number — it is a number backed by something.

Where it stands

Measured across 47 open-source projects, regenerated by python tools/coverage.py:

backend repos symbols judged coverage ceiling
Flask 5 9,689 9,056 93% 94%
Django 21 332,177 303,560 91% 94%
Express 6 4,854 3,387 69% 69%
FastAPI 5 4,997 2,908 58% 58%
NestJS 4 3,587 1,759 49% 49%
Next.js 6 3,541 1,643 46% 46%
all 47 358,845 322,313 89% 92%

Ceiling is where coverage would land if every missing reader were written; the gap between the two columns is the to-do list, and everything below the ceiling is evidence that is not in the repository at all.

Django is the one being finished first, deliberately — it is used every day against a large production codebase, so a wrong finding gets noticed the same afternoon. Django's 84% → 91% is not the same 121,248 symbols scoring better: the ORM lens now reads tables, columns and the querysets that touch them, so a Django project's Postgres half went from 65 model names to 332,177 symbols. A percentage whose denominator has tripled is not comparable to the one before it, and the honest reading is that there is far more of a project in the map, judged at about the same rate. Precision is 54% on hand-labelled findings, up from 28%, and that number moves because people tell me what it got wrong — how to check it yourself, including the ways I got it wrong.

Detail: coverage per backend · what it has actually found · how to check this yourself — the same instrument, the protocol, and the four ways a careful person gets the answer wrong (all four made here)

In CI

seamcheck check --since $BASE_SHA

Exit 1 on new findings, 0 when clean, 2 if $BASE_SHA has no stored snapshot to compare against yet (run seamcheck scan once on that commit to fix that, permanently). --since is what makes it adoptable: it fails only on what your branch added, so you can turn it on today against a codebase with three thousand open findings and it will pass. No token, no network, no model — nothing per run and nothing per repository.

More

Install, per OS · Reading the map · The commands · The data layer · Using it from an agent · Telling me it got something wrong · Checking this tool

How it differs from Knip and depcheck: they work inside one language's module graph — unused files, exports, dependencies — and do it well. This looks at the boundaries between languages. Not competitors; on a TypeScript codebase, running both is reasonable.

Contributing

Issues and pull requests welcome — CONTRIBUTING.md. The most useful thing anyone can send is a finding that is wrong, and why.

License

MIT. Take it, fork it, improve it.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

seamcheck-0.13.0.tar.gz (907.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

seamcheck-0.13.0-py3-none-any.whl (686.8 kB view details)

Uploaded Python 3

File details

Details for the file seamcheck-0.13.0.tar.gz.

File metadata

  • Download URL: seamcheck-0.13.0.tar.gz
  • Upload date:
  • Size: 907.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for seamcheck-0.13.0.tar.gz
Algorithm Hash digest
SHA256 0693fe1c6ffcfc1ddace9e4900d1baddfc5230874801f7604364b2e1fcb16ceb
MD5 eff5dbd8b56bd23d701f4723f0e6f89c
BLAKE2b-256 1d472347a6b2a776185012198532199bbd500e4e71d61f8d86fc0d1fb42efd83

See more details on using hashes here.

Provenance

The following attestation bundles were made for seamcheck-0.13.0.tar.gz:

Publisher: release.yml on dardameiz/seamcheck

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file seamcheck-0.13.0-py3-none-any.whl.

File metadata

  • Download URL: seamcheck-0.13.0-py3-none-any.whl
  • Upload date:
  • Size: 686.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for seamcheck-0.13.0-py3-none-any.whl
Algorithm Hash digest
SHA256 fe561f66c8da6191f5efda29e4fae2197073cfd567e434f94289b7369e3ea842
MD5 41bab020860861e4888b9d78a6dd1770
BLAKE2b-256 38116c946cc8a4e22ffd5d605c41c789b98455a4582033c2132ae8656f292465

See more details on using hashes here.

Provenance

The following attestation bundles were made for seamcheck-0.13.0-py3-none-any.whl:

Publisher: release.yml on dardameiz/seamcheck

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.14.0

2 files

This release

0.13.0 This release

2 files

0.12.1

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.2

2 files

0.8.1

2 files

0.8.0

2 files

0.7.1

2 files

0.7.0

2 files

0.6.1

2 files

0.6.0

2 files

0.5.0

2 files

0.3.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page