Skip to main content

codegraph

Ask your Python codebase what breaks if you change this.

One file. No dependencies. Nothing but the standard library.

curl -O https://raw.githubusercontent.com/jedisolana/codegraph/main/codegraph.py
python3 codegraph.py build .
python3 codegraph.py impact _paid_ok
callers: boardofdirectors/server.Handler.do_POST, boardofdirectors/server._board,
         boardofdirectors/server._single, boardofdirectors/server._tier
sites:   boardofdirectors/server.py:201, boardofdirectors/server.py:250,
         boardofdirectors/server.py:279, boardofdirectors/server.py:632,
         boardofdirectors/server.py:635, boardofdirectors/server.py:665,
         boardofdirectors/server.py:679, boardofdirectors/server.py:697,
         boardofdirectors/server.py:712
blast:   19 functions could be affected

That is a real answer from a real codebase, not a mock-up. Four callers, the nine exact lines to open — a caller that checks a gate five times is five places to edit — and the transitive radius, before you touch anything.

Why

grep finds the name. It cannot tell you that two files define a function called digest and only one of them is the one you are about to break. Ask this for digest and it will not guess either — it names both and waits for you to say which, because merging their callers into one answer is how you end up "fixing" a caller of the other one. Your editor's "find references" can, but it needs a language server running, and it will not give you the transitive answer: who calls the callers.

This is the question you actually have before an edit — what depends on this — answered in one command, from a file you can copy into any repo.

It tells you when it doesn't know

This is the part that matters, and it is why the tool is worth having rather than clever.

Every call edge carries a confidence label:

label what it means resolved?
LOCAL a bare call to a function this scope can see — the same module, or a function it is nested inside yes
QUALIFIED thing.load(), where thing is a module this file actually imported yes
SELF-METHOD self.helper(), resolved inside the enclosing class yes
INHERITED self.method() or super().method(), where the method lives on a base class yes
CLASS Parent.method() — the receiver is a class in this module yes
TYPED the receiver's class is known: x = Foo(), x = svc.Foo(), or an annotation that says so yes
CONSTRUCTOR Client() — the second, equally real edge, to the __new__ and __init__ it runs yes
AMBIGUOUS several definitions match; the candidates are listed and none is picked no
BUILTIN len(), open(), sorted() — certainly not yours no
EXTERNAL a library, the stdlib, or a method whose name nothing in your tree defines no
UNTYPED the receiver could not be typed, and the target might be yours no

Two of those deserve the detail the table cannot hold.

TYPED decides which Foo by what the file defines or imports, never by a tree-wide name search. It looks the method up the inheritance order, so an inherited one is found and an override wins. It refuses when two branches give x two types, when a later line rebinds it to something it cannot name, and when two classes answer to the name. Calling an instance — c(1) where c is a Client — resolves to that class's __call__, since the call site never writes the name.

CONSTRUCTOR and INHERITED exist because of the same problem: some calls never mention what they run. Client() runs an __init__; super().save() runs a parent's save; c(1) runs a __call__. Without those edges the tool answers "nothing depends on this" about the most-edited method in Python.

The language calls things the source never names. with r: runs __enter__ and __exit__; for x in r runs __iter__; len(r) runs __len__; r[k], r[k] = v and del r[k] are three different methods; r + 1 runs __add__ and 1 in r runs r's __contains__ — the one case where the receiver is the operand on the right. All of these are edges. In the standard library 2,780 definitions are reached only this way, and 98% of them used to report no caller at all.

Defining class Child(Base) runs Base.__init_subclass__. Building a dataclass runs __post_init__, from an __init__ that is generated and so is nowhere in the graph. An f-string placeholder runs __format__, or __repr__ for {x!r}.

Iteration counts however it is written — a for statement, a comprehension, a generator expression, a, b = r, [*r], yield from r. And x += y runs __iadd__ when the class has one and __add__ when it does not, which is decided against the class rather than guessed at the line.

Reading a property runs it, and writes no parentheses doing so — c.endpoint is a call with nothing in the syntax to say so. Attribute reads are counted wherever the receiver can be typed, the same reach a method call has: self and cls inside a class, or a local whose class is known, resolved through the same inheritance order. In the standard library that is 807 definitions, 331 of which really are read somewhere and used to report no callers at all.

Most call-graph tools guess and hand you one answer. This one refuses. An AMBIGUOUS edge is a real result: it means the question does not have a single answer, and a blast radius that quietly picked one would be worse than no blast radius at all.

UNTYPED is the one that matters, and it is a claim rather than a shrug: the receiver could not be typed, and the target might be in your tree. x.run() where something of yours defines a run. It is deliberately not lumped in with EXTERNAL, because that would claim knowledge it does not have.

The claim has to be true, though. rows.append(1) and text.strip() are not calls it failed to place — nothing in your tree is named append or strip, so the target cannot be here, and the tool can prove that rather than guess it. Those are EXTERNAL. On one real codebase nine of every ten "cannot tell" edges were .get(), .items(), .join() and .assertEqual().

stats reports that rate over what was winnable — resolved, plus ambiguous, plus the calls it could not type. Builtins and library calls are excluded, since counting them would only measure how much of the standard library you happen to use:

Run on this repository, so you can reproduce it — codegraph build . && codegraph stats:

{
  "call_edges": 3801,
  "call_sites": 5070,
  "edge_confidence": {"EXTERNAL": 2023, "INHERITED": 601, "BUILTIN": 444, "SELF-METHOD": 299,
                      "QUALIFIED": 257, "LOCAL": 119, "UNTYPED": 54, "CONSTRUCTOR": 3,
                      "TYPED": 1, "AMBIGUOUS": 0},
  "resolved_to_one_def": 1280,
  "could_have_been_resolved": 1334,
  "resolution_rate": 0.96
}

That 0.96 says: of the calls that could plausibly have gone to something in this codebase, it placed 96%. It is not the sum being flattered — the rule is the opposite of the usual one. A denominator that counts list.append and str.strip is not measuring how much the tool resolved, it is measuring how much of Python you happen to use, and the same reasoning that keeps builtins out keeps those out. What remains in it are the calls that genuinely might have been yours.

The number this README first published was 0.29, over a denominator that swept in every impossible call. Two things moved it: resolution bugs fixed with tests, and then that denominator being made to mean something. Both directions are in CHANGELOG.md, with the count of fabricated edges each one removed.

There used to be one more label. RESOLVED meant "a bare call, and exactly one definition of that name exists somewhere in the tree" — which is a coincidence, not a resolution. By the time every real way a bare name reaches a definition had a rule of its own, it fired fifteen times across an entire standard library and not once across 2,725 installed packages, and the fifteen were wrong: turtle builds up, down, left and right at import time rather than defining them, and those calls were being answered with functions in _pyrepl. A name this file neither defines nor imports nor stars in is a library, something injected at runtime, or a mistake — and saying EXTERNAL is the true answer to all three.

The commands

codegraph build [dir...]     build the graph (default: here); writes codegraph.json
codegraph impact <name>      callers + call sites + blast radius, and what it could not resolve
codegraph callers <name>     who calls this
codegraph calls <name>       what this calls
codegraph blast <name>       transitive callers — what could break
codegraph sites <name>       every call site as file:line, relative to what you built — all of them
codegraph where <name>       where a symbol is defined
codegraph find <substr>      fuzzy symbol search
codegraph path <from> <to>   a call path connecting two functions
codegraph deps <module>      a module's in-tree imports and importers, `import_module("x")` included
codegraph cycles             import cycles of any length, including a module importing itself
codegraph symbols            every function and class defined here
codegraph unused             every definition nothing here calls — read the caveat
codegraph stats              counts, resolution rate, never-called definitions
codegraph --selftest         32 ground-truth checks, several of them red-first
codegraph --help             the same list; a bare `codegraph` prints it too
--json                       any query, answered as data instead of prose
--only PAT / --exclude PAT   keep or drop results — a glob over the id or its module,
                             so `--exclude 'tests/*'` and `--exclude '*.Handler.*'` both work

Exit codes, because scripts and agents read them:

code meaning
0 answered — including a real function with no callers
1 this graph has never heard the name, or a search matched nothing
2 the name matches several definitions, or the command was malformed

"Nothing depends on this" and "I do not know that name" are deliberately different answers. A misspelling that returns success is how an agent talks itself into an unsafe edit.

Queries find the graph by walking up from where you are, the way git finds .git, so you can ask from anywhere in the repository. They rebuild automatically when the tree has changed — a file edited, added, or deleted. The graph records the size and timestamp of every file it read and compares them exactly, rather than asking whether anything is newer than itself: a file restored from a backup, a checkout or a container layer arrives older than the graph while holding different code. Deletion is the one people forget: removing a file changes nobody else's timestamp, so a graph that only watched timestamps would go on answering about code that is gone. The rebuild is incremental — unchanged files are reused from a cache keyed by modification time and by codegraph's own source hash, so editing the parser invalidates every stale parse instead of silently reusing it. The graph carries that hash too: upgrade codegraph and the next query rebuilds, rather than answering from the version you replaced.

For an AI agent

An agent that edits code reads it as text and guesses at the consequences. Give it structure instead:

import codegraph
g = codegraph.load()
codegraph.impact(g, "spend_cap")   # {"callers": [...], "sites": [...], "blast": [...]}

impact before an edit is the difference between "I changed a function" and "I changed a function that five others depend on, here are their line numbers." The output is small, exact, and cheap — no model call, no network.

From a shell, add --json and every verb answers in data rather than prose — including the refusals, so a misspelling and a function with nothing calling it stay distinguishable without matching on an English sentence:

$ codegraph impact leaf --json
{"query": "impact", "target": "app.leaf", "callers": ["app.mid"],
 "sites": [{"at": "app.py:4", "caller": "app.mid"}], "blast": ["app.mid", "app.top"],
 "unresolved": []}

$ codegraph callers digest --json        # exit 2
{"error": "ambiguous", "name": "digest", "detail": "2 definitions answer to that name",
 "candidates": [{"id": "pulse.digest", "at": "pulse.py:9"}, ...]}

The prose form of impact says the blast radius holds five functions; the JSON says which. That gap existed because the prose was written for a person reading it and nothing was checking the other reader.

The library refuses exactly what the command line refuses, through the same code: codegraph.Ambiguous when several definitions answer to the name, carrying the candidates, and codegraph.Unknown when this graph has never seen it. Neither is answered with an empty result, because "nothing depends on this" is the one reply a misspelling must never get.

What it cannot do

Static analysis, honestly labelled:

  • Python only. A file it cannot parse — Python 2, a template, something half-written, or a generated one nested deeper than the interpreter's own stack — is named on stderr and left out, never turned into an empty module in silence.
  • A conditional import has no single answer. try: from fast import parse / except: from slow import parse binds one name from two modules, and which one runs depends on the machine. Both are listed as candidates; neither is chosen.
  • Dynamic dispatch defeats itgetattr(obj, name)(), dispatch tables, monkeypatching, plugin registries. These land as EXTERNAL, which is the truthful answer.
  • Definition-time code counts as calls — a default value, an annotation, a decorator all run beside the def, so they are edges from the enclosing scope.
  • A decorator is counted as a call@register is an edge from the enclosing scope, since that is where it runs. But a decorator that replaces the function with a different one is not followed through: calls to the decorated name still point at the original def.
  • A call made by syntax needs a receiver the file can type, the same as a written one. with open(p) as f and with self.lock name no typed local, so neither is recorded. When the class is known and does not have the method, the edge is EXTERNAL rather than "cannot tell" — for row in rows on a list subclass runs code outside the tree.
  • An augmented assignment picks the method the interpreter would. x += y resolves to __iadd__ if the class has one and to __add__ if it does not — never to both, since only one of them runs.
  • A reflected operator is not followed. a + b records a.__add__; Python's fallback to b.__radd__ happens only when the first returns NotImplemented, which is a runtime answer.
  • A property read needs a receiver the file can type. self.thing and c.thing where c is annotated or was built here both resolve; a bare x.thing on an untyped x does not, the same way x.method() does not. Every typed attribute read is recorded while parsing and dropped afterwards unless the name turns out to be a property, since one file cannot know what another one defines.
  • A function passed by name is not a callThread(target=f), sorted(key=f), partial(f, 1). f is referenced, never invoked here, so it shows in unused. That listing says "nothing here calls this", which is not the same as dead, and says so.
  • super() resolves against the class it is written in. Python's own order for that class - so a diamond lands where the interpreter lands. What static analysis cannot know is that B.m's super() goes to C when B is reached through a D(B, C) instance.
  • A base class is found the way any other class name is — defined in this module, or imported into it, before any tree-wide search; an alias (from x import Base as B) is followed. Method lookup then uses Python's own order, C3 linearisation, so a diamond resolves where the interpreter resolves it. What is not followed: a base built by a metaclass, a base that is a variable (V = Generic[T] then class C(V)), and a subscripted one — Generic[T] names Generic, and a subscript is not a name.
  • Type inference is one line deepx = Foo() then x.method(), plus annotations, which say it outright: a parameter's, and a variable's. A container annotation is not its contents, so Dict[str, Client] stays a dict. A name rebound to anything the tool cannot name loses its type rather than keeping the old one. Nothing beyond that: no return types, no attributes, and with Foo() as c is not assumed to give you a Foo, because __enter__ may return anything at all.

Everything it cannot resolve is labelled rather than guessed, so the limits are visible in the output instead of hidden in it.

Why not something else

  • grep / ctags — finds names, not relationships. No transitive answer, and it cannot tell two same-named functions apart.
  • pyan, code2flow — call graphs, but they want Graphviz and hand you a picture. This hands you an answer to a question, and installs nothing.
  • pydeps — module imports only, not function calls.
  • pycg — more rigorous, and a heavier dependency; worth it if you need academic precision.
  • A language server / SCIP indexer — more accurate than this, and correspondingly large. If you already run one, use it. If you want an answer in a shell script or a CI job or an agent loop, that is what this is for.

Install

The point is that you don't have to:

curl -O https://raw.githubusercontent.com/jedisolana/codegraph/main/codegraph.py

If you'd rather have it on your PATH:

pipx install git+https://github.com/jedisolana/codegraph

On PyPI it will be jedi-codegraph, not codegraph — that name belongs to somebody else, and PyPI reads hyphens as if they weren't there. The command is codegraph either way.

Python 3.9+. Tested on Linux, macOS and Windows.

Proving itself

A suite that never fails is not evidence of anything, so it is checked the other way round. tools/mutation.py breaks the tool one small way at a time — flips a comparison, swaps an and for an or, drops a not, moves a number by one — and runs the suite against each change. Every one of them should make something go red, and one that does not is the interesting output: it names a behaviour nothing is checking.

The file admits 839 mutations. A full pass is hours of work, so it is run deliberately rather than on every push, and the count is checked by a test — it was published as 710 here and 208 in the changelog while the file admitted 839, because a number written twice and checked nowhere drifts in two directions. Two mutations do not make the suite fail but make it never finish — flip the comparison that ends a while — and those are caught by a timeout and counted apart, because "hung" and "failed" are different facts.

Which name shadows which is the question everything else rests on, so it is not only checked against fixtures: symtable is CPython's own scope analysis, and on the versions where its model matches this one, the two are compared scope by scope over the running interpreter's standard library. Fifteen thousand scopes, nothing missed.

python3 codegraph.py --selftest builds small trees with known answers and checks all 32 — including red-first controls that prove the naive approach fails where this one does not:

  • two modules both defining digest, and a query that must reach exactly one of them
  • open('x').write('y') — a method on a file object, which must not resolve to a same-named function in your tree
  • from .thing import load inside a package, next to a top-level thing.py — the trap that makes a lazy implementation return the wrong function with full confidence
  • Mid(1) where Mid inherits its __init__ — and the control showing that matching on the name __init__ finds no caller at all, because no call site contains the word
  • two functions in one module that each define a helper called inner, which a lookup keyed on (module, name) can only tell apart by luck
  • super().run(), whose caller never writes the name of what it calls
  • thing.load() in a file that never imported thing, next to a thing.py that would have answered — and the same trap reached by writing one dot too many in a relative import
  • two classes called Client, one import naming which, and a method call that has to land on the one the file imported
  • try: from fast import parse / except ImportError: from slow import parse — one name, two sources, and a dict that could only remember the fallback
  • def send(c: Client) beside def keyed(d: Dict[str, Client]), where one of the two says what the receiver is and the other does not
  • a class defined inside a function, which a second function cannot name — beside svc.Client(), which any function that imported svc can
  • lambda config: config.dumps(x) in a file that imports a module called config, beside a deferred import config inside a function, which is the same name meaning the opposite thing
  • [config.dumps(r) for config in rows] on one line and config.dumps(2) on the next, where the same name is the loop variable and then the module again
  • from time import sleep in a tree that happens to contain a sleep of its own, and from ops import index as _index, where the module holds index and the file says _index
  • from turtle import * followed by a bare home(), next to another module that also has one

The test suite adds 489 more. Grouped, because a list of every one of them stopped being readable a long time before it stopped growing:

  • Python's own rules, which are where the wrong answers come from: what shadows what — a parameter, a lambda's argument, a comprehension variable, a class attribute, a match capture, a loop target at the top of a file; which Client an import means when two files define one; super() through a diamond, checked against the interpreter's own order; from x import y as z; a star import; a conditional import that has two answers; a base class that is really a variable.
  • Calls that never write the name they call — a constructor, an inherited method, a __call__, a super(). Each of those once returned "nothing depends on this" about code that is called constantly.
  • Trees that fight back — dangling symlinks, a symlink into the tree, a self-linked directory, an unreadable file, a read-only directory, a byte-order mark, a file too deeply nested for the interpreter to walk, folders with dots in their names, two trees whose folders share a name, eight builds racing each other.
  • The graph on disk — cache invalidation by size and timestamp, a file restored from a backup with the timestamp it used to have, deletion (which changes nobody's mtime), a graph built by an older copy of the tool, an interrupted write, valid JSON that is not a graph, and a second build that has to reach the same graph as the first.
  • Both doors — every verb crossed with every state the graph can be in, the library refusing exactly what the command line refuses, exit codes, and a misspelled name that must never answer "nothing depends on this".
  • Its own claims — the counts in this file, the example above, the promises in SECURITY.md (no network, no execution, nothing but the standard library), the label table further up, and every shipped file checked for a private origin story.

Licence

MIT.

Built by @jedisolana.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

jedi_codegraph-0.1.0.tar.gz (148.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

jedi_codegraph-0.1.0-py3-none-any.whl (65.4 kB view details)

Uploaded Python 3

File details

Details for the file jedi_codegraph-0.1.0.tar.gz.

File metadata

  • Download URL: jedi_codegraph-0.1.0.tar.gz
  • Upload date:
  • Size: 148.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for jedi_codegraph-0.1.0.tar.gz
Algorithm Hash digest
SHA256 def081d3863cac2f9ef8335b2fd131698c334bbfa23288e80649e06c5de403f4
MD5 ece55d825c95bee2fa7f640e9d4dbfac
BLAKE2b-256 07efb41a0948a5f540c5daa8449370f7a18f6646a09df90c5ecbdbcc2fb0cd1d

See more details on using hashes here.

Provenance

The following attestation bundles were made for jedi_codegraph-0.1.0.tar.gz:

Publisher: publish.yml on jedisolana/codegraph

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file jedi_codegraph-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: jedi_codegraph-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 65.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for jedi_codegraph-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 04ab8cae1553632ac03401129d88c235d74bd30c865ff7310f81b6199664fbee
MD5 09bb545e0ebc9398bacdc9a34cf291ad
BLAKE2b-256 9739269019ed3af8c7af5254198994d3caf02609506929851e8126a0d0bdcdbc

See more details on using hashes here.

Provenance

The following attestation bundles were made for jedi_codegraph-0.1.0-py3-none-any.whl:

Publisher: publish.yml on jedisolana/codegraph

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page