Skip to main content

codedex

Find your own code by what it does, not what the folder is called.

One Python file. Standard library only. No installs, no model, no server, no network.


The problem

You wrote a Star Citizen themed internet radio player. Two years later you want it back.

It's in Cleanup/Radio_Drake.py.

Neither "Cleanup" nor "Radio_Drake" contains the word radio in a way you'd think to search for, and nothing anywhere on that path says Star Citizen. grep -r radio gives you four hundred hits across every project you've ever written. You give up and rewrite it.

But the file itself already told you exactly what it was, in its first three lines:

Drake Interplanetary — HCN Radio  MODEL R4-HCN
Star Citizen themed internet radio player.
Requires: PyQt6, python-vlc

That description exists in almost every source file you've ever written. It just isn't searchable. So codedex reads the header of every source file once, remembers it, and searches that instead.

$ codedex find radio

3 match(es) for 'radio':

  tools/Drake Radio
     5 files, 5,813 lines, python:5, 34d ago
     | Drake Interplanetary — HCN Radio  MODEL R4-HCN
     | Star Citizen themed internet radio player.
     | Requires: PyQt6, python-vlc
     -> tools/Drake Radio/Radio_Drake.py

It returns projects, not files. The question you're asking is "which thing did I make," not "which lines contain this string."


Install

pip install git+https://github.com/Levatat3/codedex.git

Or just take the single file and run it. It imports nothing outside the standard library:

python codedex.py build

Languages

62 extensions, 29 languages:

Comment style Languages
"""docstring""" Python
// and /* */ JS, TS, Java, Kotlin, Scala, C#, F#, C, C++, Go, Rust, Zig, Swift, Dart, Nim, PHP, CSS/SCSS/Sass/Less, Vue, Svelte, GLSL/HLSL/WGSL
# Ruby, Perl, Elixir, Julia, R, GDScript, shell (sh/bash/zsh/fish/PowerShell)
-- and --[[ ]] Lua, SQL
{- -} Haskell
; Clojure, Lisp, Elisp
REM batch

Non-comment preamble is skipped before the scan — PHP's <?php, shebangs, @echo off, doctypes — so the description underneath is still found.

Adding a language is one line in SOURCE_EXT if it uses one of the styles above.

Tests

python -m unittest -v

38 tests, standard library only — no pytest, nothing to install. Each builds its own throwaway tree, so the suite never touches your real index or your code.


Use

codedex build                  # scan the tree, write the index
codedex find radio             # search
codedex find star citizen      # multiple words are ANDed
codedex list                   # every project, one line each, newest first
codedex stale --days 90        # projects nobody has touched in 90+ days
codedex find qr --scores       # show the scoring, for when ranking looks wrong
Command Does
find <words> Search. Every word must land somewhere or the project is dropped.
list Every project with a one-line description, newest first.
stale [--days N] Projects untouched in N+ days (default 90).
build Rescan. Run after moving things around.

On a tree of ~950 files across 43 projects, a cold build takes about 20 seconds and a warm rebuild about 1.5. Searches after that are instant — they read one JSON file.


Nested projects

By default every top-level folder is one project. That's right for the common ~/code/thing-one, ~/code/thing-two layout.

Some trees nest, though: tools/Star Citizen/logger/main.py should be the project tools/Star Citizen, not all of tools lumped together. Tell it which folders merely hold projects:

codedex build --category tools,games,ai

After a build, codedex points out folders that look like they might need this:

  These folders may hold several projects each rather than being
  one project:  games, tools
  If so, index their children separately with:
      codedex build --category games,tools

It suggests and never applies this automatically, on purpose. Whether AutoCraft is one project with a bot/ and a brain/, or a folder holding two projects, is the author's intent — it cannot be read off the directory shape. I tried; the detector confidently got real cases wrong in both directions. Guessing wrong silently merges unrelated projects, which is worse than asking once.

Configure

Variable Default Does
CODEDEX_ROOT current directory Tree to index.
CODEDEX_CATEGORIES unset Same as --category.
CODEDEX_INDEX unset Exact index file path, overriding the default location.

The index is not written next to the script — it goes to a per-user data directory (%LOCALAPPDATA%\codedex on Windows, $XDG_DATA_HOME/codedex or ~/.local/share/codedex elsewhere), with one file per tree keyed by its absolute path. Several checkouts can each keep their own index without colliding, and nothing is ever written into site-packages.


How ranking works

Each search term scores against five places. Every term must land somewhere, or the project is dropped entirely — that's what makes multi-word searches narrow instead of noisy.

Signal Weight Why
file header / docstring 10 The author said what it is. Most trustworthy thing available.
README prose 8 Also deliberate, but often describes neighbours too.
folder name 6 Weak. Folder names are where this problem comes from.
file name 4 Weaker. utils.py tells you nothing.
body text 2 Last resort, for undocumented code.

Plus a breadth bonus: +3 per additional file mentioning the term, capped at +15.

That cap matters more than it looks. A project that name-drops "discord bot" once in a README should not outrank the actual Discord bot, which says it in file after file. Breadth of mention is the signal, not frequency.

Headers that don't count

A header only scores if it actually explains something. These are discarded:

  • separator lines (# ----------------) and bare filenames (# utils.py)
  • anything under three words
  • vendored licence text — an MIT header describes the licence, not your project, and without this check every vendored dependency matches every search

Undocumented code

Files that open straight into import math have no header to read. Rather than making them invisible, codedex keeps a 700-character slice of the body at the lowest weight, so old undocumented code stays findable by its identifiers. When a match lands there, the output says which file it came from so you know the hit is weaker.


What it is not

Being clear about this up front:

  • Not semantic. It matches words, not meaning. Searching music player will not find a file whose docstring says "audio playback." If you want meaning, you want embeddings, and that means a model and a runtime — deliberately not what this is.
  • Not a code search tool. It reads the first 4 KB of each file, not all of it. For searching inside code, use ripgrep. This answers a different question.
  • Not a documentation generator. It reads what you already wrote. Projects with no docstrings and no README get thin results, and that is a fair reflection of reality.
  • Source files only. 62 extensions across 29 languages (below); everything else, plus build output, dependencies and caches, is skipped.

The upside of all those noes: it runs anywhere Python runs, in one file, offline, with nothing installed and nothing phoning home.


Why it works because your folders are a mess

It doesn't ask you to organise anything. It works precisely because names are inconsistent: FQR2 is found by searching "qr code". Bryan Test by "spotify". Version 5 by "discord bot". The folder name is treated as the weakest signal, not the only one.

If your projects were already named well, you wouldn't need this.


Licence

MIT. See LICENSE.

Metadata

Release files for codedex 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for codedex 0.1.0
File Interpreter ABI Platform
codedex-0.1.0-py3-none-any.whl Python 3 none any Details

Release files / codedex-0.1.0-py3-none-any.whl

Download URL codedex-0.1.0-py3-none-any.whl
Size 16.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
676ad3c64d2e902f325b22f8bf597cfd538a0640b8d2dc37d95e5ea13404dad8
BLAKE2b-256 checksum
How to use checksums
5bb326e9ed1b08313b74075b39fed3bb9fe78b9ec4da9c09b9cc9144f15bdab2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page