Skip to main content

Static code analysis + knowledge-graph for AI coding agents — 12 MCP tools, 13 languages, cut token costs by 87–92 %

Project description

Aethvion Project Mapper

Give your AI coding agent a living map of your codebase.
Scans once, builds a knowledge graph, answers architecture queries in milliseconds — at 87–91% fewer tokens than reading raw files.
Runs entirely on your machine. No data leaves your computer.


Quick Start

Step 1 — Install uv (one-time, manages Python automatically):

# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Step 2 — Install Project Mapper:

uv tool install "aethvion-project-mapper[languages]" --python 3.10

Step 3 — Connect to your agent:

Add Project Mapper to your agent's MCP config, then restart the agent. The exact one-liner or JSON block per agent is below under Connect your agent.

For Claude Code:

claude mcp add -s user project-mapper -- pm-mcp --db workspace

Step-by-step per-agent guides: docs/howto/

New here? The docs/explained/ folder covers what Project Mapper is, what MCP tools are, exactly what PM reads and stores on your machine, and the full tools reference.


Why it exists

AI coding agents (Claude Code, Cursor, Copilot, etc.) search through your files on every task — reading source files, following imports, grepping for context. That's expensive and slow, and context windows fill up with semi-relevant content before the agent sees what actually matters.

Project Mapper scans your codebase once, builds a structured knowledge graph of every module, class, function, and their relationships, and lets agents query only what they need — in milliseconds, reducing token costs by 87–91% on average.


Benchmark numbers

Measured across 11 real-world codebases — Python, Java/Kotlin, C#, PHP, C, Go, Ruby, TypeScript/JS, Rust, C++, Swift — ranging from 57 to 10,437 files.

Summary (Geometric Mean across all 11 projects)

Mode Token Reduction Speedup vs Normal
PM Full ~87% ~380× faster
PM Slim ~91% ~380× faster

At 100,000 input tokens, PM typically uses ~13,000 (Full) or ~9,000 (Slim) tokens.

PM Slim returns name + file path + line number only — enough for navigation and refactoring tasks. PM Full returns complete entity context. See the full benchmark suite for per-codebase numbers and the security benchmark for a 3-test audit comparison.

Query latency (measured)

Query Latency
Context query 1–100 ms (warm cache)
Impact query < 1–10 ms

Session startup — entity map load (measured)

The entity map is stored as a single snapshot file built at the end of each scan.

Codebase size Load time
~400 entities < 50 ms
~12,000 entities ~145 ms
~33,000 entities ~300 ms

Financial impact at scale (modelled)

Modelled from the measured ~87% Full / ~92% Slim token reduction (geomean, 11 codebases). Assumes 10 tasks/dev/day, 8 turns/task, Claude Sonnet pricing. See the cost calculator for your own numbers.

Team size Monthly AI coding cost (est.) Savings with PM
Solo developer $80 $74
10-person team $2,400 $2,230
100-person team $48,000 $44,600
Enterprise (1,000 devs) $480,000 $446,400

What it does

  1. Static scan — walks your project, extracts every module / class / function via AST analysis. No AI needed for this step.
  2. Knowledge graph — stores entities + relationships (imports, calls, extends, depends_on, …) in a local JSON database.
  3. Agent queries — 12 MCP tools that agents call instead of reading raw files:
Tool What it answers
pm_context "What should I know before touching the auth system?"
pm_impact "What breaks if I change UserService?"
pm_path "How does RateLimiter connect to the payment flow?"
pm_find "Where is validateToken defined?"
pm_orphans "What code is never called?"
pm_visualize "Show me the dependency graph for UserService"
pm_security "Are there security vulnerabilities in this codebase?"
pm_security_triage "Mark this finding as a false positive (or a confirmed bug)"
pm_contribute "Record that I added rate limiting to endpoint X"
pm_stats "What's already indexed in this database?"
pm_delta "What changed since the last scan?"
pm_scan "Scan this project directory right now"

See the PM Tools Reference for full documentation on each tool.

Security scanning

pm_security is a standalone SAST-style scanner built into Project Mapper. One call checks your entire codebase — no scan dependency, no setup beyond installation.

  • 140+ patterns across OWASP Top 10 (A01–A10) in 8 languages: Python, TypeScript/JS, Java, Go, C#, PHP, Ruby, C/C++
  • Runs in 1–5 seconds on codebases of any size
  • Route-reachability taint tracking — ⚡ flags findings confirmed reachable from HTTP handlers
  • CWE mapping on every finding, stable finding IDs for triage persistence, snapshot delta across scans

It's designed to work alongside pm_context: pm_security catches pattern-detectable vulnerabilities across 100% of files in seconds; pm_context then closes the logic-flaw and IDOR gaps that patterns can't detect. See the security benchmark for a 3-test comparison on OWASP Juice Shop.


Other access methods

HTTP API

# Install
pip install aethvion-project-mapper

# Start server
pm-server
# pm-server --host 127.0.0.1 --port 8080 --workers 2

# Scan your project
curl -X POST http://localhost:7474/api/project-mapper/scan \
  -H "Content-Type: application/json" \
  -d '{"project_root": "/path/to/your/project", "enrich": false}'

# Query context for a task
curl -X POST http://localhost:7474/api/project-mapper/query/context \
  -H "Content-Type: application/json" \
  -d '{"q": "add rate limiting to auth endpoints", "detail_level": "medium"}'

Docs at http://localhost:7474/docs

Developing from a repo clone instead of a pip install? Use uvicorn server:app --reload --port 7474 from the repo root — server.py is a dev shim, not part of the installed package.

Docker

docker compose up
# Server running at http://localhost:7474

Mount your projects:

# docker-compose.yml — set PROJECTS_DIR to your code root
PROJECTS_DIR=/home/you/code docker compose up

MCP stdio (Claude Code / Cursor / Antigravity / Codex)

Detailed step-by-step setup guides (including Windows and Linux/macOS paths) are in docs/howto/.

A single global config gives every session access to Project Mapper. The AI passes the project root when it calls pm_scan, so you don't need to specify it upfront — just tell Claude (or Cursor, etc.) to scan the current project and it handles the rest.

Step 1 — install uv (one-time, skippable if you already have it):

# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

# Already have Python? Alternatively:
pip install uv

uv manages its own Python environment — the curl/PowerShell installers above work even if Python is not installed.

Step 2 — install Project Mapper (one-time):

uv tool install "aethvion-project-mapper[languages]" --python 3.10

This installs pm-mcp as a global command and downloads all language parsers. Takes ~30 seconds on first run, instant on subsequent starts.

--python 3.10 tells uv to fetch and pin the exact Python version Project Mapper is tested on, into an isolated environment — so it never depends on (or collides with) whatever Python is already on your system. This is what makes the install reproducible: every user runs the same interpreter and the same locked dependencies.

Step 3 — connect your agent:

Add the Project Mapper entry to your agent's MCP config below, then restart the agent. These are the only files involved — nothing edits your config for you, so you stay in control of what changes.

Claude Code — run this (registers Project Mapper for all your projects):

claude mcp add -s user project-mapper -- pm-mcp --db workspace

⚠️ Do not put mcpServers in ~/.claude/settings.json — Claude Code ignores MCP definitions there. It only reads them from ~/.claude.json (which claude mcp add writes) or a project-level .mcp.json.

If you don't have the claude CLI on your PATH, create a .mcp.json in your project root instead:

{
  "mcpServers": {
    "project-mapper": {
      "type": "stdio",
      "command": "pm-mcp",
      "args": ["--db", "workspace"]
    }
  }
}

Cursor — add to ~/.cursor/mcp.json:

{
  "mcpServers": {
    "project-mapper": {
      "command": "pm-mcp",
      "args": ["--db", "workspace"]
    }
  }
}

Antigravity (Google) — add to ~/.gemini/antigravity/mcp_config.json:

{
  "mcpServers": {
    "project-mapper": {
      "type": "stdio",
      "command": "pm-mcp",
      "args": ["--db", "workspace"]
    }
  }
}

Codex CLI — add to ~/.codex/config.json:

{
  "mcpServers": {
    "project-mapper": {
      "type": "stdio",
      "command": "pm-mcp",
      "args": ["--db", "workspace"]
    }
  }
}

Alternative — no install, run directly with uvx:

{
  "command": "uvx",
  "args": ["--from", "aethvion-project-mapper[languages]", "pm-mcp", "--db", "workspace"]
}

Restart the agent after editing the config.

Optional — pin to a single project:
If you always work on one codebase, add PM_PROJECT_ROOT so the AI never needs to specify it:

{
  "mcpServers": {
    "project-mapper": {
      "...",
      "env": { "PM_PROJECT_ROOT": "/absolute/path/to/your/project" }
    }
  }
}

The workspace database is shared — scanning a new project overwrites the previous one. This is fine for single-project sessions; incremental scans on a pre-indexed repo typically finish in under 2 s.


Configuration

Variable Default Description
PM_DATA_DIR ~/.aethvion/project-mapper Root directory for all databases
PM_LOG_LEVEL INFO Log level: DEBUG / INFO / WARNING / ERROR
PM_DB_NAME default MCP server: database name
PM_DB_PATH (unset) MCP server: explicit database path
PM_PROJECT_ROOT (unset) MCP server: default project root for pm_scan

Project structure

The package is layered: analyzers → core → {mcp, http}, with config and db as shared infrastructure. Interface layers (mcp, http) are thin adapters over the transport-agnostic core; core never imports from them.

project_mapper/
├── config.py             — DATA_DIR / environment config
├── analyzers/            — language extractors (one file per language)
│   ├── base.py               shared types (CodeAnalysis) + language dispatch
│   └── python.py  typescript.py  java.py  kotlin.py  go.py  rust.py
│       c.py  cpp.py  csharp.py  ruby.py  php.py  swift.py
├── core/                 — transport-agnostic engine
│   ├── scanner.py            async background scan engine
│   ├── ingestor.py           CodeAnalysis → AethvionDB entities
│   ├── query_cache.py        query-result cache (snapshot-mtime invalidation)
│   ├── cleanup.py            incremental scan maintenance
│   ├── delta.py              filesystem diff (no DB writes)
│   ├── watcher.py            auto-scan file watcher (--watch mode)
│   ├── query/               graph queries, one module per primitive
│   │   ├── _common.py            entity-map build, resolution, shared constants
│   │   └── context.py  impact.py  path.py  find.py  orphans.py  contribute.py
│   └── security/            standalone SAST scanner (OWASP Top 10, 140+ rules)
│       ├── patterns.py          the pattern catalog (data)
│       └── scanner.py           the scan engine (logic)
├── db/                   — storage
│   ├── entity_schema.py      entity data model + validation
│   ├── pm_store.py           in-memory entity store (PMEntityStore)
│   ├── name_index.py         thread-safe name → ID index
│   ├── file_manifest.py      file ↔ entity provenance tracking
│   ├── snapshot.py           single-file snapshot persistence
│   ├── db_registry.py        named database registry
│   └── utils.py              logging + atomic-write helpers
├── mcp/                  — MCP (stdio) interface
│   ├── server.py             JSON-RPC 2.0 stdio MCP server
│   └── tools/               12 tools, one module per tool (+ base.py)
└── http/                 — HTTP (FastAPI) interface
    ├── app.py                app factory + pm-server CLI entry point
    └── routes.py             FastAPI router (/api/project-mapper/*)

server.py                 — FastAPI dev entry shim (uvicorn server:app)

Incremental scanning

Subsequent scans only process files whose SHA-256 hash has changed since the last run. On a 10,000-file repo that's been scanned before, incremental mode typically processes < 1 % of files — scan time drops from ~60 s to < 2 s.

# Full scan (first time or force refresh)
curl -X POST .../scan -d '{"project_root": "...", "incremental": false}'

# Incremental scan (default — only changed files)
curl -X POST .../scan -d '{"project_root": "..."}'

License

Open-source core: GNU AGPL v3
Free to use, modify, and self-host. Network use requires open-sourcing your modifications.

Commercial license: COMMERCIAL_LICENSE.md
Available for teams that need a proprietary license, SLA, or integration support.


Contributing

Pull requests are welcome. By submitting a PR you agree to the Contributor License Agreement (§2).

git clone https://github.com/Aethvion/Aethvion-ProjectMapper
cd Aethvion-ProjectMapper
pip install -e ".[dev]"
pytest

Support development

If Project Mapper saves you tokens, time, or money — consider sponsoring. It keeps the project maintained and new languages / features coming.

GitHub Sponsors

See the full sponsors list.

Built with care by the Aethvion team.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

aethvion_project_mapper-2.1.0.tar.gz (176.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

aethvion_project_mapper-2.1.0-py3-none-any.whl (210.7 kB view details)

Uploaded Python 3

File details

Details for the file aethvion_project_mapper-2.1.0.tar.gz.

File metadata

  • Download URL: aethvion_project_mapper-2.1.0.tar.gz
  • Upload date:
  • Size: 176.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.11

File hashes

Hashes for aethvion_project_mapper-2.1.0.tar.gz
Algorithm Hash digest
SHA256 1c8e1c8ab9908dc2995c4eabb47fc33416aeb9dd7246d7287fd1b8cbfaf32900
MD5 e60d4e19717da8f77f3064f761881033
BLAKE2b-256 23061636c28a9354696853ff493d407c5b408e9aa9d7f5747856d7fff7093757

See more details on using hashes here.

File details

Details for the file aethvion_project_mapper-2.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for aethvion_project_mapper-2.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 43552e8fd63e560b2a082ab08bdc5049dafd444c64379071b2d791acbd6880e7
MD5 154d9ae454ef24aefb7e341f33994902
BLAKE2b-256 5c4fc775076fd5f063d8d2f074fa4593d2e31e8b8ff1fcfd6ac1d38aaa92a02e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page