Local runtime evidence CLI and MCP server for coding agents
Project description
flameox
Runtime evidence for coding agents
Give an agent profiler traces, benchmark results, memory captures, and execution evidence it can query, compare, and audit—without uploading your code or captures.
Quick start · What flameox investigates · How it works · CLI and MCP · Documentation
Connect your agent: npx flameox setup
flameox is a permanently local CLI and Model Context Protocol server. It connects coding agents to maintained tools such as pyperf, py-spy, Perfetto Trace Processor, coverage.py, Memray, and torch.profiler, then keeps their native artifacts and the provenance needed to check a conclusion later.
flameox is not another profiler and it is not an automatic bug finder. It gives agents a consistent evidence workflow: collect from approved workloads, preserve what happened, compare compatible runs, and keep observations separate from inferences.
Quick start
Connect an agent
Run the guided setup:
npx flameox setup
The wizard detects Claude Code, Cursor, OpenCode, Codex, Gemini CLI, and Antigravity. Nothing is selected by default. It previews every configuration file it will change, installs and verifies a versioned local runtime, and activates only the clients you approve.
Restart the configured client, open a project, and ask:
Initialize flameox in this project and show me which profiling capabilities are available.
flameox can initialize its local .diagnostics/ workspace through MCP. Capturing
code requires one additional safety boundary: declare the command as a named
workload in flameox.toml, inspect its canonical form, and approve it once. The
named workload example shows that flow.
Use the CLI from source
Python 3.12 or newer and uv are required:
uv sync --extra dev --extra python --extra execution --extra memory --extra trace --extra cpu
uv run flameox init .
uv run flameox status
What flameox investigates
| Question | Evidence |
|---|---|
| Where does this workload spend CPU time? | Sampled stacks, frames, callers, callees, and trace windows |
| Does runtime grow with input size? | Repeated measurements, scaling fits, uncertainty, and correlated hotspots |
| Why does memory grow? | Allocation records, retained memory, phases, threads, and processes |
| Which execution paths changed? | Coverage contexts, files, functions, branches, and two-run differences |
| What does PyTorch spend time on? | Operators, shapes when captured, CPU or accelerator time, and memory |
| Are failures clustered rather than isolated? | Failed attempts grouped by environment, source, workload, and error |
Profiles point to candidates; they do not prove causality or correctness. A confirmatory comparison also needs a representative workload, declared metric, compatible source and environment identities, preserved samples, and a semantic oracle.
How it works
- Declare. You approve a repeatable workload. flameox never exposes arbitrary shell commands or SQL to an agent.
- Capture. A maintained profiler or benchmark tool runs while flameox records the exact tool, command, environment, source identity, limits, and outcome.
- Preserve. flameox keeps the native artifact and publishes normalized evidence into an immutable local corpus.
- Analyze. The CLI and MCP server expose the same bounded operations for hotspots, scaling, memory, execution, failures, and comparisons.
- Record. Findings remain tied to the runs, measurements, validation, and analysis that support them. Failed attempts remain visible.
Native artifacts and Parquet evidence are authoritative. DuckDB is a rebuildable local query layer, and Perfetto Trace Processor handles detailed trace queries.
Boundaries
flameox supports local investigation of performance, memory, execution, concurrency, and reliability behavior. It does not continuously monitor production, upload evidence, provide accounts or synchronization, modify source code, install system tools, or delete artifacts automatically.
flameox also does not reimplement profilers, trace databases, native viewers, or private format decoders. It reports a workload as contained only when an active backend enforces that boundary, and it never treats statistical non-significance as proof of equivalence.
Setup and installation details
Run setup again to connect or disconnect clients, verify the active runtime,
update to the npm package's matching version, or roll back to a previously
installed version. npx flameox@latest setup resolves the newest setup release.
Automation can select clients and inspect the plan explicitly:
npx flameox setup --codex --claude --yes
npx flameox setup --all --dry-run --json
npx flameox setup --verify --yes --json
The npm package is only a bootstrap. It delegates to the exactly matching
flameox Python release and supplies the maintained jsonc-parser
helper needed to preserve comments in OpenCode configuration. MCP clients launch
the installed runtime directly; they do not invoke npx, uvx, or a
network-dependent installer at startup. Setup itself does not initialize a
project or create .diagnostics/.
Optional Python extras are independent:
python: pyperf capture and importcpu: py-spy capturetrace: Perfetto Python API; a local Trace Processor binary must also be configuredexecution: coverage.pymemory: Memraytorch: PyTorch captureall: all runtime integrations
flameox pins the official Python MCP SDK to mcp==2.0.0b2.
Local data model
Initialize a project-local workspace:
uv run flameox init .
uv run flameox status
.diagnostics/ contains content-addressed native artifacts, immutable JSON
records, append-only Parquet generations, and a rebuildable DuckDB catalog.
Parquet and generation manifests are authoritative; deleting
catalog.duckdb is safe because flameox catalog rebuild reconstructs it.
Artifacts are deduplicated by SHA-256 without collapsing their contextual registrations. A run records source and environment identity, workload and measurement identities, lifecycle state, validation, process evidence, and artifact roles. Investigations, hypotheses, experiments, variants, trials, frozen run sets, comparisons, and findings remain separate domain records.
Named workloads and capture
Repeatable commands live in flameox.toml. Templates accept declared scalar
parameters only—there is no shell expansion:
schema_version = 1
[workloads.scan]
argv = ["python", "bench.py", "--implementation", "{implementation}"]
cwd = "."
timeout_seconds = 60
[workloads.scan.parameters]
implementation = ["baseline", "candidate"]
[workloads.scan.oracle]
strength = "cross_treatment_equivalence"
argv = ["python", "validate.py", "--implementation", "{implementation}"]
[experiments.scan_comparison]
workload = "scan"
variants = ["baseline", "candidate"]
design = "randomized_complete_blocks"
blocks = 10
primary_metric = "pyperf.workload"
polarity = "lower_is_better"
estimand = "median_paired_log_ratio"
practical_threshold = 0.05
confidence_level = 0.95
random_seed = 1984
Inspect and approve the exact canonical workload before exposing it through MCP:
uv run flameox workload show scan --json
uv run flameox workload approve scan
uv run flameox capture plan pyperf --workload scan \
--parameters '{"implementation":"baseline"}' --json
uv run flameox capture run pyperf --workload scan \
--parameters '{"implementation":"baseline"}' --json
Editing a command, environment, parameter domain, timeout, working directory,
or oracle changes the canonical digest and revokes that approval. Execution
uses argument arrays through one subprocess broker, bounded output, timeout and
cancellation cleanup, and optional Linux bubblewrap containment. Perfetto
parsing also runs in a broker-owned worker so a long trace cannot block or
outlive the MCP request. A truthful uncontained result is never relabeled as
sandboxed.
Investigations and experiments
Create an investigation and optionally attach a falsifiable hypothesis before running a predeclared experiment:
uv run flameox investigations create \
'{"question":"Does the candidate remove reverse-scan overhead?"}' --json
uv run flameox hypotheses record @hypothesis.json --json
uv run flameox experiment plan scan_comparison \
--investigation <investigation-id> --adapter pyperf --json
uv run flameox experiment run scan_comparison \
--investigation <investigation-id> --adapter pyperf --json
Experiment execution randomizes treatment order within complete blocks, persists the declared protocol before collection, registers every attempted trial—including cancellation and failure—freezes one trial-aware run set per variant, and only runs the automatic paired comparison when the blocks, measurements, source identity, environment, and cross-treatment validation support it. Failed trials stay in the evidence rather than disappearing from the denominator.
Useful read-only analyses include:
uv run flameox analyze hotspots <run-or-artifact>
uv run flameox analyze scaling <experiment-id>
uv run flameox analyze compare @comparison-request.json
uv run flameox analyze memory <run-or-artifact>
uv run flameox analyze execution <run-or-artifact>
uv run flameox analyze pytorch <run-or-artifact>
uv run flameox analyze failures
Those commands are deterministic read-only previews. Persist a recipe result and its typed provenance explicitly:
uv run flameox analyze record \
'{"recipe":"memory","input_id":"<run-id>"}'
uv run flameox analyze record-comparison @comparison-request.json
Hotspots can be followed into normalized trace structure without arbitrary SQL:
uv run flameox stacks callers <run-or-artifact> <frame-id> [--cursor CURSOR]
uv run flameox stacks callees <run-or-artifact> <frame-id>
uv run flameox stacks examples <run-or-artifact> <frame-id>
uv run flameox trace window <artifact-id> --start 0 --end 1000000 [--cursor CURSOR]
uv run flameox open <artifact-id>
flameox open only prints a native viewer plan. --launch is an explicit
consequential action and cannot be combined with --json.
CLI and MCP
Start the permanently local stdio server with a fixed project root:
uv run flameox mcp serve --project-root .
The MCP layer is a thin adapter over the same application services as the CLI.
It offers approved named capture and experiment plans, bounded evidence and
drill-down queries, pure analysis previews, explicit record_analysis and
record_comparison operations, typed records, resources, progress, structured
domain errors, and cancellation cleanup. It does not expose shell strings,
arbitrary SQL, raw artifact bytes, approval mutation, deletion, or viewer
launching.
Plan tokens are 256-bit, in-memory, short-lived, bound to the current workspace, approval, executable, policy, adapter, parameters, and experiment definition, and atomically single-use. Restarting the server invalidates them.
Inspect the protocol surface with a real stdio client:
uv run flameox mcp inspect --project-root . --json
Integrity and recovery
uv run flameox validate
uv run flameox validate --full
uv run flameox catalog validate
uv run flameox catalog rebuild
uv run flameox catalog compact
uv run flameox recover
uv run flameox gc
uv run flameox gc --apply
Full validation hashes native artifacts and Parquet files. Recovery closes only
runs whose exact boot/PID/process-start lease has disappeared. Garbage
collection is dry-run by default and --apply moves eligible objects into
recoverable trash instead of unlinking them.
Documentation
- Architecture: process model, package boundaries, dependencies, and platform policy
- Storage and evidence: authoritative data, identity, provenance, publication, and schemas
- Investigations and analysis: workloads, experiments, recipes, statistics, and evidence quality
- Adapters and capabilities: profiler integration, compatibility, probing, and approval behavior
- Runtime safety: concurrency, recovery, retention, integrity, security, privacy, and local observability
- CLI and MCP boundaries: human and agent interfaces and their trust boundaries
- Architectural decisions: settled choices and open design questions
- Acceptance and verification: completion criteria and representative proof
Development
uv sync --extra dev --extra python --extra execution --extra memory --extra trace --extra cpu
uv run ruff check src tests
uv run mypy src tests
uv run pytest -q
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file flameox-0.1.1.tar.gz.
File metadata
- Download URL: flameox-0.1.1.tar.gz
- Upload date:
- Size: 2.1 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.7 {"installer":{"name":"uv","version":"0.11.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2c109ebd9ed33fc5d8cee404c537da0076da5b543065ce38e1d4054db68ed72c
|
|
| MD5 |
026663172288c65e0b2f0d033316fa41
|
|
| BLAKE2b-256 |
7cc765cbf1aad880f8df2e2e1596213403e741c8e435a3e9ba487de921f5370a
|
File details
Details for the file flameox-0.1.1-py3-none-any.whl.
File metadata
- Download URL: flameox-0.1.1-py3-none-any.whl
- Upload date:
- Size: 181.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.7 {"installer":{"name":"uv","version":"0.11.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e0302a85a7260dc3585ef0dd9ec5dc80481644cea928c3fe7c4f8ec89911d5c1
|
|
| MD5 |
e18e5760903bbf613495fed8b7975755
|
|
| BLAKE2b-256 |
6fa543adf7767fb7b95bacf5eef188a04079c6dfedbc355a4a1b102249fe8805
|