Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

MLCA: Machine Learning Copilot Agent (v2)

MLCA guides a machine-learning study step by step. Instead of an open-ended chat, a study moves through a fixed sequence of events. At each step you choose from a few natural options, or give a custom instruction limited to the current event. Every step becomes a standalone, numbered script with recorded outputs. Rules that protect the science are enforced by code, not by prompts: the test set stays locked, models are fitted on training rows only, and every number has a source.

onboard → data → eda → split 🔒 → preprocess → features → model ⇄ validate → test (unlock once) → interpret → report
                                     └──────── iteration loop (validation only) ────────┘

How a step works

  1. Status: the engine rebuilds the project state from its event log.
  2. Options: two to three options from the current event's action catalog, plus custom. When an event is done, the next event is offered.
  3. Choice: you choose (or an agent chooses; see Two ways to use MLCA).
  4. Script: a short, standalone Python script is written for the action and run in a sandbox:
    • only declared inputs are readable;
    • no network;
    • the test set is invisible until unlocked.
  5. Record: outputs, metrics and facts (each citing the file it came from) go into the event log. The event's sub-goals are then checked by code.

Each event has a goal, sub-goals, an action catalog and a custom-instruction scope. These are defined in a workflow template:

  • templates/tabular_classification/: logistic baseline vs gradient boosting, ROC AUC;
  • templates/survival/: Cox vs penalised Cox, C-index.

The template content is short by design; see templates.

Two ways to use MLCA

Standalone (available now) Docked to another agent through MCP (available now)
Interface terminal UI (mlca), CLI, browser view of the TUI (mlca web) Claude Code, Codex, Hermes, OpenClaw, Cursor or any MCP client, connected with mlca mcp
Who proposes options and writes scripts MLCA's own Navigator and Worker models (your API keys: OpenAI, OpenRouter or a local OpenAI-compatible server) Brain mode (default): the connected agent proposes options and writes each script; MLCA makes no model calls. Driver mode: the agent only chooses; MLCA's Navigator and Worker do the work
Rules all engine rules the same engine rules; the agent cannot confirm sub-goals, unlock the test set or approve packages; it can only request them, and you approve in MLCA
Attribution decisions recorded as user decisions recorded as agent:<client>/<version>, so reports separate human and agent work

In brain mode, a script written by Claude or Codex goes through exactly the same checks as one written by MLCA's Worker:

  • option validation;
  • static checks;
  • the sandbox;
  • the training-only fit rule;
  • the test lock.

Connect an agent

mlca mcp install --client claude --with-skill       # preview
mlca mcp install --client claude --with-skill --yes # register only with explicit consent
mlca mcp install --client codex                    # preview
mlca mcp doctor                                   # real stdio tool discovery

The server is included in the Python package. See MCP docking for Hermes, OpenClaw, Cursor, work orders, polling, privacy and approval commands. Human-only enforcement is a Session/MCP boundary; MCP client identity is attribution, not OS authentication.

Install

The distribution name is mlcagent; the command and Python imports remain mlca.

Python ≥ 3.11 and uv are required. Install the beta with an explicit version:

uv tool install mlcagent==0.2.0b1

Plain uv tool install mlcagent does not select this prerelease. Shell download, npm and Homebrew channels remain unpublished.

From this checkout: uv sync --group dev, then uv run mlca. See installation.

Isolation:

  • Linux: full isolation (level A) with bubblewrap: sudo apt install bubblewrap.
  • macOS and native Windows: only the weaker level B, which must be enabled explicitly (sandbox.allow_level_b: true).
  • Windows: WSL2 is recommended.

Configure models (standalone mode)

Copy examples/config.yaml to ~/.config/mlca/config.yaml (or run mlca config edit). Choose a provider and model for each role. Keys are read from environment variables (OPENAI_API_KEY, OPENROUTER_API_KEY) and never written to disk. No default models or prices are assumed. See providers.

Quick start

mlca init ./study --data data.csv --template tabular_classification   # asks the brief; resolves the target column
mlca -p ./study                                                        # terminal UI
# or step by step from the CLI:
mlca -p ./study options
mlca -p ./study choose 1
mlca -p ./study confirm brief_confirmed "Aim, target and unit checked"
mlca -p ./study next

Human decisions are explicit:

  • confirm <sub-goal> "<note>" for author checks;
  • unlock-test --phrase "UNLOCK TEST" --reason "..." once, in the test event;
  • package approvals.

After the test set is unlocked, fitting, validation loops and package changes are closed for the project.

Supporting inspections follow the main actions in fallback menus. Once an event is complete, the menu offers next first, then unrun supporting actions and scoped custom instructions. Those steps preserve completed checks. To try a primary alternative, branch before its step.

Trust boundary

MLCA enforces decisions and data access only through its own API and sandboxed steps. A connected agent's native file or shell tools can bypass MCP: they may read the original data, .mlca/data/splits/, the test key under the MLCA user-data directory, or invoke human approval commands. Generated tables/, figures/, reports/ and scripts/ are also local copies of evidence. Reports therefore describe agent_file_access as unknown (outside MLCA).

mlca -p ./study mcp install --client claude --scope project previews recommended permissions; --yes backs up and merges them into ./study/.claude/settings.json. The block denies native shell execution and reads/edits of registered data, private state, keys and evidence copies. Other clients receive equivalent guidance. These are client restrictions, not OS authentication: disable other filesystem tools and bypass modes, or use a separate OS identity/container. See MCP trust boundary before connecting a client.

With allow_sample_rows: false, private salted fingerprints check values across all columns, regardless of their names, at six and four significant digits. The additional table heuristic flags at least min(3, n_columns) distinct cell matches to one source row; exact shared-column checks also remain. Full-row facts are checked without requiring matching column names. Matching outputs are flagged in reports; genuine summaries normally remain readable, but coincidental matches may be refused. Existing indexes upgrade from unchanged registered data without decrypting the test set.

This guard catches accidental and simple copies, not a deliberate agent. In brain mode the agent writes the code and can encode data in strings, noise, figures or other derived outputs. Those results are sent to the agent's provider by design. For confidential data, use standalone driver mode with every model local and no external agent receiving outputs; the heuristic cannot guarantee confidentiality for brain mode.

Scientific guard rails

Rule Where it is enforced
Options are only actions from the current event's catalog, with validated parameters engine/options.py
An event can be entered only when its gate passes; sub-goals are ticked only by code checks or your confirmation engine/gates.py, engine/checks.py
Custom instructions must stay in the event's scope engine/scope_guard.py
The split is done by the engine; the test partition is encrypted until unlocked engine/split.py, sandbox/vault.py
Scripts run sandboxed: declared inputs only, no network, writes only to their own step folder runner/, sandbox/
In preprocess, features and model, every fit goes through mlca_runtime.fit_record on the exact frozen training rows; other events may not fit at all runner/static_checks.py, mlca_runtime
Every fact cites the file it came from engine/facts.py

Limits: level B (macOS, Windows) is an audit hook, not an operating-system boundary. Code checks cannot prove the meaning of every computation. Domain review of targets, units, preprocessing and interpretation remains yours. See security model.

Evaluation

mlca metrics writes sourced completion, recovery, intervention, time, usage and provenance measurements without model calls. After a brain-mode run, mlca usage import imports only client-reported usage fields; unknown costs stay unavailable. mlca bench run creates independent driver runs from a task card and keeps failures; scripted policy decisions are labelled separately from humans. See metrics.

Evidence and reproducibility

Everything lives in the project's .mlca/ folder:

  • the append-only, hash-chained event log;
  • the frozen template;
  • the analysis environment and its lock file;
  • the immutable step folders;
  • the encrypted test partition;
  • the usage and egress logs.

The scripts/, figures/, tables/, reports/ and PROJECT_LOG.md folders are regenerated views; edits to them are detected.

mlca log --verify        # integrity of the event log
mlca replay --verify     # re-run every step script in a fresh environment and compare outputs
mlca report all          # methods draft, provenance, figure index, usage (no model calls)
mlca branch alt --from <record or step>   # try another path; mlca checkout main to return

Status

  • PyPI beta: upload rejected with HTTP 400 (project-name similarity); nothing published. Release source: ac4bcc8. A replacement name or registry resolution is required.

  • Core (plan.md): P0–P8 built; beta version 0.2.0b1. 106 offline tests pass, 1 skipped (PowerShell unavailable); see validation.

  • Not yet done:

    • runs with real providers;
    • the study on real data;
    • native macOS and Windows checks.
  • MCP docking (plan_mcp.md): M0–M6 implemented: 14 stdio tools, brain/driver modes, human approvals, filtered evidence and usage records. Manual runs in real clients (M7) remain pending.

  • Licence: CC BY-NC-ND 4.0 (CC-BY-NC-ND-4.0), chosen by the author. The approved package name is mlca.

Documentation

Document Contents
plan.md Core design, phases and current status
docs/mcp.md Client setup, tool loop, approvals and manual validation
plan_mcp.md MCP server: docking modes, tools, safety, phases
docs/architecture.md Components and data flow
docs/templates.md Writing workflow templates (short-content rules)
docs/event-log.md Record types and replay
docs/security-model.md Isolation levels and their limits
docs/providers.md Model providers and roles
docs/install.md, docs/releasing.md Install channels; author-only release setup
docs/m7-vm-brain-mode.md Manual end-to-end test on a fresh VM with Claude Code in brain mode
docs/testing.md Running the offline tests (prime the uv cache once: uv run python scripts/prime_test_cache.py)
docs/DECISIONS.md, CHANGELOG.md Decisions and history

Metadata

Release files for mlcagent 0.2.0b1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mlcagent 0.2.0b1
File Size Uploaded
mlcagent-0.2.0b1.tar.gz 784.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mlcagent 0.2.0b1
File Interpreter ABI Platform
mlcagent-0.2.0b1-py3-none-any.whl Python 3 none any Details

Total release size: 1.1 MB

Release files / mlcagent-0.2.0b1.tar.gz

Download URL mlcagent-0.2.0b1.tar.gz
Size 784.2 kB
Tags Source
SHA-256 checksum
How to use checksums
3f91399736cb8a80a6a0208671928b39f6166726a2cc34feee632acd284b336f
BLAKE2b-256 checksum
How to use checksums
1b89abfc54761ca17ab7918300ebe03f9e105206e9999bf34c641f953b862e62
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / mlcagent-0.2.0b1-py3-none-any.whl

Download URL mlcagent-0.2.0b1-py3-none-any.whl
Size 272.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a437b0705183a5655e350fb5a5bd6d125d0474cc0502e45d1be09ede53e727c5
BLAKE2b-256 checksum
How to use checksums
1d22d1abb8553c6878b6447ed3e38f32d0fe8160be07d3605557ba1a522a2dc5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.2.0b1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page