Discriminative MCTS Agent
An autonomous reasoning and execution agent powered by Discriminative Monte Carlo Tree Search (MCTS). The system couples TypeSafe Jev System One Primitives (Noul, Choice, Score) for lightning-fast, structured discriminative evaluations with Gemini 3.8 Flash for action generation in a closed-loop execution environment.
Includes a local MCTS Agent Web UI for prompting the agent, following its output and image artifacts live, and loading saved run logs.
Key Highlights
- Fast Discriminative Scoring: Choice assigns action priors and Score evaluates hypothetical states. Low early scores do not remove proposed paths.
- Low Tree Reuse Threshold: After execution, grounded surviving paths are rescored. The tree is rebuilt only when no actionable path reaches the reuse threshold.
- Dynamic Prime Branching (
Choice): Dynamically samples expansion widths from prime numbers ($2, 3, 5, 7, 11, 13$) based on state uncertainty and goal complexity. - Closed-Loop Execution & Grounding: Follows a strict cycle of Plan & Choose -> Execute -> Review -> Adapt -> Assess. Only the immediate winning action is executed; git diffs and terminal observations are folded back into context for subsequent decisions.
- Agent Web UI: Prompt the agent, follow its live text output, preview changed image files, and browse or import run logs.
Architecture Overview
┌──────────────────────────┐
│ Goal & Initial State │
└────────────┬─────────────┘
│
┌──────────────────────▼──────────────────────┐
│ 1. PLAN & CHOOSE (Discriminative MCTS) │
│ ├── Selection: PUCT + First-Play Urgency │
│ ├── Expansion: Gemini proposes actions │
│ ├── Keep all proposed actions │
│ ├── Priors: Choice policy distribution │
│ ├── Value: Score rubric evaluation │
│ └── Backpropagation: Value sum & visits │
└──────────────────────┬──────────────────────┘
│ Best Immediate Action
┌──────────────────────▼──────────────────────┐
│ 2. EXECUTE │
│ Execute ONLY the selected immediate step │
│ in the designated workspace. │
└──────────────────────┬──────────────────────┘
│ Returncode, stdout, stderr
┌──────────────────────▼──────────────────────┐
│ 3. REVIEW │
│ Inspect environment changes, git status, │
│ new files, and terminal outcomes. │
└──────────────────────┬──────────────────────┘
│ Observation
┌──────────────────────▼──────────────────────┐
│ 4. ADAPT │
│ Ground real-world execution observation │
│ into updated state description. │
└──────────────────────┬──────────────────────┘
│ Grounded State
┌──────────────────────▼──────────────────────┐
│ 5. ASSESS (Early Stopping) │
│ Noul evaluates: Is the goal complete? │
│ - Yes (conf >= 0.85) -> Finish! │
│ - No (conf < 0.85) -> Next MCTS step │
└─────────────────────────────────────────────┘
TypeSafe Jev System One Primitives
Large language models (LLMs) are System Two text-generators. When software needs fast, deterministic judgments, coercing an LLM into producing structured JSON and parsing strings introduces latency, format failures, and context rot.
TypeSafe Jev is a System One model built for instant, structured judgments consumed directly in software:
| Primitive | Role in MCTS | Mechanism & Benefit |
|---|---|---|
Noul |
Execution Checks & Early Stopping | Evaluates whether a proposed command matches the selected action, whether execution evidence supports completion, and whether the overall goal is complete. It does not prune proposed search paths. |
Choice |
Action Priors & Dynamic Branching | Returns normalized probability distributions across discrete candidates. Assigns initial policy prior probabilities $P(s, a)$ used in the PUCT formula. Also selects optimal branching factor from primes ($2, 3, 5, 7, 11, 13$). |
Score |
State Value Heuristic | Evaluates states on a continuous 1–10 rubric. Replaces costly, random Monte Carlo rollouts with a fast discriminative heuristic value estimate $V(s)$. |
PUCT Formula with First-Play Urgency (FPU)
Nodes are selected during tree traversal using the predictor upper confidence bound applied to trees (PUCT):
$$\text{PUCT}(s, a) = Q(s, a) + c_{\text{puct}} \cdot P(s, a) \cdot \frac{\sqrt{N(s)}}{1 + N(s, a)}$$
- $Q(s, a)$ is normalized to $[0, 1]$ (divided by max score $10.0$) to maintain scale balance with the exploration term.
- First-Play Urgency (FPU): Unvisited child nodes inherit their parent's estimated value rather than defaulting to 0, preventing starvation of unvisited siblings when one child discovers a high score early.
Agent Web UI
The local web UI is served by mcts-agent visualize.
Start and use the UI
-
Via the CLI tool:
mcts-agent visualize --port 8000
This serves the project workspace and opens the app in your browser.
-
Prompt and run:
- Enter a goal and optional context, then choose the search depth and branching factor.
- Mock mode is enabled by default. Workspace execution is optional.
- Follow the agent's console output and MCTS events live. Changed image files appear in the artifact gallery.
-
Review prior runs:
- Open Run history to browse saved JSON logs in
logs/. - Import an event log from disk with Load log.
- Open Run history to browse saved JSON logs in
Quickstart
Prerequisites
- Python $\ge$ 3.10
- Git
- TypeSafe API Key
- Antigravity CLI (
agy) or Gemini API key
Installation
Once the package is published, install it as a command-line tool with uv or
pipx:
uv tool install mcts-agent
# or
pipx install mcts-agent
For development from source:
# 1. Clone the repository
git clone https://github.com/lhemerly/mcts-agent.git
cd mcts-agent
# 2. Set up a virtual environment
python3 -m venv .venv
source .venv/bin/activate
# 3. Install in editable mode with CLI entry point
pip install -e .
Check for a newer PyPI release with mcts-agent version, or explicitly upgrade
with mcts-agent update. Normal CLI commands check PyPI at most once every 24
hours and only display a notice; they never update automatically. Use
mcts-agent --no-update-check ... or set
MCTS_AGENT_DISABLE_UPDATE_CHECK=1 to disable that check.
Environment Configuration
Create a .env file in the project root:
TYPESAFE_API_KEY=your_typesafe_api_key_here
# Optional model overrides:
AGY_MODEL=gemini-3.8-flash-medium
Usage
Research mode: open-ended investigations
Research mode adds a structured notebook of scope, assumptions, validation criteria, claims, and captured evidence. It uses the existing harnesses and MCTS engine to choose investigations, then evaluates actual artifacts before proposing a scoped candidate answer. System One judgments remain explicitly heuristic.
mcts-agent research --query "Investigate an open question" --workspace . --mock --max-steps 2 --json
# Resume with the checkpoint path returned by the command:
mcts-agent research --resume .mcts-research/RUN_ID/checkpoint.json --max-steps 3 --json
Remove --mock for live work with configured providers. This mode is available
through the CLI and Python API; see Research mode for
validation adapters, evidence handling, budgets, and interruption behavior.
The CLI provides a modern, Rich-stylized terminal experience with colored panels, search rollout progress, and step tables, along with full headless / machine-parseable support for AI agents.
1. Run Closed-Loop MCTS
Execute search and execution with a custom goal and workspace:
mcts-agent run --goal "Build a REST API" --workspace ./my-project --max-steps 5
2. Live Demo Run
Run against the built-in demo goal:
mcts-agent demo
3. Hermetic Mock Mode (Zero Keys Required)
Run the full closed-loop search loop with mock primitives (no network calls or API keys needed):
mcts-agent demo --mock
# or
mcts-agent run --mock --goal "Create a hello world app"
4. Interactive Human-in-the-Loop Mode
Interactively inspect MCTS action recommendations and approve, skip, or manually redirect states:
mcts-agent interactive
5. Launch the Agent Web UI
Open the local app to prompt the agent, follow live output, preview image artifacts, and load saved logs:
mcts-agent visualize --port 8000
6. Backward Compatibility
You can continue invoking the CLI via main.py directly:
python main.py --mock --max-steps 3
AI-Friendly & Headless Features
For automated pipelines, orchestrators, and AI agents, mcts-agent provides clean stdout modes:
Machine-Parseable JSON Output (--json)
Emits a clean, valid JSON summary to stdout with zero ANSI escape codes, progress banners, or decorative styling:
mcts-agent run --mock --goal "Refactor user authentication" --json > result.json
Or pipe directly into tools like jq:
mcts-agent run --mock --json | jq '.steps[] | {step, action, score, visits}'
Example JSON response:
{
"goal": "Refactor user authentication",
"completed": true,
"completion_confidence": 0.92,
"total_steps_run": 2,
"steps": [
{
"step": 1,
"action": "Analyze existing authentication handlers in auth/login.py",
"score": 8.75,
"visits": 10,
"mcts_log": "logs/mcts_20260917_214342_step1.json",
"execution": {
"success": true,
"stdout": "...",
"returncode": 0
}
}
],
"summary_file": "logs/run_20260917_214342_summary.json"
}
Quiet and Headless Flags
--quiet/-q: Suppresses informational banners and progress spinners while retaining essential results.--no-color: Disables ANSI coloring and setsNO_COLOR=1for clean plaintext logging.
CLI Options Reference
Commands
mcts-agent run [OPTIONS]: Run closed-loop MCTS search with custom goal, state, and execution options.mcts-agent demo [OPTIONS]: Run pre-configured task management demo.mcts-agent interactive [OPTIONS]: Step-by-step human-in-the-loop search with manual approval.mcts-agent visualize [OPTIONS]: Launch the local agent web UI and run history.
Common Options (run & demo)
| Option | Short | Default | Description |
|---|---|---|---|
--goal |
-g |
Demo goal | Custom task goal string |
--state |
-s |
Demo state | Initial state / context description |
--workspace |
-w |
. |
Target directory for workspace actions |
--max-steps |
5 |
Maximum outer loop execution steps | |
--iterations |
-n |
10 |
MCTS search iterations per reasoning step |
--sim-depth |
Dynamic | Fixed lookahead simulation depth (primes 2, 3, 5) | |
--actions |
-a |
Dynamic | Fixed candidate actions per node (primes $\le 13$) |
--mock |
False |
Run hermetically with mock primitives (no API keys required) | |
--no-early-stop |
False |
Disable Noul-based completion early stopping | |
--no-execute |
False |
Plan only; skip executing actions in the workspace | |
--json |
False |
Output clean machine-parseable JSON summary to stdout | |
--quiet |
-q |
False |
Suppress informational messages and progress banners |
--no-color |
False |
Disable ANSI color output |
visualize Options
| Option | Short | Default | Description |
|---|---|---|---|
--port |
-p |
8000 |
Port for the local agent web UI |
--host |
127.0.0.1 |
Host interface to bind server to | |
--browser / --no-browser |
True |
Automatically open default web browser |
Project Structure
mcts-agent/
├── agent/
│ ├── __init__.py
│ ├── logger.py # Structured JSON event emission for live output and logs
│ ├── mcts.py # Selection, Expansion, Simulation, Backpropagation loop
│ ├── node.py # Search tree Node data structure & PUCT calculations
│ └── primitives.py # TypeSafe Jev primitives (Noul, Choice, Score) wrappers
├── logs/ # Search event logs and run summaries
├── prompts/
│ └── rubrics.txt # Evaluation rubrics and reference criteria
├── tests/
│ ├── __init__.py
│ └── test_mcts.py # Unit tests for primitives, PUCT, and search loop
├── .github/
│ └── workflows/
│ └── ci.yml # GitHub Actions CI matrix testing
├── .gitignore # Clean git exclusion rules
├── CONTRIBUTING.md # Contribution guide & development setup
├── LICENSE # MIT License
├── pyproject.toml # PEP 517/621 packaging metadata
├── requirements.txt # Python dependencies
├── main.py # Main CLI entrypoint for closed-loop execution
├── run_pure_mcts.py # Standalone script for deep pure MCTS exploration
├── run_test.py # Integration benchmark runner
└── agent/visualizer.html # Local prompt, output, artifacts, and run history UI
Testing
Run unit tests across all components using Python's standard unittest or pytest:
# Using unittest
python3 -m unittest discover tests -v
# Or using pytest
pytest -v
All test cases execute hermetically in mock mode and verify:
- Node initialization, PUCT computation, and First-Play Urgency
- Low-score path preservation and tree reuse
- Prior probability distribution generation (
Choice) - Continuous state scoring (
Score) - MCTSLogger event emission and JSON serialization
- Multi-step closed-loop search execution
License
This project is licensed under the MIT License.
Release files for mcts-agent 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mcts_agent-0.1.0.tar.gz | 113.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mcts_agent-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 193.7 kB
Release files / mcts_agent-0.1.0.tar.gz
| Download URL | mcts_agent-0.1.0.tar.gz |
|---|---|
| Size | 113.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
37adb88a79831cd33669df58485bf93d4c4044e0fb6d59cf96e115d14980a99b
|
|
BLAKE2b-256 checksum How to use checksums |
e5cd4fa55dd325829435b75640d268040036ed425ee6408631c10c98d2020a4b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency logRelease files / mcts_agent-0.1.0-py3-none-any.whl
| Download URL | mcts_agent-0.1.0-py3-none-any.whl |
|---|---|
| Size | 80.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
32ade8c417337bcfd5174c34116494847e8e63726f69f26abad7e4efe6c72ca6
|
|
BLAKE2b-256 checksum How to use checksums |
d1921752feef039ed1b40696c5d8a9f05315ef4156fea008dce037a45a583d6e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency log