Skip to main content

From papers to proof — a Claude-first research and experimentation CLI.

Project description

ResearchForge

From papers to proof.

ResearchForge turns a research question — or an "improve my repository" goal — into evidence. It finds relevant papers, helps Claude form testable hypotheses, benchmarks competing implementations against a frozen baseline in isolated local workspaces, and delivers the strongest supported result as a clean branch, an engineering report, or a research package.

Inspired by Andrej Karpathy's autoresearch, where an agent autonomously runs training experiments against a fixed benchmark overnight — ResearchForge generalizes that loop to any repository with a measurable benchmark, and grounds it in literature: papers → hypotheses → controlled experiments → a validated, reproducible result.

How it works

Three parties with strict roles:

Who Does what
You set the objective; approve the benchmark contract, each experiment plan, and anything that ships
Claude (in Claude Code) reads the papers, writes the research landscape, hypotheses, and experiment patches; explains results
The Python engine everything that must be trustworthy: search, validation, isolated execution, ranking, shipping — no prompt can bypass it

No model API key is required — Claude Code is the model. Every artifact Claude writes is schema-validated before it is stored, every experiment runs in a detached git worktree (your checkout is never touched), and "validated" is only ever earned by repeated benchmark runs.

Get started (two minutes)

Everything runs on your machine — nothing is uploaded anywhere.

# 1. Install ResearchForge (Python 3.12+):
pip install "researchforge[serve]"       # or: pipx install "researchforge[serve]"

# 2. Make ResearchForge available in EVERY Claude Code session (recommended):
researchforge claude install --user      # skills -> ~/.claude/skills/

That's the whole setup. Now open any Claude Code session and say what you want:

/researchforge-start"Can uncertainty-aware routing outperform fixed routing?" — or — "Improve this repo's F1 without hurting latency — work in ~/projects/my-repo."

Claude asks where the project should live (any folder — a cloned repo to improve, or an empty directory for pure research), initializes it there, and walks the whole journey: it runs the CLI, shows you what it found, and asks before anything is approved, executed, or shipped. If you ever wonder where things stand, researchforge status names the exact next step — and the hub at http://127.0.0.1:9000 shows every project, where it lives, and what's happening, all the time.

Prefer per-project skills instead of the global install? cd into the folder and run researchforge init --claude, then start Claude Code in that folder (desktop app: open the folder as the session's project — a home-screen session without a folder can't see project skills).

Everything is directory-scoped: the database, worktrees, artifacts, and dashboard live under the project folder. Run commands from any subfolder (they walk up to the project root like git), point at another project with researchforge -C /path/to/project <command>, and print the full location map with researchforge paths.

Working on the ResearchForge source itself?

git clone https://github.com/forger-labs-hq/researchforge
cd researchforge
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"

See it work first

docs/demo.md is a ten-minute, fully offline walkthrough against examples/simple-python — a deterministic benchmark where one variant genuinely improves F1, one is rejected for violating the latency budget, and one fails (all three stay in the record). examples/docker-python is the same demo under Docker isolation.

Watch your experiments (dashboard + live monitor)

Two ways to see how experiments perform against the baseline:

researchforge dashboard --open       # one self-contained HTML snapshot

A single static file (no server, no JS libraries, nothing leaves your machine) with an autoresearch-style progress chart — every experiment as a chronological dot, kept improvements annotated, a running-best step line — plus per-experiment bars vs the baseline, the trade-off scatter with the hard-constraint line, the funnel with drop-offs, and validation spread.

The static dashboard also opens with summary stats (best score and its delta vs baseline, experiments kept · discarded · errored, the Pareto frontier) and the experiment tree — a graph from the baseline through every branch of experiments (see "experiments on experiments" below).

pip install "researchforge[serve]"
researchforge serve --background     # live monitor for THIS project (URL printed)

A local web monitor that follows runs as they happen: overview with the next action, collapsible research sessions with the full recorded detail (directions, evidence, limitations, underexplored aspects), per-run execution timelines with work locations on disk, the experiment tree with click-through drill-down pages for every experiment (lineage, decision, executions, artifacts on disk), the live chart dashboard, and a JSON API (/api/state). Once the extra is installed, experiment run/start auto-start it and print the URL. Manage it with researchforge serve --status / --stop. The server opens the database read-only and binds 127.0.0.1 only by default — watching can never interfere with a run.

The hub — every project, one dashboard

researchforge hub --background       # http://127.0.0.1:9000 — all projects

Projects live in whatever folders you initialize them in, and it is easy to forget which. The hub is one machine-wide page listing every project with its folder location, status, and live activity; click through to any project's full monitor (sessions, runs, experiment tree, drill-downs). Every researchforge init registers the project, so new projects appear automatically — and once the serve extra is installed, any researchforge command quietly ensures the hub is running, so http://127.0.0.1:9000 is simply always there (set RESEARCHFORGE_NO_HUB=1 to opt out). Commands run in a subfolder of a project also walk up to find it (like git) and print Using project at <root> so you always know which project you're in.

Journey A — research an idea

You have a question; ResearchForge grounds it in literature. From Claude Code, /researchforge-start with your question does all of this; the CLI equivalent (where you write the synthesis files against exported schemas) is:

researchforge project create --mode explore_research_idea --objective "..."
researchforge research search        # arXiv: fetch -> dedup -> rank -> store
researchforge research context       # exports context.json with the schemas
# Claude (or you) writes landscape.yaml + hypotheses.yaml — validated on import:
researchforge research landscape --import .researchforge/synthesis/landscape.yaml
researchforge hypotheses import .researchforge/synthesis/hypotheses.yaml
researchforge report build           # citation-backed report
researchforge paper package          # optional: BibTeX, outline, evidence matrix

You end up with: a research landscape (directions + graded evidence: published claim vs interpretation vs speculation), testable hypotheses, a citation-backed Markdown report, and optionally a full research bundle. Details: docs/research-mode.md

Journey B — improve a repository

Everything in Journey A, then benchmarked experiments on your code. Every consequential step is your typed approval — Claude cannot approve, run, or ship anything by itself.

1. Define and freeze the evaluation:

researchforge project create --mode improve_repository --objective "..."
researchforge repo scan
researchforge contract generate      # edit researchforge.yaml, then:
researchforge contract approve       # YOUR approval -> immutable contract
researchforge baseline run           # frozen reference measurement

2. Run experiments (Claude writes plan.yaml + one patch per variant after researchforge experiment plan hyp-001; manually you write them against the exported schema):

researchforge experiment start .researchforge/experiments/plan.yaml
# = import + ONE typed approval (shows worst-case wall time) + run

Experiments on experiments (branching): a plan entry can declare parent: exp-001 (or the key of another entry in the same plan) to build on top of a previous experiment instead of the baseline — refine a winner, or rescue a rejected idea by combining it with something else. The engine validates the whole chain at import (a child that doesn't apply on its parent's state is refused), executes it root-first in isolation, and still measures honestly against the frozen baseline. Ship a branched winner and the chain is composed into one clean commit. The dashboard draws the whole tree, like Karpathy's autoresearch UI.

3. Inspect and validate:

researchforge results show run-001   # ranking, trade-offs, rejections
researchforge validate run-001       # repeated runs earn "validated"

4. Accept the result → what ships (and how the PR happens):

researchforge ship branch    # clean LOCAL branch on the frozen baseline —
                             # one commit, nothing pushed, inspect with git
researchforge report build   # engineering report: the full evidence chain
researchforge ship pr        # OPT-IN: push + open a DRAFT PR on YOUR repo

ship pr is how the pull request gets created, and it only works when all three gates open: shipping.allow_draft_pr: true in the approved contract, the gh CLI authenticated to your remote, and a typed push confirmation from you. It pushes exactly one branch and opens a draft PR for human review — nothing is ever pushed without those gates. Prefer full control? Stop after ship branch and push/PR yourself with plain git. Details: docs/experiment-mode.md

Open-source repositories (no push access): cloned someone else's repo? ship pr detects that you can't push to origin and switches to the fork workflow — behind a typed fork confirmation that spells out exactly what happens: a public fork is created (or reused) under your GitHub account, one branch is pushed to the fork, and a draft PR is opened on the upstream repository. One caveat to review first: if you committed a benchmark locally to establish the baseline, that commit is part of your branch's history and will appear in the PR — check git log (the CLI warns you when the branch carries extra commits). And read the project's CONTRIBUTING.md before opening PRs on repos you don't maintain — a draft PR with a reproducible, validated improvement is a good contribution; ten of them are spam.

Managing experiment runs

I want to… Command
start a batch (one command) researchforge experiment start plan.yaml
watch it live researchforge serve --background (URL is printed)
see ALL projects + their folders researchforge hub --backgroundhttp://127.0.0.1:9000
stop a running batch Ctrl-C — always safe (isolated worktrees)
continue an interrupted run researchforge experiment resume run-001
discard an interrupted run researchforge experiment abandon run-001
run another batch researchforge experiment plan hyp-002 → start again
build on a previous experiment parent: exp-001 in the next plan entry
see what's next, always researchforge status
see where everything lives researchforge paths
reset the whole project rm -rf .researchforge researchforge.yaml

Everything is local; nothing outside your repository is ever created. To redefine just the objective on existing data: researchforge project create --force-update. docs/claude-mode.md explains exactly what the Claude skills do and cannot do.

Supported repositories (beta)

The improve-repository journey currently expects:

  • a git repository with user-owned or trusted code (isolation is local, not a hostile-code sandbox — docs/security.md);
  • a Python 3.11+ single project or single target service;
  • an existing Dockerfile or simple Python dependency metadata (requirements.txt / pyproject.toml);
  • a machine-readable benchmark (a command that writes JSON metrics) with bounded runtime;
  • no production infrastructure required.

The explore-research-idea journey works anywhere. Repositories outside this matrix are reported honestly by researchforge repo scan rather than half-supported.

Beta feedback

This is a narrow-but-complete beta — reports shape what gets built next. Use the issue templates (bug, setup failure, beta feedback). Optionally, researchforge analytics enable records local-only coarse events — nothing is transmitted — and researchforge analytics show computes the beta metrics you can choose to include in a report.

More documentation

License

Apache-2.0 — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

researchforge-0.1.0.tar.gz (255.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

researchforge-0.1.0-py3-none-any.whl (215.2 kB view details)

Uploaded Python 3

File details

Details for the file researchforge-0.1.0.tar.gz.

File metadata

  • Download URL: researchforge-0.1.0.tar.gz
  • Upload date:
  • Size: 255.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for researchforge-0.1.0.tar.gz
Algorithm Hash digest
SHA256 c1e76a122ebc212a36d5e5a02be34ff8152858e17139723977996dd1d750b8f5
MD5 fd536a3eab667a2571a90dfbb31385cf
BLAKE2b-256 f1832b5e9f86ab0550d15e9343d62fe38fdc2e396b468b40e8aed640b25516f5

See more details on using hashes here.

Provenance

The following attestation bundles were made for researchforge-0.1.0.tar.gz:

Publisher: release.yml on forger-labs-hq/researchforge

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file researchforge-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: researchforge-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 215.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for researchforge-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e0634b59370f8664b03ff1ce7f7f7f6ca42025c8cc54fad013e23264566be2cd
MD5 11af6e07629f7476bbe7c39f3ad91e07
BLAKE2b-256 85fa8c66da431b269b51a3844db07cc94c6717d330238a970fc6a58ce8b653f1

See more details on using hashes here.

Provenance

The following attestation bundles were made for researchforge-0.1.0-py3-none-any.whl:

Publisher: release.yml on forger-labs-hq/researchforge

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page