Skip to main content

hillclimb

Hillclimbing on verifier-defined problems. You give it a problem — a folder whose verifier.sh scores a solution.py — and a budget. It spawns coding agents (Claude Code, Codex or pi) as operators that draft, debug and improve solutions, scores every candidate through your verifier, and keeps the best. The harness is fixed. The climber — what to try next and how each attempt is prompted — is a bundle you can swap, edit and share, so two methods can be compared on the same problem under the same budget.

Get started

Eight commands, ten minutes. Each step says what you should see.

1. Install and connect an agent.

pip install hillclimb
hillclimb connect claude       # logs in; operator calls bill your Claude subscription

connect codex and connect pi do the same for the other backends. No agent yet? Add --backend dummy to any run below: it climbs with no LLM at all.

2. Make a hillclimb dir, get a problem, read it.

hillclimb init                         # hillclimb/ here: config.yaml, problems/, runs/
hillclimb problem list                 # the bundled starter problems, with the best known value
hillclimb problem get heilbronn-11

Everything hillclimb writes lives under that one hillclimb/ folder. The problem lands in hillclimb/problems/heilbronn-11/. Open verify.py: the verifier is the problem. Everything the agents will be told is in description.md.

3. Check the verifier.

hillclimb verify heilbronn-11 --repeat 3

Scores the problem's floor three times and prints the spread. Starter problems are exact, so the spread is 0 and any improvement is real. On a noisy problem of your own, this number is the noise floor the search must beat.

4. Climb.

hillclimb run heilbronn-11 --budget 10m

One search in the foreground: a baseline, then drafts, debugs and improves, each scored as it lands. Ctrl-C stops it; the best solution so far is kept.

5. Watch it, in a second terminal.

hillclimb watch      # every agent, what it is doing, its candidate's score
hillclimb chart      # best score so far against time, every candidate a dot
hillclimb tree       # the exploration tree: what was expanded, what was left

The result is hillclimb/runs/<run-id>/searches/heilbronn-11/best/: solution.py and the submission.csv it wrote.

6. Stop everything.

hillclimb stop --all

7. Next.

  • Your own problem. Copy a starter and edit verify.py, or hillclimb init for a blank scaffold. See docs/problems.md.
  • Another climber. hillclimb run heilbronn-11 --climber openevolve. hillclimb climber list shows the bundled ones; hillclimb climber new mine --from greedy copies one into hillclimb/climbers/mine/ for editing. See docs/climbers.md.
  • Compare two. hillclimb run heilbronn-11 --climber greedy --climber openevolve --parallel-searches 2, then hillclimb experiment report <run-id>. See docs/experiments.md.

Starter problems

Construction problems in one shape: the submission is a small CSV of numbers, the verifier checks the constraints and computes the score exactly, there is no dataset, no holdout split and no noise. Each ships with the best known value as a reference line on the chart, and each family is a ladder, so a ten-minute run visibly climbs on the small instance and an hour does not saturate the large one.

Family Instances Score
Circle packing (AlphaEvolve's benchmark) circle-packing (26), circle-packing-32 sum of radii, maximize
Heilbronn triangles heilbronn-11, -14, -17, heilbronn-convex-13 smallest triangle area, maximize
Low-autocorrelation binary sequences labs-40, labs-60 sidelobe energy, minimize
Tammes (points on a sphere) tammes-30, tammes-50 minimum angle, maximize
Thomson (charges on a sphere) thomson-50, thomson-100 Coulomb energy, minimize
Autocorrelation inequalities (AlphaEvolve) autocorr-1, autocorr-3, erdos-overlap the inequality's constant
Kissing configuration in dimension 11 kissing-11 number of points, maximize
Golomb rulers golomb-20, golomb-27 ruler length, minimize
Travelling salesman on a fixed instance tsp-200 tour length, minimize

hillclimb problem list prints the catalog with the best known value and who found it. Problems that need data, a hidden split or a provider (Kaggle competitions through MLE-bench, energy forecasting through emflow) are documented in docs/providers.md.

Climbers

A climber is a directory with a climber.yaml naming a search policy (or a whole loop), the operators it may use, their prompts and a tuner. Three are bundled: greedy (debug failing tips, draft a few branches, improve the best), openevolve (MAP-Elites over hillclimb's operators) and gepa (reflective Pareto search, brings its own loop). A one-file climber is a .py with one policy class. A search snapshots its climber, so editing the live copy never changes a running search, and hillclimb climber check replays recorded journals through an edited climber before an agent hour is spent on it.

What is where

Path What it is
src/hillclimb/harness/ The fixed core every search runs on: core.py (the Harness), the loop, evaluation and the executor, the journal, candidates and the store, budgets, slots and the control queue. Never a research surface.
src/hillclimb/modules/ What a climber exchanges, one subpackage per kind, each with its contract in base.py: policies/ (what to try next), operators/ (how one attempt is made), tuners/ (which parameter values), similarity/ (how alike two solutions are), memory/ (the file-based memory, and the graph module that indexes it). Implementations import only hillclimb.sdk.
src/hillclimb/sdk/ The one import a climber needs: the contracts and the read-only views of the search.
src/hillclimb/climbers/ The bundled climbers, greedy, openevolve and gepa, each a climber.yaml naming its modules and prompts. hillclimb climber new copies one for you to edit.
src/hillclimb/tui/ Every terminal view (watch, chart, tree, archive, surface, similarity, graph) and the layout it draws. Reads the store, imported by nothing else.
src/hillclimb/cli/ The hillclimb command, one module per command group.
src/hillclimb/backends/ The agents that write code: Claude Code, Codex, pi, and the dummy and fake backends for tests.
src/hillclimb/integrations/ Problem providers and libraries that bring their own loop: emflow, MLE-bench, Einstein Arena, GEPA.
src/hillclimb/prompts/ The operator prompt templates. A climber may shadow them by name.
src/hillclimb/runtime/ The managed venv the verifier and the solution run in, and the shim that makes hillclimb.spaces importable there.
src/hillclimb/demo/ The starter problems as package data, so hillclimb problem get works from a bare install.
src/hillclimb/spaces.py The output-format contract a problem's interface.py is written in, and the params.json contract. Stdlib only, byte-copied into runtime venvs.
src/hillclimb/{api,config,problem,climber,experiment,connect}.py The public surface: run a search, the config schema, load a problem or a climber, experiments, and connecting an agent.
problems/ The starter problems' source of truth, one make_<family>.py generator per family; the bundled copies under demo/ are stamped from here.
hillclimb/ This repo's own hillclimb dir: config.yaml, experiments/ (specs and seeds), knowledge/ (the graph, cards, credit, playbooks).
tests/ The suite (uv run pytest). test_layout.py pins which package may import which, test_sdk_imports.py that climber code imports only the sdk, golden/ the prompt bytes and every --help screen.
docs/ The topic docs linked below, and the dated design plans.

Docs

  • Problems — the verifier contract, floors, unit tests, tunable parameters, per-instance scores, noise
  • Climbers — bundled climbers, the manifest, one-file climbers, mixed fleets, climber check, GEPA
  • Agents — Claude Code, Codex, pi; connect; billing through OpenRouter; sampling
  • Experiments — arms, repeats, matched budgets, the report
  • Providers — emflow, MLE-bench, Einstein Arena
  • Operators and memory — operator scaffolds, model routing, the knowledge graph
  • The hillclimb dir — config precedence, run specs, how runs are laid out, the store, pruning
  • Commands — every command, the TUIs, the demo suite
  • Development
  • Changelog

Release files for hillclimb 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hillclimb 0.4.0
File Size Uploaded
hillclimb-0.4.0.tar.gz 484.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hillclimb 0.4.0
File Interpreter ABI Platform
hillclimb-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.1 MB

Release files / hillclimb-0.4.0.tar.gz

Download URL hillclimb-0.4.0.tar.gz
Size 484.1 kB
Tags Source
SHA-256 checksum
How to use checksums
acd97c067264d1f11c5fb744414b8952e851bae5a0114f6204ae9408d48739ce
BLAKE2b-256 checksum
How to use checksums
047cf0c26dc811c2dbb8aafb43ca6d9e1154da619211bbdfd1ac3935de93d679
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.2

Release files / hillclimb-0.4.0-py3-none-any.whl

Download URL hillclimb-0.4.0-py3-none-any.whl
Size 642.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a87d4947ea36fff5991f2bee8e44a2ac000bc12afd103392185dfc3a4190713d
BLAKE2b-256 checksum
How to use checksums
143776d3719829f0f9e59529246eef732b0f213044997b77f9381b0cfd6fbffa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.2

Release history Release notifications | RSS feed

0.5.0

2 release files

This release

0.4.0 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page