repo-bug-hunter
An AI coding agent that fixes real GitHub bugs, and a lab that measures how well it does it.
Give it a GitHub issue and it works the way a developer would: it explores the repository, reproduces the bug, edits the code and runs the tests until it can submit a fix. The fix is then graded by the project's own hidden tests from SWE-bench Verified, and every step the agent took can be replayed in the browser.
What it does
- Fixes bugs from real GitHub issues inside a Docker sandbox, using six simple tools.
- Grades each fix with the tests the project's maintainers wrote for the real fix.
- Replays every run step by step on a static web page.
- Compares two versions of the agent on the same tasks, with paired statistics.
- Builds new tasks from merged GitHub pull requests, to test on bugs newer than the model.
- Works with Claude, any OpenRouter model, local models in Ollama, or any OpenAI-compatible server.
The agent is deliberately small: one file of under 200 lines, a plain loop around a chat model, with no agent framework, so every prompt and tool call is easy to follow. The sandbox is the same Docker image the grader uses and has no network, so the agent can't install packages or look up the real fix.
Quick start
pip install repo-bug-hunter # Python 3.10+; or: uv tool install repo-bug-hunter
repo-bug-hunter demo # replay example runs in your browser: no key, no Docker
To run the agent yourself, it needs two things pip can't install:
- A model. A free key from OpenRouter is enough:
export OPENROUTER_API_KEY=.... Claude and local models work too; see Models. - Docker, running, with about 30 GB of free disk space, since each task's image is a few GB. On Apple Silicon, turn on Rosetta in Docker Desktop's settings; the images are x86-64.
repo-bug-hunter doctor # checks both, and says what to fix
repo-bug-hunter smoke # fixes a toy bug end to end, in a few minutes
No Docker yet? repo-bug-hunter smoke --local runs the toy bug in a temporary folder on your machine instead. It runs on macOS and Linux; on Windows, use WSL2 (wsl --install) or Codespaces.
Then run it on real bugs. Results go to runs/ in the current folder:
repo-bug-hunter run --name pilot --n 5 --difficulty "<15 min fix"
repo-bug-hunter evaluate runs/pilot # grade with the official SWE-bench harness
repo-bug-hunter analyze runs/pilot # resolve rate, cost, steps, why tasks failed
repo-bug-hunter viewer runs/pilot # build the replay site in ./site
python3 -m http.server -d site 8000 # open http://localhost:8000
Without installing anything
- GitHub Actions: fork this repository and enable workflows in the fork's Actions tab. Add
OPENROUTER_API_KEYunder Settings → Secrets and variables → Actions, and set Settings → Pages → Source to GitHub Actions. Then run the experiment workflow. Each task runs on its own GitHub machine, the official grader scores it, and the replay site is published to your GitHub Pages. - Codespaces: click the badge at the top for a ready-made environment in the browser, then run the commands above with
uv runin front, such asuv run repo-bug-hunter doctor.
The test-first experiment
The question this project was built to answer: does an agent fix more bugs if it must first reproduce the bug with a failing test?
In test_first mode, the harness enforces it, not the prompt:
- Source files are locked until the agent registers a test command that fails on the current code.
- When the agent submits, the harness runs that test again. If it still fails, the submission is sent back once.
- So the agent can't get stuck: after three tests that don't fail, the source unlocks anyway, and a second submit is always accepted.
Run both modes on the same tasks and compare:
repo-bug-hunter run --name baseline --variant baseline --n 50
repo-bug-hunter run --name test_first --variant test_first --n 50
repo-bug-hunter evaluate runs/baseline && repo-bug-hunter evaluate runs/test_first
repo-bug-hunter analyze runs/baseline runs/test_first --out results.md
The report pairs the two runs task by task, using McNemar's exact test and a bootstrap confidence interval for the difference, so noise from a small sample isn't mistaken for an improvement.
How it works
- agent.py is the loop. It sends the conversation to the model, runs the tools the model asks for, and repeats until the agent submits or hits a step or cost limit. With Claude it uses prompt caching and adaptive thinking.
- tools.py has the tools (
read_file,search,edit_file,write_file,bash,submit), the test-first gate, and the code that turns the final repository into a patch. - env.py runs everything in the task's own SWE-bench Docker image, without network access, and caps output, file reads and processes so a runaway command can't hurt the machine.
- evaluate.py grades with the official harness. A bug counts as fixed only if the hidden tests for the real fix pass and nothing that passed before breaks.
- analyze.py reports the resolve rate, cost and steps, and gives every failure one reason: gave up, wrong file, broke other tests, and so on.
- viewer.py builds the replay site, and pr.py builds tasks from merged GitHub pull requests:
repo-bug-hunter pr more-itertools/more-itertools#1305 --name fresh
Models
| Provider | How to choose it | API key |
|---|---|---|
| Anthropic | --model claude-sonnet-5-5 |
ANTHROPIC_API_KEY |
| OpenRouter | --model vendor/model |
OPENROUTER_API_KEY |
| Ollama (local) | --provider ollama --model NAME |
none |
| Other OpenAI-compatible servers | --base-url URL --model NAME |
LLM_API_KEY |
Without --model, a free OpenRouter model is used, so you can try everything without paying. repo-bug-hunter doctor lists the free models that currently support tool calling. Free models allow about 50 requests a day, roughly two tasks, or 1,000 a day after buying $10 of OpenRouter credits once. When a limit runs out, the run stops cleanly, and running the same command later continues where it stopped. Free providers may log prompts, so don't point them at private code.
Local models need tool calling and a context window of at least 32k tokens. Ollama's default is 4k, so create a variant first:
printf 'FROM gemma4\nPARAMETER num_ctx 32768\n' > Modelfile && ollama create gemma4-32k -f Modelfile
Limitations
- SWE-bench Verified is public, so strong models may have seen some of the real fixes; OpenAI stopped reporting it in February 2026 for this reason. Paired comparisons are less affected, and
repo-bug-hunter prtests on fresh bugs. repo-bug-hunter prsupports Python projects tested with pytest, and the pull request must name the issue it fixes.- Failure reasons are heuristics. For example, "wrong file" compares the patch with the files the real fix changed, so a correct fix in another file would be miscounted.
Development
git clone https://github.com/Manishmaurya89/repo-bug-hunter && cd repo-bug-hunter
uv sync
uv run pytest # the Docker tests run only when Docker is up
- IMPLEMENTATION.md: how each part is built and why, including the bugs found along the way.
- LEARNING_GUIDE.md: every idea explained from zero, with exercises.
License
Metadata
Release files for repo-bug-hunter 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| repo_bug_hunter-0.3.0.tar.gz | 147.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| repo_bug_hunter-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 291.8 kB
Release files / repo_bug_hunter-0.3.0.tar.gz
| Download URL | repo_bug_hunter-0.3.0.tar.gz |
|---|---|
| Size | 147.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8e74acff6147bacd7da75d6d5fe2840dd97037a2517da362c04b03090071b2ad
|
|
BLAKE2b-256 checksum How to use checksums |
88f9499d6a926dd5a6ffdb149477a32e00d89f17c44e492dcc17e6e3cb9a35f5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / repo_bug_hunter-0.3.0-py3-none-any.whl
| Download URL | repo_bug_hunter-0.3.0-py3-none-any.whl |
|---|---|
| Size | 144.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4194bfc8e95ac1068c2e9e0495a4b218e9c6129a2abbcdbdc2c6c2a500b0950e
|
|
BLAKE2b-256 checksum How to use checksums |
7b2db6c75f19b3bd4f974177a0f78c215755b5f91bb5cdae7ed81d5e6baf87eb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log