This release is a pre-release and may not be stable for production use.
Outerloop
Autoresearch agents that improve your benchmark.
Outerloop runs AI agents on your own research code. An agent proposes a change, runs the experiment on your cluster, and opens a pull request only when your benchmark actually improves. Every attempt is written up, including the ones that failed.
You run it yourself: your keys, your compute, your repos. Nothing reports back to us. It is built and used every day by the Agentic Learning AI Lab at NYU, where it co-develops our research codebases.
How it works
- Propose. An agent picks a hypothesis and writes the code change.
- Experiment. It runs the training on your cluster and reads the results.
- Measure. Outerloop scores the change against the base tree at the same seed. Noise does not count as an improvement.
- Review. Reviewers read the change and the claim. If both hold up, a pull request opens.
- Record. Every attempt gets a short report: hypothesis, change, outcome, next step. Negative results included.
Agents cannot touch the benchmark, the budgets, or your CI. Your branch protection and required checks apply to them as to any contributor. By default a pull request waits for a human; a repo can also let clean ones merge themselves.
Get started
Three commands and one file. You need a repo with a benchmark command, an API key for the model that will write the code, and a Slurm cluster or one machine with a GPU.
pip install outerloop-science
outerloop init # where the loop runs, which repo, which model and its key, your GitHub identity
The wizard asks for a GitHub identity for the agents to open pull requests
as. Pick app and it walks you through creating a GitHub App, under your
account or under an organization you name, and installing it on the repo:
two browser pages and a code pasted back. It then checks that the App can
write the repo and tells you if it cannot. Pick pat if you already have a
token. It writes the config and the key files; nothing to edit by hand. Then
add one file, .outerloop.yaml, to the repo you want improved:
benchmarks:
- name: my-benchmark
command: uv run python -m mypkg.eval --json # prints {"success_rate": 0.42}
metric: success_rate
direction: max
budgets:
gpu_hours_per_run: 8
runs_per_week: 10
scope:
allowed: [src/] # the only paths an agent may change
roadmap: docs/roadmap.md # what the agents read for direction; never written
outerloop start # on a Slurm login node this submits the loop; without Slurm it runs in the foreground
Step by step, other model backends included: docs/install.md. Everything the contract can say: docs/contract.md.
Only want pull request reviews?
The reviewer works on its own. One workflow file and an API key, about five minutes, no bot account and no cluster. It comments on pull requests with concrete findings and never approves, blocks, or fails your build. See docs/reviewer.md.
Where it runs
The first-class home is a Slurm cluster. There is no daemon: the loop is a chain of short jobs that resubmit themselves, so nothing listens and no inbound SSH is needed. Experiments and evaluations run inside your container image with no credentials, and GPU-hours are metered against the contract's budget. A single machine with a GPU works too, for cheap benchmarks. Details: docs/compute.md.
Safety by design
- Opt-in and contract-bound. A repo takes part by granting the bot access
and committing a contract. The contract, your roadmap, and
.github/are never writable by an agent. - Nothing on trust. Outerloop measures every claim itself, on committed trees, and re-verifies before a pull request exists.
- Untrusted input. Pull request text, diffs, issues, web pages, and job output are data, never instructions. Agents run without credentials.
- Budgets in code. Launches, GPU-hours, and runs per week are enforced by the kernel, not left to the agent.
- No model lock-in. Claude Code, Codex, and hermes-agent are wired today; backends are swappable.
Full design: docs/design/architecture.md · Roadmap: docs/roadmap.md
Developing
uv sync
uv run pre-commit install
uv run pytest
| Path | Purpose |
|---|---|
src/outerloop/ |
The kernel: contract, tick (the Slurm chain), attempt/orchestrator (the climb), measure/dispatch (evals as jobs), syscall (the author's tool), panel/verifier/review, github, harness backends |
tests/ |
Tiers: unit (default), slow, llm, slurm markers |
scripts/ |
Committed operational scripts (the tick chain, provisioning) |
docs/ |
Install guide, architecture and design notes, roadmap |
License
Release files for outerloop-science 0.2.0rc3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| outerloop_science-0.2.0rc3.tar.gz | 1.0 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| outerloop_science-0.2.0rc3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.4 MB
Release files / outerloop_science-0.2.0rc3.tar.gz
| Download URL | outerloop_science-0.2.0rc3.tar.gz |
|---|---|
| Size | 1.0 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
efd24d9c43dcf0a154e27ab3d2b38809c52924598d65896c3c8d3a28df935938
|
|
BLAKE2b-256 checksum How to use checksums |
5157cd227636ccee95526ef0ae904d97d3d1a1791906c475e7e745969efb6253
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency logRelease files / outerloop_science-0.2.0rc3-py3-none-any.whl
| Download URL | outerloop_science-0.2.0rc3-py3-none-any.whl |
|---|---|
| Size | 430.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7c0b9acd33ef48db4b378e6ec0343ac8220fb976fc43d2d3cc3b4010ffdcf347
|
|
BLAKE2b-256 checksum How to use checksums |
7b27a78596d392e46a4fe97026d2c8fd522f14d9a45f92b5a3ec27163baed5b6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.
Transparency log