Rabbit Brain
Release review for iterative perception models: optical flow, stereo, depth, anything that refines an answer over iterations.
You have a current checkpoint and a candidate. Give Rabbit Brain the per-case errors of both, plus one number per refinement iteration if your model refines, and it ranks what to look at: the cases that got worse, and the cases whose error improved while the model never stopped changing its answer. Each flagged case comes with the numbers, the reasoning in plain language, and the evidence rendered as an image. The cases you cared about become saved checks, so the next checkpoint gets the same review and CI fails when one of them breaks.
pip install rabbit-brain
rb onboard # describe your setup once; get a configured project and an integration ladder back
If your model is RAFT-family, or you already know what goes in the config:
pip install "rabbit-brain[raft]"
rb init --project my-flow --adapter raft --model-code ./raft --dataset ./data/kitti2015/training
rb doctor && rb verify-hook --checkpoint ckpt/raft-things.pth
rb run --baseline ckpt/raft-things.pth --candidate ckpt/raft-small.pth # both checkpoints, trajectories recorded for you
rb findings <run> --top 5 # the ranked queue and a verdict
rb check save <run> <case> # keep a case for the next checkpoint
rb check run <run> # exit 1 in CI when something regressed
rb import results.json # or bring per-case results your evaluator already produced (rb schema example)
rb init --demo && rb run --baseline ckpt/synth-current.json --candidate ckpt/synth-candidate.json # a synthetic dry run on any machine
Nothing is sent anywhere. Every run leaves a receipt (report.md, record.json) with the exact command, the input's hash and the definitions in force, so a colleague, or you after your coding agent ran it, can reproduce the queue and see what every number was computed from. (rb share writes a file of anonymised statistics you may choose to send; the tool never sends it.)
The label-free half. How much a model is still moving its answer at the last iteration is measurable without a label, and it correlates with that case's error: Spearman 0.88 to 0.92 on 200 cases for each of the four public RAFT checkpoints. That is a correlation, not a verdict. A case that is still moving is worth looking at; it is not thereby wrong, and a settled case is not thereby right. So the same regression test runs twice: once on the error, and once on how much each model was still moving at the end, candidate against current, on the same case. The second one works on cases you have no ground truth for, and it is the reason a case that improved on error can still end up in the queue.
Why you should believe the numbers. Before a run reports anything, the adapter has to reproduce the model repository's own evaluation: rb verify-adapter runs your model through its own loader, forward call and metric formula and compares, and a disagreement beyond 0.001 relative stops the run and writes nothing. A harness that loads or scores a model slightly differently produces numbers that look plausible and are not about your model; this is the check for that, and every receipt records the result. Findings that turn on a margin thinner than the difference between two machines are marked borderline, so a colleague re-running your review on their own box knows which differences are real.
What it needs. Either a checkpoint pair and a case set for a supported adapter (RAFT-family optical flow today, RAFT-Stereo-style depth through a small custom adapter, and any other iterative model the same way; rb onboard writes that adapter's scaffold from a description of your setup, with the parts only you can write marked one by one), or one lower-is-better error per case for both models from your own evaluator (any metric with a unit: endpoint error in px, depth error in cm), computed against the same ground truth and valid mask. Trajectories are recorded by the adapter with a forward hook, or by the one-line TrajectoryRecorder in your loop. Without trajectories you get the error half and the tool says stability was not assessed.
What it doesn't do. Train anything, or certify a model. The stability limits are generic heuristics and a starting point; a scorer fitted to your model is a separate, paid evaluation. It does not replace your evaluator either: the error is whatever your metric says it is.
What a flagged case looks like. For each case it flags, rb renders the evidence: the inputs, where the two models disagree (no ground truth needed), both flow fields, the per-pixel error of each and where it changed, the per-iteration filmstrips and the trajectory. The decision is a look, not a number.
Below is the one regression in a real review of raft-sintel against raft-things on 200 KITTI-2015 pairs (examples/raft-kitti): +0.41 px on this case, and the sheet shows it is one object, the pole nearest the camera, which both checkpoints keep moving through all twelve iterations. Those two checkpoints were trained on different data, so the mean improvement is a domain gap rather than a release decision; the point of the example is the one case that got worse anyway, and that the receipt names it.
In CI. examples/ci/github-actions.yml is a working release gate: it imports your evaluator's results, runs your saved checks, writes the verdict into the pull request and keeps the run as an artifact. It exits non-zero when a saved check fails, and also when your policy declares a limit this version cannot evaluate, so a required measurement that was never taken does not quietly pass. One thing that file is careful about and worth repeating here: a red job is not a blocked merge. Someone with repository admin has to add the job to branch protection as a required status check before the gate enforces anything.
For coding agents. rb docs prints AGENTS.md: the workflow, the file format, the JSON output (--json on every command), exit codes and every error code with its fix. Tell your agent: "Review candidate checkpoint B against A on this case set with Rabbit Brain."
Install and versions. pip install rabbit-brain (core, pydantic only), pip install "rabbit-brain[raft]" (torch and the RAFT adapter's needs), pip install "rabbit-brain[evidence]" (numpy and pillow, for evidence sheets with a custom adapter). Python 3.10 or later. CHANGELOG.md for what each version changed.
The trajectory diagnostic comes from a paper that is not public yet; the reference goes here when it is.
Apache-2.0.
Metadata
Release files for rabbit-brain 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| rabbit_brain-0.4.0.tar.gz | 164.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| rabbit_brain-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 285.5 kB
Release files / rabbit_brain-0.4.0.tar.gz
| Download URL | rabbit_brain-0.4.0.tar.gz |
|---|---|
| Size | 164.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
da334865828013c409d08be89824f8a7dad396ff23fcdb97c1fe06bdfee77685
|
|
BLAKE2b-256 checksum How to use checksums |
5913d9ea6790f9ca4f81706b6776b5c6fb307a5705757a70064ee7567ed165a7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency logRelease files / rabbit_brain-0.4.0-py3-none-any.whl
| Download URL | rabbit_brain-0.4.0-py3-none-any.whl |
|---|---|
| Size | 121.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c8bd4ac9595b17011b60a1f89142f00b1c8e6004ac046a83dfa18999f5433972
|
|
BLAKE2b-256 checksum How to use checksums |
c0d21227d915b7355eb39a156b1401bc1438218a151fff8703aaf93ea1869d45
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency log