llm-panel
Put the same question to several models independently, then read every answer in full.
Judges run in parallel, never see each other's work, and answer from their own reading of your repo. An optional second round shows each of them the others' findings — anonymised — and asks them to defend or withdraw. The output is a single self-contained HTML page where that second round is grouped by the finding being argued about, so comparing what five models said about one line of code doesn't mean holding five documents in your head.
It is not a voting machine. A panel generates candidate defects; it does not establish
truth by counting agreements. Every finding still has to be checked against the code — and
the tool's other half, recall/, exists to measure what the
panel misses rather than assert what it catches.
llm-panel --diff "Which of these changes is most likely to be wrong?"
panel-report --open # render the newest run and open it
panel-triage --bad # which runs went wrong, across every run root
A real run, with the waiting compressed: three judges asked in parallel (two free-tier,
one on a ChatGPT plan) landing as they finish, the scoreboard from panel.md, and
panel-report rendering it (recording):
The rendered report — the scoreboard counts spend and names who answered; the
citation-overlap tables show where the panel's attention landed (three judges reviewing a
cline PR, converging on one line of TerminalProcess.ts):
Contents: What's here · Install · Configure your roster · Using it · What it actually catches · On real PRs · Tests · Known limitations
What's here
| tool | what it does |
|---|---|
llm-panel |
asks the judges, in parallel, and writes the run to disk |
panel-report |
renders a run as one self-contained HTML page, grouped by claim |
panel-triage |
finds the runs that failed, which a listing shows as ordinary rows |
recall/panel-recall |
measures what the panel misses, against a corpus of planted defects |
recall/aacr-upstream |
runs the panel over AACR-Bench PRs and hands the findings to upstream's evaluator |
recall/aacr-score |
invokes that evaluator, and refuses to report a number from a judge that isn't running |
claimlib.py |
the one measurement boundary: reviews → span-grounded observations |
*-controls |
the regression suites — 1089 controls, every one tied to a defect that shipped |
Install
Pure Python 3.11+ standard library on Linux, macOS or WSL — it needs POSIX file locks and process groups, and says so on Windows instead of tracing back. No dependencies, no build step. Each tool is one readable file, so either install route runs identical code:
# as a package (entry points: llm-panel, panel-report, panel-triage)
uv tool install llm-panel # or: pipx install llm-panel
# or as the files themselves
git clone https://github.com/musharna/llm-panel ~/llm-panel
ln -s ~/llm-panel/{llm-panel,panel-report,panel-triage} ~/.local/bin/
3.11 is a hard floor (the link renderer uses atomic groups, added to re in 3.11);
panel-report says so at startup rather than failing part-way through a render.
Judges reach models through command-line tools you install separately — none are bundled, and you need at most one to start:
| tool | who it is | billing |
|---|---|---|
codex |
OpenAI's CLI | a ChatGPT plan, not metered API |
opencode |
multi-provider CLI most judges route through | your OpenRouter / HuggingFace keys |
claude |
Anthropic's CLI | a claude.ai subscription (setting ANTHROPIC_API_KEY switches it to metered) |
ollama |
local models | free, and no tool loop — see the caveat below |
If none are present the panel still runs, fails loudly, exits 4, and tells you what to install. A missing tool is one judge's problem, never the whole panel's.
Two transports skip the CLI: ollama uses its local HTTP API, and orvision calls
OpenRouter's HTTP API directly with your OpenRouter key so that the vis-* judges (grok,
kimi, gemini, gpt) can look at an --image.
The read-only agent for opencode judges
opencode judges run as an agent named panelist that can read the repo but not write
to it, and llm-panel refuses to start an opencode judge until that agent is defined and
verified read-only (exit 9) — opencode's default build agent will happily edit the tree
it is reviewing. The definition ships as
opencode.jsonc: merge its
agent.panelist block into ~/.config/opencode/opencode.jsonc, or keep the file at the
root of a repo you review with it — opencode reads project-local config too.
Configure your roster
The built-in judge list is a default, not a fixture — it names the author's accounts. Yours will be different. Point the roster at models you actually have:
Copy roster.example.json to
~/.config/llm-panel/roster.json ($XDG_CONFIG_HOME honoured; $LLM_PANEL_CONFIG wins).
It is strict JSON — no comments, no trailing commas — and a malformed config is fatal
and names the offending key, because quietly falling back to the built-in roster would run
a panel you didn't ask for, and bill you for it:
{
"default": ["codex", "nemotron", "glm", "kimi", "or-deepseek", "or-grok"],
"judges": {
"my-gpt": {
"transport": "opencode",
"model": "openrouter/openai/gpt-5.6",
"family": "OpenAI"
},
"big-pickle": null
}
}
null drops a shipped judge. default is the panel run when --judges is absent.
llm-panel --list shows the roster offline and marks config-defined judges.
llm-panel --check actually pings each one. llm-panel --help-config prints this schema.
Picking judges
"One per vendor" is not the answer. It is tempting to treat vendor labels as a proxy for independent opinions. The evidence says they aren't: Kohli 2026 measured cross-family judge correlation at φ̄=0.389 against same-family 0.437 — barely different — with the three most correlated pairs being cross-family, and found that restricting to one judge per family made effective independence worse (n_eff 1.93 vs 2.18). Family is display metadata here, not policy.
The six-judge set above did score 6/6 against the planted-defect corpus described under What it actually catches, where a two-vendor panel scored 4/6, but treat that as debugging evidence, not as a result: the roster was repaired because of what happened on those very fixtures, so the comparison is in-sample, and the six defects live in only two files (effective n≈2, 95% CI 61–100%).
What to actually do: pick judges by what they find on your code, and use panel-recall to
measure it. The quantity worth maximising is each judge's marginal rescue rate — how
often it catches something every other judge on the roster missed — not how many logos are
represented.
Two practical constraints: wall-clock is the slowest judge, not the sum, so one slow
model sets the pace for every run; and claude-* are deliberately absent from the default,
because when Claude wrote the code under review a Claude judge shares the author's blind
spots. Add it explicitly when that isn't the case — it is strong.
Using it
--diffattaches the working-tree diff, so nobody has to describe the change — including you, who would describe it favourably.--rebutadds the anonymised second round. Worth it whenever a finding would trigger real work: the first run of it killed three confident findings that were simply wrong. To be precise about the report's grouping of that round: it keys on the rebuttal letter each finding is given (A1, B2 …), so it collects the discussion around one judge's finding. It is not semantic clustering — two judges independently raising the same underlying defect stay two findings, and without--rebutthere is no grouping at all.--judges a,b,coverrides the default panel.codex~2runs the same model a second time as a full, separate judge — its own file, its own letter, its own row. Collapsing repeats would hide exactly the disagreement that makes them worth running.--thread NAMEkeeps a persistent conversation per judge. For design questions, not review.--image PATH(repeatable) attaches an image. Only thevis-*andclaude-*judges can look at it; every other judge reportsunavailablerather than answering blind, and--vision-check TEXTmakes each judge quote something visible before it is believed.--liveprints each answer the moment it lands instead of waiting for the slowest judge;--streamechoes tokens as they arrive, which only ollama and claude judges can honour. Without either, a heartbeat still names who is still working.--runslists past panels for this repo (--all-reposfor every repo) and--showprints the latest report.--effort {low,…,max}sets reasoning effort where a judge has the setting;--timeout SECONDScaps each judge, and a judge over the deadline is killed as a whole process group and reportedharness.-f FILEreads the prompt from a file.- Long questions go via stdin:
llm-panel - <<'ASK' … ASK. --synthesize JUDGEasks one judge to fold every answer into a single synthesis after the round;--cwd DIRreviews a repository other than the current one;--save-herewrites the panel into the reviewed repo as well as the cache;--agent NAMEpicks the opencode agent (defaultpanelist) and--keep-alivethe ollama model residency.panel-reporttakes--repo SUBSTRto pick a run root,--out FILE,--webfontsand--max-image-kb;panel-triagetakes--since HOURS,--repo,--limitand--json.--usageshows what thecodexjudge is spending: the ChatGPT plan's 5-hour and weekly windows, when each resets, and the "Full reset (Weekly + 5 hr)" credits OpenAI banks on the account.--reset-usageredeems one of those credits — it prints the same screen, then asks you to typeRESET, because a credit is finite and a script should not be able to spend one by passing a flag. Both read the account throughcodex app-server, the same channel the interactive/statusscreen uses, and cost no quota themselves.
On pull requests
The repository doubles as a GitHub Action. It builds a prompt from the PR's diff, runs the panel in the checked-out tree, and posts every judge's answer in full as one comment, edited in place on each push rather than added to:
# .github/workflows/panel.yml
on: pull_request
permissions: { contents: read, pull-requests: write }
jobs:
panel:
if: github.event.pull_request.head.repo.full_name == github.repository # forks have no secrets
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: musharna/llm-panel@v0.1.6
with:
openrouter-api-key: ${{ secrets.OPENROUTER_API_KEY }}
# judges: or-glm,or-kimi,or-deepseek timeout: "600" extra-args: --rebut
The default judges are the three OpenRouter ones, so one key is the whole setup. They
read the checked-out tree, not just the diff: on this repository's own 9-file PR the three
spent 300k–990k input tokens each and billed $0.68 and $1.12 for the panel on two
runs, 4.5 minutes wall clock, with kimi-k3 the largest share both times. The job fails on
exit 9 — the PR's tree carries .opencode/ or claude hooks the judges would run — and
posts the panel on 0 or 4. This repository runs it on its own pull requests
(.github/workflows/panel.yml, installing from source).
Judges reading through codex/opencode/claude can read your repo. ollama judges
answer from the prompt alone with no tool loop, so they cannot verify a claim against code.
Treat their findings accordingly.
Exit codes are deliberate and llm-panel --help lists them: 0 every judge answered ·
1 usage, config, or a failure of this program · 2 --file could not be read · 3 --check
found a judge that could not answer at all · 4 degraded panel (a judge never ran — our
failure, reported as such) · 7/8 --diff could not produce a diff / had nothing to review ·
9 the opencode agent or the reviewed tree is not verified safe · 10 --thread is locked by
another run · 11/12 --repeat out of range / a repeat suffix typed by hand · 13 illegal
judge name · 14 none of the selected judges has its CLI installed, with the roster path in
the message · 130 interrupted (Ctrl-C or SIGTERM), with whatever landed kept in the run
directory.
The rebuttal round as rendered — every position each judge took on each finding, grouped
by the finding under dispute, disagreements marked CONTESTED. This run: four free-tier
judges asked to review llm-panel's own failure-classification code; one failed and is
reported as harness, the other three upheld 7 findings, rejected 4, and missed 6:
What it actually catches
recall/panel-recall is the part most tools like this don't have: a corpus of defects
planted in real code, each one proven to misbehave by execution, so "the panel missed
it" is a measurement rather than an impression.
At least one of four independent passes (codex ×2 + claude-opus ×2) matched 25 of 27 known targets in this controlled, single-file Python corpus. That is a keyword-matched lower bound on an easy corpus — not an estimate of real-world code-review capability. 95% CI 76.6–97.9%, and that is before accounting for defects clustering within fixtures.
Read that next to a real-world number. CR-Bench
(Nutanix, 2026) builds review tasks from real bugs git blamed out of merged PRs in
django, sympy, astropy and scikit-learn, and reports GPT-5.2 + Reflexion at 32.8% recall
and 5.1% precision. The gap between that and 25/27 is the corpus, not the panel: hand-
planted single-mechanism defects in ~40-line files are far easier than real defects in
mature codebases, and the two numbers are not even the same estimand — different agents,
different context, different definitions of a hit.
So this corpus is a development instrument, good for controlled A/Bs where ground truth must be known and iteration must be cheap (the abstention experiment below is exactly that). It is not evidence of absolute capability, and no number from it should be quoted as one.
Three results worth knowing before you trust any of the output:
-
Recall was limited by the roster, not by the models. The two defects that panel never found — a
.get(k, default)that doesn't apply to an explicitnull, and a corrupt cache file silently becoming empty — are both found by a six-vendor panel (OpenAI / NVIDIA / Zhipu / Moonshot / DeepSeek / xAI): 4/6 → 6/6 on those two fixtures. The best two judges there, at 4/6 each, beat codex at 2/6 — and both were broken or out of credit until the roster was repaired. If your panel is missing things, check who is actually answering before concluding the models can't see it. -
Running the same model twice recovered nothing. First passes 25/27, with repeats 25/27. The repeat-passes idea is well supported in the literature and did not reproduce here. An earlier grader bug reported +1 and it was an artifact. Adding a different vendor did what adding a second pass of the same one could not.
-
Letting judges say "nothing is wrong here" is a precision/recall trade, not a free win either way. One sentence of abstention licence is the whole difference.
findings/fixture false positives licence on 0.42 0 / 6 judges licence off 2.17 2, from 1 of 5 judges Findings-per-fixture is measured on fixtures that do contain defects, where the extra findings were verified true — so the licence suppresses real findings (one judge went 3.00 → 0.00 on files with genuine defects). False positives are measured on
p01-exhaustive-codec, the one fixture with proven absence rather than verified scope — which is what makes a false-positive rate computable at all. There, the same judge on the same code abstained with the licence and produced two demonstrably false findings without it (it claimed int and str subclasses were rejected;encode(MyInt(1))returns'A').So: the licence costs true findings and prevents false ones. Which you want depends on whether chasing a false lead costs you more than missing a real defect. Caveat worth stating: one proven fixture, eleven reviews.
On real PRs (AACR-Bench)
The planted corpus above is a development instrument; the real-world numbers come from
running the panel over AACR-Bench PRs and
scoring the findings with upstream's own evaluator — a real LLM judge doing
path → line → semantic matching, so the numbers are theirs, not a self-graded matcher's.
The prompt style is a flag of that harness, recall/aacr-upstream --prompt-style, not of
llm-panel. On 18 PRs at full roster. The default row is the shipped prompt measured at
e2ad666 (2026-09-06, recall/benchmarks/results-0.1.4-3judge/); the other two are
extractor-3 re-measurements from 2026-08-28:
--prompt-style |
semantic recall | precision | findings read per validated hit |
|---|---|---|---|
defect (default) |
9.8% | 9.5% | 10.5 |
broad |
26.0% | 13.2% | 7.6 |
volume |
25.2% | 7.9% | 12.6 |
broad — asking for what a careful maintainer would actually raise — doubles the recall
of the defect arm it was paired against (McNemar on paired references, p = 0.0005; that
arm was the pre-rewrite prompt at 12.2%, not the row above). But the volume control
shows what that class of gain is made of: it is the defect prompt plus one
exhaustiveness clause, reaches the same recall (p = 1.0 vs broad), and pays for it with
half of broad's precision. On a 35-PR replication the ordering holds on both transports
while every arm's precision falls (broad ~9.7%, volume ~5.5–6.1%, ~16–18 findings read
per hit). A declared cost cut over all of it settled the product default: it stays
defect; the only candidate for a future default change is broad
(recall/benchmarks/cost-cut/README.md).
Against the paper's own baselines, the panel's precision is ordinary and its recall is low. AACR-Bench's Table 3 (v3, 2026-01-30) reports single models on all 200 PRs under a "No context" condition — the diff plus the PR title and description, no retrieved repository code — which is the closest published condition to the diff-in-prompt arm above:
| paper, "No context", all 200 PRs | recall | precision |
|---|---|---|
| GPT-5.2 | 47.1% | 7.0% |
| Claude-4.5-Sonnet | 42.9% | 8.7% |
| DeepSeek-V3.2 | 36.5% | 5.6% |
| GLM-4.7 | 27.6% | 11.3% |
| Qwen-480B-Coder | 27.4% | 9.4% |
this panel, defect, 18 PRs |
9.8% | 9.5% |
this panel, broad, 18 PRs |
26.0% | 13.2% |
The rows are not directly comparable and the gap should be read with that in mind:
ours is an 18-PR subsample, scored by upstream's evaluator code with claude-opus-4.5 as
the judge where the paper used Qwen3-235B, without the PR title and description, at line
tolerance k = 1 where the paper says only "overlaps", and it is a three-judge panel of one
subscription model and two free-tier ones where every paper row is a single frontier
model. The paper's agentic condition (Claude Code with repository access) scores 10.1%
recall at 39.9% precision, so the paper itself shows recall and precision trading against
each other by an order of magnitude across conditions. What can be said: the defect
prompt sits at the low-recall end of that spread, broad sits inside the paper's
no-context recall range at better-than-paper precision, and nothing here has been measured
on the full 200.
Three things to know before quoting any of it:
- The variance floor is measured. Re-running the same judge on the same 35 PRs moves up to ±3 human-reference matches of 150, with an evaluator replicate at exactly zero — so effects under ~5–7 pp of recall are re-run noise at this n, which every subgroup claim so far was.
- Three earlier readings were withdrawn on re-measurement — a DEFECT/IMPROVEMENT split (the classifier was circular), "broad finds different hits" (pre-registered replication on 35 fresh PRs, p = 0.40), and a transport effect that did not survive a re-run. Nothing above rests on a withdrawn claim.
- Location agreement overstates semantic agreement ~2.5x (25.2% of references had a finding at the right file and line; 9.8% had one a judge called the same concern) — which is why scoring is delegated upstream instead of done by a local matcher.
The rest — the repo-checkout arm, what a degraded roster costs, accepted-vs-rejected
comments, why unlocated findings are withheld from upstream — with every run ledger and
the data licensing, is in
recall/benchmarks/README.md.
Tests
./claimlib-controls # 90
./llm-panel-controls # 427
./panel-report-controls # 319
./panel-triage-controls # 19
./recall/aacr-upstream-controls # 96
./recall/aacr-recut-controls # 27
./privacy-controls # 10
cd recall && ./panel-recall selftest && python3 validate_corpus.py
CI runs all seven suites on every push (Python 3.11, 3.12 and 3.13).
Every control corresponds to a defect that shipped, and each asserts the fixed behaviour and — where the pre-fix input is representable — that the broken version would have failed on it. An assertion that passes on both the broken and the fixed code tells you nothing.
Known limitations
- The judge roster's shipped defaults will not work for you until you configure it.
- Judges can read the working tree.
--diffsends untracked file contents to remote APIs. Don't point it at a repo holding secrets you haven't gitignored. - Reviewing a repository means trusting its
.opencode/,opencode.json[c],.claude/settings*.jsonhooks and.mcp.json. opencode loads plugins, tools and agent definitions from the tree it is pointed at, so a repository can ship code that a judge would run as you.llm-panelrefuses (exit 9) when the tree carries any of that.claude -pwas measured to run a tree's.claude/settings.jsonhooks with no trust prompt, so a claude judge is refused the same way when the tree declares hooks or.mcp.json;--unsafe-agentoverrides all of it. - A prompt over 128 KB is written to
<repo>/.llm-panel-material/for the run so judges' read tools can reach it; it is removed when the run ends, on any exit. On a shared host, the prompt is also visible in the judge processes' command lines while they run. - A panel is not a jury. Independent models generate candidates; verification against code, tests, and execution is still yours to do.
- Recall is measured on a 27-defect Python corpus. That number does not transfer to other languages or to defect classes the corpus doesn't contain.
- The headline recall numbers are prompt- and condition-specific. They move with
--prompt-style, roster health, and diff-vs-checkout context — see On real PRs before quoting any of them. - Every other fixture has verified scope, not proven absence. Their known unplanted
defects are recorded in each
truth.jsonand re-checked by execution invalidate_corpus.py, so a judge that finds one is not scored as wrong. Anything not yet recorded still depresses the recall floor by making a true finding look like noise.
License
MIT — see LICENSE. The MIT
grant covers the code in this repository; the benchmark data under recall/benchmarks/
contains third-party material (PR diffs and review-comment text) that stays under its
upstream terms — see
recall/benchmarks/PROVENANCE.md.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llm_panel-0.1.6.tar.gz.
File metadata
- Download URL: llm_panel-0.1.6.tar.gz
- Upload date:
- Size: 151.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.3 {"installer":{"name":"uv","version":"0.11.3","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cf60c5e52e39aec851dc4265669fdd127a44d258fe0de74bed3b38b648c07e07
|
|
| MD5 |
8310ba2de4b6f16466d5e27351a42190
|
|
| BLAKE2b-256 |
6bc07b5967893450a96ebf728aa3ab33b4a8dc80fe65758db02596b4f0b4c616
|
File details
Details for the file llm_panel-0.1.6-py3-none-any.whl.
File metadata
- Download URL: llm_panel-0.1.6-py3-none-any.whl
- Upload date:
- Size: 138.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.11.3 {"installer":{"name":"uv","version":"0.11.3","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
16a0e544931fdd90f34a96c5ccc164cb909c18609d135051cff32fecb44b0319
|
|
| MD5 |
5015fdeb0ac145990fbe8a4d49cf5a78
|
|
| BLAKE2b-256 |
79e2b104f6ec76aee5e4f31c97a9b19b4f719ebf5f7621a817c051c8ba63a7d2
|