grounded-judge-gate
A small offline example of one rule:
An LLM judge may suggest a value, but deterministic code decides whether that value is accepted.
The project uses short numeric answers because the full decision path fits in a few files and can be tested without a live model.
The problem
Suppose the canonical answer is 3/4.
| Answer | What happens | Result |
|---|---|---|
3/4 |
deterministic check matches the canonical value | accept |
0.75 |
normalization turns it into 3/4 |
accept |
6/8 |
normalization turns it into 3/4 |
accept |
0.7 |
the value is valid but different | reject |
three quarters |
the judge extracts 3/4, then code checks it again |
accept |
about 0.7 |
the judge extracts 0.7, which fails the second check |
manual review |
I do not know |
no safe value can be confirmed | manual review |
The important part is not the prompt or the model. It is who has authority to decide.
The rule
The system has three roles:
- Deterministic authority handles every answer it can parse.
- The LLM judge is called only when authority cannot parse the answer.
- Manual review receives everything that cannot be decided safely.
Authority runs twice:
- before the judge, so standard forms do not need an LLM call;
- after the judge, so an extracted value cannot bypass the original rule.
An unambiguous deterministic reject is final. The judge cannot overturn it.
Decision flow
answer
|
v
normalize and compare
|
+ matches canonical value: accept
+ parsed but does not match: reject
+ cannot parse: ask the recorded judge
|
v
judge result
|
+ does not confirm equivalence: manual review
+ extracts a candidate value: compare again
|
v
second deterministic check
+ matches: accept
+ does not match: manual review
The judge is intentionally placed between two calls to the same deterministic check.
Quickstart
Requirements: Python 3.12 and uv.
git clone https://github.com/ErnisBadmae/grounded-judge-gate.git
cd grounded-judge-gate
uv sync
uv run judge-gate run scenarios/short_answer.yaml --report report.md
Expected result:
scenario: short_answer_probability cases: 15
route=authority: 9 route=judge: 2 manual_review: 4
PASS 15/15 (verdict+route+grounding matched contract)
The command returns exit code 0 when every case matches its contract and 1 when any
verdict, route, or grounded value differs.
No network call is made during the run. The judge responses come from
scenarios/fixtures/judge_responses.json.
The scenario contract
Each case declares three things:
- the final verdict;
- the route that produced it;
- the value extracted by the judge, when that route was used.
- id: verbal-form
answer: 'three quarters'
expect: {verdict: accept, route: judge, grounded: '3/4'}
- id: illegal-rescue-trap
answer: 'about 0.7'
expect: {verdict: needs_manual_review, route: judge, grounded: '0.7'}
Checking the route matters. A correct final verdict reached through an illegal path is still a defect.
The fixture contains adversarial cases where the recorded judge claims equivalence and extracts the wrong value. These cases verify that the second deterministic check cannot be skipped.
Coverage floors (v0.2)
Per-case expectations cannot express one thing: whether the machinery under test ran at
all. If the judge is unreachable, every case degrades into a legal needs_manual_review
row. Nothing in the per-case contract is violated, and the run reports green while
proving nothing.
expect_totals:
route_min:
authority: 2 # decided without the judge
judge: 2 # the judge branch actually executed
reason_max:
judge_error: 0 # an unreachable judge is a broken run, not a full queue
timeout: 0
route_min is a floor on how many cases took each branch; reason_max is a ceiling on
failure reasons. The two catch different things: a floor catches a branch that never ran,
a ceiling catches a branch that ran and failed on every case. A floor alone does not
prove the branch worked — cases that entered the branch and crashed still satisfy it.
judge_error and timeout are capped at 0 by default, whether or not a scenario
declares expect_totals. Making that opt-in was a bug: a scenario that omitted the block
could run with a completely unreachable judge and report green. Raise a ceiling
explicitly (reason_max: {judge_error: 3}) when a scenario is deliberately testing
failure handling.
Unknown route or reason names are rejected at load time, because a typo'd floor is a constraint that can never fire — the same silent pass this block exists to prevent.
The run below satisfies every per-case contract and still exits 1:
scenario: judge_branch_never_ran cases: 3
route=authority: 3 route=judge: 0 manual_review: 0
PASS cases 3/3 (verdict+route+grounding matched contract)
FAIL coverage 1 aggregate violation(s)
route 'judge' taken by 0 case(s), contract requires at least 2: that branch was not exercised
Calibration
uv run python scripts/calibrate.py
The script prints two separate results:
- Agreement and Cohen's kappa for a small synthetic judge-vs-label dataset.
- Route coverage, manual-review rate, and grounding checks for the three-route gate.
These are different measurements and are not combined into one score. If kappa is not
defined because both sides used only one class, the script prints N/A instead of a
misleading perfect score.
What this project demonstrates
- deterministic work happens before probabilistic work;
- an LLM can help with forms that normal code cannot parse;
- an LLM claim is not enough to accept a result;
- the decision path can be part of the test contract;
- a gate must also assert that the path it tests was taken at all;
- uncertainty can be routed to a person instead of being hidden.
Output validators, retries, and human review are established engineering patterns. This repository does not claim to invent them. It demonstrates one strict authority hierarchy in a form that is small, offline, and easy to inspect.
What this project does not prove
The example has important limits:
- The recorded adapter is not a real LLM. It only replays fixed responses.
- The gold set has 15 synthetic cases from one numeric domain.
- Passing the second check does not prove faithful extraction. A model could fabricate
3/4from an unrelated answer, and the authority would see only a matching value. Coverage floors do not touch this: they prove a branch ran, not that it was honest. - A live judge would need its own calibration and adversarial tests.
- This is not a production library, an agent framework, or an evaluation platform.
The project proves that its routing contract works on the included cases. It does not prove that an LLM judge is safe in general.
Project layout
src/judge_gate/ core authority, judge, runner, contract, report, CLI
scenarios/short_answer.yaml 15-case scenario
scenarios/fixtures/ recorded judge responses
data/calibration_set.json synthetic calibration data
scripts/calibrate.py agreement and route metrics
tests/ deterministic regression tests
Development
uv sync
uv run pytest -q
Provenance and license
This is a clean-room implementation. No code, data, tests, or internal structure were copied from a private or employer-owned repository. All examples and calibration data are synthetic.
MIT license. See LICENSE.
Русская версия
Это маленький офлайн-пример одного правила:
LLM-судья может предложить значение, но решение о принятии остаётся за обычным кодом.
Для демонстрации используются короткие числовые ответы. На таком примере весь маршрут решения помещается в несколько файлов и проверяется без живой модели.
В чём проблема
Пусть правильный ответ - 3/4.
| Ответ | Что происходит | Результат |
|---|---|---|
3/4 |
детерминированная проверка видит точное совпадение | принять |
0.75 |
нормализация приводит значение к 3/4 |
принять |
6/8 |
нормализация приводит значение к 3/4 |
принять |
0.7 |
число корректно записано, но не совпадает с каноном | отклонить |
три четверти |
модель извлекает 3/4, затем код проверяет значение ещё раз |
принять |
примерно 0.7 |
модель извлекает 0.7, повторная проверка не проходит |
ручная проверка |
не знаю |
безопасно подтвердить значение нельзя | ручная проверка |
Главный вопрос здесь не в модели и не в промпте. Важно заранее решить, кто имеет право принимать окончательное решение.
Как распределены полномочия
В системе три роли:
- Детерминированный арбитр обрабатывает все ответы, которые умеет разобрать.
- LLM-судья вызывается только для неразбираемой текстовой формы.
- Человек получает всё, что система не смогла решить безопасно.
Арбитр работает два раза:
- до модели, чтобы не тратить вызов LLM на обычные дроби, проценты и десятичные числа;
- после модели, чтобы извлечённое значение не могло обойти исходное правило.
Если арбитр разобрал ответ и увидел несовпадение, это окончательный отказ. Модель не может его отменить.
Схема
ответ
|
v
нормализация и сравнение
|
+ совпало с каноном: принять
+ разобрано, но не совпало: отклонить
+ разобрать не удалось: спросить записанного судью
|
v
ответ судьи
|
+ эквивалентность не подтверждена: ручная проверка
+ извлечено значение: проверить ещё раз
|
v
повторная детерминированная проверка
+ совпало: принять
+ не совпало: ручная проверка
Модель намеренно зажата между двумя вызовами одного и того же арбитра.
Быстрый запуск
Нужны Python 3.12 и uv.
git clone https://github.com/ErnisBadmae/grounded-judge-gate.git
cd grounded-judge-gate
uv sync
uv run judge-gate run scenarios/short_answer.yaml --report report.md
Ожидаемый результат:
scenario: short_answer_probability cases: 15
route=authority: 9 route=judge: 2 manual_review: 4
PASS 15/15 (verdict+route+grounding matched contract)
Код возврата 0 означает полное совпадение с контрактом. Любое расхождение по
вердикту, маршруту или извлечённому значению даёт код 1.
Во время прогона сеть не используется. Ответы судьи записаны заранее в
scenarios/fixtures/judge_responses.json.
Что проверяет контракт
Для каждого случая зафиксированы:
- итоговый вердикт;
- маршрут решения;
- значение, извлечённое моделью, если использовался маршрут судьи.
- id: verbal-form
answer: 'три четверти'
expect: {verdict: accept, route: judge, grounded: '3/4'}
- id: illegal-rescue-trap
answer: 'примерно 0.7'
expect: {verdict: needs_manual_review, route: judge, grounded: '0.7'}
Проверка маршрута нужна потому, что правильный итог можно получить неправильным путём. Для такого гейта это тоже дефект.
В фикстуре есть враждебные ответы. Записанный судья уверенно заявляет эквивалентность, но извлекает неправильное значение. Эти случаи проверяют, что повторный вызов арбитра нельзя незаметно пропустить.
Нижние границы покрытия (v0.2)
Покейсовые ожидания не умеют выразить одну вещь: отработала ли вообще проверяемая
машинерия. Если судья недоступен, каждый случай честно вырождается в законный
needs_manual_review. Ни одно покейсовое ожидание не нарушено, гейт зелёный, доказано
ничего.
expect_totals:
route_min:
authority: 2 # решено без судьи
judge: 2 # ветка судьи реально выполнилась
reason_max:
judge_error: 0 # недоступный судья - это сломанный прогон, а не полная очередь
timeout: 0
route_min - пол по числу случаев, прошедших через ветку; reason_max - потолок по
причинам отказа. Ловят они разное: пол ловит ветку, которая не выполнялась вообще,
потолок - ветку, которая выполнилась и упала на всех случаях. Неизвестные имена маршрутов
и причин отвергаются при загрузке: опечатка в нижней границе даёт ограничение, которое
никогда не сработает, то есть ровно тот тихий зелёный, против которого блок и написан.
Прогон ниже удовлетворяет всем покейсовым контрактам и всё равно возвращает 1:
scenario: judge_branch_never_ran cases: 3
route=authority: 3 route=judge: 0 manual_review: 0
PASS cases 3/3 (verdict+route+grounding matched contract)
FAIL coverage 1 aggregate violation(s)
route 'judge' taken by 0 case(s), contract requires at least 2: that branch was not exercised
Калибровка
uv run python scripts/calibrate.py
Скрипт выводит два независимых результата:
- Agreement и каппу Коэна на небольшом синтетическом наборе judge-vs-label.
- Покрытие маршрутов, долю ручной проверки и структурные нарушения grounding.
Эти измерения нельзя объединять в одну метрику. Если каппа математически не определена,
скрипт выводит N/A, а не создаёт видимость идеального согласия.
Что здесь полезного
- детерминированная работа выполняется раньше вероятностной;
- модель помогает только с формой, которую обычный код не разобрал;
- уверенного заявления модели недостаточно для принятия ответа;
- в тестовый контракт входит не только итог, но и путь решения;
- неопределённость становится обычным маршрутом к человеку.
Валидаторы, повторные запросы и ручная проверка существовали задолго до этого проекта. Репозиторий не претендует на изобретение нового класса систем. Он показывает одну жёсткую иерархию полномочий в маленьком, полностью воспроизводимом примере.
Чего пример не доказывает
- Записанный адаптер не является настоящей LLM. Он только воспроизводит готовые ответы.
- В наборе всего 15 синтетических случаев из одного числового домена.
- Повторная проверка не доказывает честность извлечения. Модель может выдумать
3/4из постороннего текста, а арбитр увидит только совпавшее значение. - Живой судья потребует отдельной калибровки и враждебных тестов.
- Это не промышленная библиотека, не агентный фреймворк и не платформа оценки.
Проект доказывает работу маршрутизации на включённых сценариях. Он не доказывает безопасность LLM-судьи вообще.
Структура проекта
src/judge_gate/ арбитр, судья, раннер, контракт, отчёт и CLI
scenarios/short_answer.yaml сценарий из 15 случаев
scenarios/fixtures/ записанные ответы судьи
data/calibration_set.json синтетические данные калибровки
scripts/calibrate.py метрики согласия и маршрутов
tests/ детерминированные регрессионные тесты
Разработка
uv sync
uv run pytest -q
Происхождение и лицензия
Это clean-room реализация. Код, данные, тесты и внутренняя структура не копировались из приватных репозиториев или проектов работодателя. Все примеры и данные калибровки синтетические.
Лицензия MIT. См. LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file grounded_judge_gate-0.3.0.tar.gz.
File metadata
- Download URL: grounded_judge_gate-0.3.0.tar.gz
- Upload date:
- Size: 29.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
609e50d16b83db489245201a62453821d585756e62be8ed1ed4b630cbaf8e8c2
|
|
| MD5 |
e75859f859341cad9eb5d623e3aa27f2
|
|
| BLAKE2b-256 |
2ac8ea6dc0d93cbfd0a87a25be6a6021d2a97160d4524e7ba563e3cf39c45eed
|
File details
Details for the file grounded_judge_gate-0.3.0-py3-none-any.whl.
File metadata
- Download URL: grounded_judge_gate-0.3.0-py3-none-any.whl
- Upload date:
- Size: 23.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.14.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
af0a19917b3825633a499df7db0fd19b2b69518315f8f9aa7c62521af982bd80
|
|
| MD5 |
86ffe689ed0515d52f5a60788f9db0c4
|
|
| BLAKE2b-256 |
3f7b0c64da3a36063c25122ba38f2bb7338b98ade88e0d07072611f69f3666f3
|