ConstraintLoop
ConstraintLoop is an evidence-based completion gate for AI coding agents. Instead of relying on a human to inspect every generated line, it requires the agent's work to pass an explicit, versioned contract of tests, static checks, metrics, artifacts, and independent model rubrics.
The central distinction is deliberate:
- Deterministic constraints produce reproducible evidence: exit codes, parsed metrics, and validated artifacts. Required deterministic failures block autonomous completion. A human may explicitly waive an exact local evidence snapshot with a non-empty reason; CI ignores every waiver and remains blocking.
- Non-deterministic constraints apply a written rubric through OpenAI, Anthropic, or any command that speaks ConstraintLoop's JSON protocol. They are advisory by default. A required rubric must run at least twice and declare a majority quorum.
ConstraintLoop supports Claude Code, Codex, and Gemini CLI through their hook lifecycles. CI is the final authority: it ignores local caches and human waivers.
Bounded convergence loops are implemented. The design keeps ConstraintLoop in control of evidence, budgets, and stopping while native Claude or Codex loops perform at most one requested repair per transition. See docs/convergence-loops.md.
Completion loops can also require the current coding session to generate and investigate N domain-grounded failure scenarios before stopping. Claude Code, Codex, and Gemini CLI use their existing session context and tools; this challenge gate requires no separate model evaluator. See session challenge gates.
At a glance
| Question | ConstraintLoop answer |
|---|---|
| What decides that work is complete? | Fresh evidence from a committed contract |
| What can block locally? | Required deterministic failures and undisposed advisory findings |
| What can block CI? | Every required CI constraint; local caches and waivers are ignored |
| Does it replace pytest, Ruff, or CI? | No. It turns their outputs into one completion decision |
| Does it run an autonomous agent? | No. It owns evidence and stopping; native agents own repairs |
| Can it review design? | Yes, through optional OpenAI, Anthropic, Codex, Claude Code, or command evaluators |
| Can it loop forever? | No. Every convergence loop has repair, unchanged-result, and time budgets |
flowchart LR
G[User goal] --> A[Coding agent]
A --> C[Versioned contract]
C --> D[Commands and metrics]
C --> R[Optional rubric review]
D --> E[Fresh evidence snapshot]
R --> E
E -->|pass| S[Completion allowed]
E -->|fail| F[One focused repair]
E -->|pending| W[Wait without repair]
F --> A
W --> E
E -->|budget reached| H[Human decision]
Choose your path
| I want to… | Start here |
|---|---|
| Add tests, coverage, and lint gates | Quick start and task-oriented recipes |
| Understand when each gate runs | Lifecycle |
| Configure every schema field | Configuration reference |
| Use Codex or Claude Code for design review | Native CLI evaluators |
| Use OpenAI or Anthropic directly | Provider privacy |
| Add a bounded repair or monitoring loop | Convergence loops |
| Diagnose a failure or stale cache | FAQ and troubleshooting |
| Evaluate the security boundary | Threat model |
Quick start
uv tool install constraintloop
# Or: pipx install constraintloop
constraintloop init
constraintloop setup --adapter all
# Or add --pre-push above to wire heavyweight push gates locally.
constraintloop run
constraintloop ci
constraintloop init detects existing Python and Node tooling and writes a
plain constraintloop.yml. It does not install tools or silently invent gates.
Review and commit the contract.
Install ConstraintLoop as an isolated tool instead of adding it to the target
project's environment. This avoids dependency conflicts with the application.
If you intentionally run setup through uvx, generated hooks pin the current
ConstraintLoop version. You can choose another persistent invocation with, for
example, constraintloop setup --hook-executable "pipx run constraintloop".
The five core commands above establish this flow:
sequenceDiagram
participant U as User
participant A as Agent
participant CL as ConstraintLoop
participant T as Project tools
U->>CL: init + review contract
U->>CL: setup hooks
A->>CL: run change/stop/push phase
CL->>T: execute ready constraints
T-->>CL: exit codes, metrics, artifacts
CL-->>A: pass, repair, wait, or escalate
CL->>T: ci reruns without cache/waivers
Contract
version: 1
settings:
max_auto_retries: 2
hook_output_limit: 4096
constraints:
tests:
kind: command
command: [python, -m, pytest, -q]
phases: [stop, push, ci]
watch: ["src/**/*.py", "tests/**/*.py", pyproject.toml]
retry:
max_attempts: 3
exit_codes: [1]
delay_seconds: 2
integration_tests:
kind: command
command: [python, -m, pytest, -q, -m, integration]
phases: [push, ci]
watch: ["src/**/*.py", "tests/**/*.py", pyproject.toml]
needs: [tests]
timeout_seconds: 900
coverage:
kind: metric
command: [python, -m, pytest, --cov, "--cov-report=json:coverage.json"]
parser:
type: json
source: file
file: coverage.json
path: totals.percent_covered
threshold: {operator: gte, value: 85}
needs: [tests]
phases: [stop, ci]
database_consumers:
kind: ratchet
description: Do not add legacy database consumers during migration
command: [python, scripts/inventory_consumers.py, --json]
parser: {type: json, path: counts.database_consumers}
mode: must_not_increase
phases: [stop, ci]
design_review:
kind: rubric
enforcement: advisory
evaluator: independent_review
rubric: >
Fail when the patch introduces an unjustified public API, crosses an
existing architectural boundary, or omits handling for a named failure
case. Cite concrete files in every finding.
include: ["src/**/*.py"]
phases: [stop, ci]
evaluators:
independent_review:
type: openai
model: YOUR_PINNED_MODEL
See examples/constraintloop.full.yml for all constraint types.
Ratchets store their numeric baseline and evidence SHA-256 in the committed
constraintloop-baselines.json; baseline changes are therefore explicit code
review events. JSON artifact constraints can map selected dotted paths through
evidence so human-readable runs and constraintloop status report useful
counts and changes while the complete machine-readable run remains available
through --json. The published
JSON Schema provides editor validation and
autocomplete for every contract field.
The pre-release engineering and open-source checklist is tracked in docs/release-readiness.md. Participation is governed by the Code of Conduct. Maintainer release setup and Trusted Publishing invariants are documented in RELEASE.md. The strict schema is documented in docs/configuration.md, and remote evaluator disclosure and cost controls are documented in docs/provider-privacy.md. OpenAI request-shape, failure, SDK-compatibility, and semantic-corpus checks are documented in docs/openai-evaluation.md. Optional local Codex and Claude Code command evaluators are documented in docs/native-cli-evaluators.md.
OpenAI evaluator setup
Install the optional provider SDK:
uv sync --extra dev --extra openai
For local development, paste the key into the gitignored
.constraintloop/secrets.env file:
OPENAI_API_KEY=your-key-here
Process environment variables take precedence over the local file. In CI, use
the CI platform's secret store and expose OPENAI_API_KEY; do not create or
commit a credential file. ConstraintLoop parses the local file as plain
KEY=VALUE data and never evaluates it as shell code. Agent hook writes to this
file are denied.
This repository dogfoods an advisory native-agent design rubric. OpenAI and Anthropic remain optional provider integrations. Keep probabilistic gates advisory until their false-positive and false-negative rates are measured.
Lifecycle
| Phase | Typical trigger | Intended work |
|---|---|---|
change |
After a file-changing tool action | Fast syntax, formatting, or diff checks |
stop |
When the agent attempts to finish | Tests, build checks, and advisory review |
push |
Explicit local run or opt-in Git pre-push hook | Full integration and platform suites |
ci |
Protected hosted workflow | Authoritative uncached and waiver-free verification |
SessionStarttells the coding agent which required gates exist.- The prompt hook records the user's goal as review evidence.
- Before tool execution, agent attempts to edit the contract or create a waiver are denied.
- After a main-agent tool execution,
changegates run and fresh results are injected. Subagent tool and stop events are ignored because their working tree may be intentionally transient. - Before compaction, the completion policy is restated.
- At
Stop/AfterAgent, requiredstopgates block completion. The agent receives compact evidence and may repair the code a bounded number of times. ConstraintLoop defers this evaluation while background tasks or scheduled wakeups are active. Full retained output remains available throughconstraintloop debug ID. - Advisory failures require either passing fresh evidence or an explicit snapshot-bound explanation; delivery alone never counts as review.
- Repeated required failure stops autonomous repair and requests a human decision. A trusted human can record a reasoned, snapshot-bound local waiver; hooks deny observed agent waiver commands, any relevant change invalidates it, and CI ignores it. The local CLI cannot authenticate whether its caller is human.
constraintloop cireruns every CI gate without local evidence or waivers.
Every command constraint has a finite timeout (300 seconds by default). On POSIX,
timeout cleanup terminates the entire spawned process group so TestContainers,
Docker clients, and other descendants cannot keep inherited output pipes open.
Commands run from their configured project-contained cwd, with the selected
project root prepended to PYTHONPATH.
Evidence is keyed by the constraint definition and the bytes of every file
matched by watch. A source change therefore makes old evidence and waivers
stale without a mutable invalidation list. Local state lives under the
gitignored .constraintloop/state directory; set CONSTRAINTLOOP_CACHE_DIR to
override it.
For stronger machine-local gates, create a gitignored
constraintloop.local.yml. ConstraintLoop recursively merges mappings over the
repository contract and rejects changes that could weaken committed gates.
The authoritative constraintloop ci command ignores this overlay.
init and setup add local state, the uninstall tombstone, generated agent hook
settings, and overlay names to the selected project's .gitignore, and warn if
state is already tracked. Claude uses its dedicated
.claude/settings.local.json path. Explicit uninstall records a local tombstone,
so a checkout that restores old committed hook wiring does not silently
reactivate ConstraintLoop; setup clears the tombstone.
Verdicts and what they mean
| Verdict | Meaning | Can complete? |
|---|---|---|
pass |
Fresh evidence satisfies the constraint | Yes |
fail |
The tool or rubric found a concrete violation | No when required |
pending |
External or delayed evidence is not ready | No |
uncertain |
An evaluator could not produce a reliable verdict | No when required |
error |
ConstraintLoop could not evaluate safely | No |
waived |
A human accepted one exact local deterministic snapshot | Locally only; never in CI |
skipped |
A dependency prevented execution | Only when no required result is missing |
Commands
constraintloop init— generate a reviewable initial contract.constraintloop setup --adapter claude|codex|gemini|all [--pre-push]— merge agent hook entries while preserving existing hooks; optionally install an owned Git pre-push hook forpushgates.constraintloop uninstall --adapter claude|codex|gemini|all— remove only ConstraintLoop hook entries while preserving unrelated settings; pass--pre-pushto remove an owned Git hook too.constraintloop run --phase change|stop|push— run local gates with fresh caching. Push gates do not honor local waivers.constraintloop ci— authoritative, uncached, waiver-free run.constraintloop cycle NAME --json— execute one journaled loop transition.constraintloop supervise NAME— poll pending evidence under a recoverable single-writer lease and exit whenever repair or termination is required.constraintloop loop-prompt NAME --adapter claude|codex|gemini— print the bounded native-agent repair and challenge protocol without launching an agent.constraintloop challenge show NAME— inspect saved scenarios and the submission schema.constraintloop challenge submit NAME --file PATH— record session-authored discovery or verification for the current request and input snapshot.constraintloop status— inspect evidence without executing commands.constraintloop explain --phase change|stop|push|ci— show why each constraint runs or is skipped, including matched and changed watch paths, cache state, and dependency chains.constraintloop baseline update ID|--all— initialize or strengthen native ratchet baselines; weakening requires the explicit--allow-regressionflag.constraintloop debug ID— explain evidence freshness, evaluator configuration, executable resolution, and native CLI availability without running an evaluator or consuming model quota.constraintloop acknowledge ID --reason "..."— record an explicit snapshot-bound advisory disposition without changing its verdict.constraintloop doctor— validate and fingerprint the contract;--deepalso checks executables, Python invocations, referenced environment files and variables, worktree environment templates, virtual environments, container runtimes and daemons, ratchet baselines, empty watch globs, and local-state hygiene.constraintloop waive ID --reason "..."— human-local, snapshot-bound waiver for fresh non-passing deterministic evidence. Rubrics cannot be waived.constraintloop enhance— write a review-only proposal for stronger tooling.constraintloop author— write a review-only QA/test-authoring proposal.
enhance and author intentionally do not install dependencies or modify the
active contract. Their proposal files make the future self-improvement
path auditable.
Documentation
| Guide | Contents |
|---|---|
| Recipes | Copyable Python, native-review, CI, and bounded-loop setups |
| FAQ | Caching, failure modes, providers, hooks, security, and troubleshooting |
| Configuration | Strict schema, defaults, constraints, evaluators, and loops |
| Convergence loops | State machine, budgets, leases, and native-agent protocol |
| Native evaluators | Codex and Claude Code read-only rubric execution |
| Provider privacy | Data flow, disclosure, credentials, cost, and failure behavior |
| Threat model | Trusted inputs, controls, residual risks, and non-goals |
| Release readiness | Compatibility, quality, security, and publishing gates |
Evaluator command protocol
A command evaluator receives an EvaluationBundle JSON object on stdin and must
write exactly one object to stdout:
{
"verdict": "pass",
"score": 0.91,
"rationale": "The patch satisfies the rubric.",
"findings": []
}
Valid verdicts are pass, fail, and uncertain. Provider errors and malformed
responses become uncertain; a required rubric therefore fails closed.
Compatibility boundary
The supported v0.4 surfaces are the CLI and exit codes, configuration schema, evaluator command protocol, native hook responses, and schema-versioned evidence and cycle JSON. Python submodules are internal during initial development and are not covered by semantic-versioning compatibility promises. Migration notes will accompany changes to supported schemas and protocols.
Security model
Hooks are policy automation, not a security sandbox. A sufficiently privileged agent process can bypass local hooks or alter local files. The trusted boundary is a protected, reviewed contract plus an independent CI run. See docs/threat-model.md.
Frequently asked questions
Why not just tell the agent to run tests? Because a prompt is not durable policy. ConstraintLoop records which contract ran, which inputs it covered, and whether the evidence is still fresh.
Why do some constraints run after every action? Put only fast feedback in
the change phase. Keep unit checks in stop; put heavyweight integration
suites in push and ci.
Can I use Codex or Claude Code instead of an API evaluator? Yes. The native evaluator adapter prefers the active supported CLI and remains read-only.
How do I test failure behavior? Use deterministic commands or fixtures that return known failure, pending, malformed, timeout, or corruption outcomes. Do not spend provider quota merely to manufacture an error.
See the complete FAQ and troubleshooting guide.
License: MIT.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file constraintloop-0.5.0.tar.gz.
File metadata
- Download URL: constraintloop-0.5.0.tar.gz
- Upload date:
- Size: 142.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8bfeee7fd729579ae1b8f8cefd29a66ff4950a569185b807f46e8e01e5f90c98
|
|
| MD5 |
37916384c099e27b38d94147aa216154
|
|
| BLAKE2b-256 |
626a9db9e1721e2348831d1c66b731fa4712929599e3b92c049b8bb8337738dd
|
Provenance
The following attestation bundles were made for constraintloop-0.5.0.tar.gz:
Publisher:
publish.yml on mauhpr/constraintloop
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
constraintloop-0.5.0.tar.gz -
Subject digest:
8bfeee7fd729579ae1b8f8cefd29a66ff4950a569185b807f46e8e01e5f90c98 - Sigstore transparency entry: 2742106117
- Sigstore integration time:
-
Permalink:
mauhpr/constraintloop@906a31de4a456c1394524d811c8493f34f7313cc -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/mauhpr
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@906a31de4a456c1394524d811c8493f34f7313cc -
Trigger Event:
release
-
Statement type:
File details
Details for the file constraintloop-0.5.0-py3-none-any.whl.
File metadata
- Download URL: constraintloop-0.5.0-py3-none-any.whl
- Upload date:
- Size: 75.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8cb9ffafd895e630ecb7d710e5f57464ea276b1f7ed5156f1f61e20918536412
|
|
| MD5 |
367642a18af4b457afd0db4656e82660
|
|
| BLAKE2b-256 |
be33cbd31fb09011e471b7560a08c7dedbb66dcc2050169c05b45d8e57cd2f5f
|
Provenance
The following attestation bundles were made for constraintloop-0.5.0-py3-none-any.whl:
Publisher:
publish.yml on mauhpr/constraintloop
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
constraintloop-0.5.0-py3-none-any.whl -
Subject digest:
8cb9ffafd895e630ecb7d710e5f57464ea276b1f7ed5156f1f61e20918536412 - Sigstore transparency entry: 2742106152
- Sigstore integration time:
-
Permalink:
mauhpr/constraintloop@906a31de4a456c1394524d811c8493f34f7313cc -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/mauhpr
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@906a31de4a456c1394524d811c8493f34f7313cc -
Trigger Event:
release
-
Statement type: