Skip to main content

AdversaryGate (v2.0.0)

Uso alto de IA ≠ confiança alta. O custo de um pipeline com agentes de código não está na inteligência do modelo, está no autoengano do pipeline.

An evidence-based fail-closed verification gate for AI coding agents where uncertainty is a first-class result (INCONCLUSIVE) instead of a silent approval.


🎯 The Thesis

  1. Uso alto de IA ≠ Confiança alta: Quanto mais um time depende de agentes no fluxo real de engenharia, mais aparece o custo do "parece certo". Agentes geram código fluente e aparentemente correto, mas pipelines ingênuos que colapsam erros de infraestrutura aprovam patches com testes quebrados ou pulados.
  2. Troca de Modelo como Sintoma: Times trocam de modelo (Claude → GPT → Gemini) buscando credibilidade nos Pull Requests. Isso é falta de verificação determinística, não falta de modelo. O AdversaryGate executa exatamente o mesmo harness de teste sem invocar LLMs no verificador, tornando os modelos comparáveis empiricamente.
  3. Perda e Reprocessamento: O prejuízo financeiro das empresas é concreto: merged regressions, tarefas "concluídas" que não estão, rollbacks e horas de code review humano repassando o mesmo PR.
  4. Menos Autoengano do Pipeline: O produto não vende "IA mais inteligente". Vende menos autoengano no pipeline, medido numericamente pelo self_deception_index.

🔒 The Fail-Closed Model & Double-Filter Rigor

A single invariant governs the entire system:

The gate only reports what a healthy test harness actually executed. Unproduced proof is never proof of clean code.

State Outcome Meaning
Executed & Passed Outcome.VERIFIED Test ran to completion and evidence confirms clean execution.
Executed & Failed Outcome.REFUTED Test ran to completion and evidence condemns the patch.
Harness / Error Outcome.UNVERIFIED Collection error, syntax error, missing file, timeout or flaky signal. Never mergeable.

Patch Decision Matrix

  • Decision.MERGE: Requires every verdict to be VERIFIED, diff_coverage >= 80%, suite_strength >= 75%, and the full repository test suite to pass.
  • Decision.BLOCK: Triggered if any claim is REFUTED or if the full test suite fails (collateral regression).
  • Decision.INCONCLUSIVE: Triggered on UNVERIFIED outcomes, open circuit breakers, or weak test suites (suite_strength < 0.75). Never merges.

📊 Measured Pytest Exit Code Taxonomy

Exit Code Pytest Meaning ExecState Gate Behavior
0 Tests passed PASS Evaluated against baseline comparison
1 Tests failed FAIL The ONLY exit code counted as evidence
2 Collection error (import crash) UNRUNNABLE UNVERIFIED $\rightarrow$ INCONCLUSIVE
3 Internal error (harness crash) UNRUNNABLE UNVERIFIED $\rightarrow$ INCONCLUSIVE
4 Usage error (bad path / node ID) UNRUNNABLE UNVERIFIED $\rightarrow$ INCONCLUSIVE
5 No tests collected UNRUNNABLE UNVERIFIED $\rightarrow$ INCONCLUSIVE
-1 Sandbox timeout TIMED_OUT UNVERIFIED $\rightarrow$ INCONCLUSIVE

📦 Installation

Install via PyPI:

pip install adversary-gate

Or run directly from source:

python3 -m cli --help

🚀 Quick Start (CLI & GitHub Action)

Command-Line Usage

adversary-gate \
  --baseline /path/to/before \
  --patch /path/to/after \
  --test-path tests/test_auth.py \
  --test-id test_token_expiry \
  --coverage-ratio 0.85 \
  --model "${AGENT_MODEL_NAME}" \
  --evidence-log evidence.jsonl \
  --report

Exit Codes for CI Integration:

  • 0 — MERGE: Every claim executed cleanly and cleared coverage & suite strength floors.
  • 1 — BLOCK: Regressions or collateral suite failures detected.
  • 2 — INCONCLUSIVE: Infrastructure error, missing test, or weak suite (blocks merge).
  • 3 — USAGE_ERROR: Bad arguments or malformed JSON.

GitHub Action Integration (action.yml)

Add AdversaryGate to your GitHub Workflow:

name: Verification Gate

on: [pull_request]

jobs:
  verify-agent-patch:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Run AdversaryGate
        uses: adversary-gate/action@v2
        with:
          baseline: './baseline'
          patch: './patch'
          test-path: 'tests/test_token_expiry.py'
          test-id: 'test_token_expiry'
          coverage-floor: '0.80'
          suite-strength-floor: '0.75'
          model: '${{ matrix.model }}'
          evidence-log: 'evidence.jsonl'

📈 Provenance & self_deception_index

Every execution logs ctx_model and ctx_commit into the audit trail. Running patches through the same harness allows compare_models() to report verification rates side by side:

$$\text{self_deception_index} = \frac{\text{unverified_merges}}{\text{merge_count}}$$

If your product pitch is "menos autoengano no pipeline", this is the dashboard tile that proves it and the metric to watch drop to zero.


📄 License

Distributed under the MIT License.

Metadata

Release files for adversary-gate 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for adversary-gate 2.0.0
File Size Uploaded
adversary_gate-2.0.0.tar.gz 30.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for adversary-gate 2.0.0
File Interpreter ABI Platform
adversary_gate-2.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 59.6 kB

Release files / adversary_gate-2.0.0.tar.gz

Download URL adversary_gate-2.0.0.tar.gz
Size 30.4 kB
Tags Source
SHA-256 checksum
How to use checksums
4a5703d9b6510ee8e09d2b6ef84fe1f2650b1a1d37a01004ed0db5496176ed1e
BLAKE2b-256 checksum
How to use checksums
abccb9b8e65b95886cc86214f67f199c7f41b5f0bfe22f97773263ba2b63ac4b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / adversary_gate-2.0.0-py3-none-any.whl

Download URL adversary_gate-2.0.0-py3-none-any.whl
Size 29.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
deafa97ba60ad39a6d842cd42727642c5a9343fedd3df9cbc37ccd353c60a0d9
BLAKE2b-256 checksum
How to use checksums
5f491ce7753067fae6ea2f0cfc5e7156b128baa1805dc72db187db76cd81806e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

2.9.0

2 release files

2.8.0

2 release files

2.7.0

2 release files

2.6.0

2 release files

2.5.0

2 release files

2.4.0

2 release files

2.1.0

2 release files

2.0.1

2 release files

This release

2.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page