🛡️ VibeHarness-MAS: Multi-Agent Testing Harness for Vibe Coding
正體中文(臺灣) | English (US)
An AI software quality-assurance toolkit that runs white-box first (heterogeneous model cross-review + executable test verification), then black-box (STRIDE/OWASP threat-model-driven attack testing).
Designed for the failure modes of today's Vibe Coding (natural-language-prompt-driven programming): code that looks fine on the surface (the happy path passes) but lacks defensive boundaries and ships BOLA/IDOR holes and uncaught exceptions.
🏛️ Architecture & Pipeline Overview
vibeharness/core/harness.py drives the sub-agents through six linear phases (no state-machine branching or rollback). Every model call goes through vibeharness/core/llm_gateway.py, and every agent that reads target code first passes it through the untrusted-code channel in vibeharness/core/sanitize.py.
flowchart TD
CLI["main.py CLI<br/>--target / --mode / --desc / --base-url<br/>--format / --sarif / --max-files / --max-model-calls"] --> CMD["🧠 Commander<br/>vibeharness/core/harness.py (six linear phases)"]
CMD --> AST["Phase 1 · Symbol extraction<br/>ast_extractor.py (AST for Python; lexer for JS/TS)"]
AST --> SAST["Phase 1 · Deterministic AI-security rules<br/>ai_static_scanner.py (VH-AI-001..009)<br/>also detects ai_components"]
GW["LLM gateway llm_gateway.py<br/>LiteLLM routing, fallback_model, simulation fallback<br/>max_model_calls, max_concurrent_calls, call_log"]
SAN["Untrusted-code channel sanitize.py<br/>strip invisible chars / redact secrets / <untrusted_target_code><br/>every agent that reads target code goes through it"]
subgraph WB ["Phases 2–3 · White-box: multi-model review → cross-verification → sandboxed execution (model calls in parallel)"]
R1["Logic & State Reviewer<br/>Claude"]
R2["Safety & Defensive Reviewer<br/>DeepSeek"]
R3["Specification & Contract Reviewer<br/>GLM"]
R4["AI Integration & Supply-Chain Reviewer<br/>Gemini (OWASP LLM Top 10 2026)"]
DEDUP["Deduplication<br/>deduplicator.py"]
DEB["Cross-verification: votes from non-author models<br/>confirm / reject / unsure"]
GEN["White-box test generation test_generator.py<br/>invalid tests repaired up to 2 times;<br/>WB-PROBE probes when no usable test"]
SBX["Sandboxed execution (with line coverage)<br/>test_executor.py → local_subprocess / docker"]
R1 --> DEDUP
R2 --> DEDUP
R3 --> DEDUP
R4 --> DEDUP
DEDUP --> DEB
DEB -- "confirmed + ties" --> GEN --> SBX
end
AST --> SAN
SAN --> R1
SAN --> R2
SAN --> R3
SAN --> R4
SAN --> DEB
SAN --> GEN
subgraph BB ["Phases 4–5 · Black-box: threat-model-driven attack testing (model calls in parallel)"]
TM["STRIDE / OWASP API/LLM threat modeling<br/>(+ MITRE ATLAS mapping) threat_modeler.py<br/>symbols sent in batches, cross-batch dedup"]
ATK["Attack script generation<br/>scenario_builder.py"]
BEX["Black-box test execution<br/>test_executor.py → same sandbox; HTTP / in-process"]
TM --> ATK --> BEX
end
AST -- "routes & symbols" --> TM
SAST -- "ai_components" --> TM
CLI -- "--desc" --> TM
SAN --> TM
SAN --> ATK
GATE["Phase 6 · Quality gate<br/>READY_TO_SHIP / NEEDS_FIXES /<br/>INCONCLUSIVE / BLOCKED_CRITICAL_RISK"]
SAST --> GATE
SBX --> GATE
BEX --> GATE
GW -- "call_log: simulated, invalid, budget exhausted" --> GATE
GATE --> OUT["Audit report report_generator.py<br/>md / json / SARIF"]
WB -. "all model calls" .-> GW
BB -. "all model calls" .-> GW
🤖 Agent Responsibilities
Every agent's model and prompt live in vibeharness/config/agents_config.yaml and can be reassigned freely. The reviewer roster is config-driven: every whitebox_reviewer_* entry is used with its own name, model_alias, temperature, and focus (the focus text goes into its prompt), so adding or removing an entry changes the ensemble (field details in docs/configuration.en.md):
| Sub-agent (agent_id) | Default model | Responsibility | temperature |
|---|---|---|---|
Logic & State Reviewer (whitebox_reviewer_claude) |
claude | Deep logical edge cases, unhandled exceptions, state inconsistencies, null-pointer/type bugs | 0.2 |
Safety & Defensive Reviewer (whitebox_reviewer_deepseek) |
deepseek | Input sanitization, injection (SQL/command/path), defensive guards, resource exhaustion, unsafe concurrency | 0.1 |
Specification & Contract Reviewer (whitebox_reviewer_glm) |
glm | API signature compliance, business-spec adherence, contract violations, authorization checks (BOLA/BFLA) | 0.2 |
AI Integration & Supply-Chain Reviewer (whitebox_reviewer_gemini) |
gemini | OWASP Top 10 for LLM Applications 2026 (LLM01–LLM08, LLM10) in code that calls models: prompt injection, secrets in prompts/logs, unchecked tool agency, model and package supply chain, unbounded consumption, unvalidated model output; also hard-coded credentials and unpinned third-party dependencies | 0.2 |
Cross-verification votes (whitebox_verifier_* ×4) |
claude / deepseek / glm / gemini | Vote only on findings reported by other models: confirm / reject / unsure; one vote per model alias (a second verifier on the same alias is skipped) | 0.0 |
White-box test generator (whitebox_test_generator) |
deepseek | Generates executable pytest tests targeting confirmed defects | 0.3 |
STRIDE threat modeler (blackbox_threat_modeler) |
claude | Extracts trust boundaries and assets; builds the STRIDE + OWASP API (and, for AI-integrated targets, OWASP LLM Top 10 2026) attack matrix and tags MITRE ATLAS techniques | 0.2 |
Attack scenario synthesizer (blackbox_test_synthesizer) |
deepseek | Turns threat scenarios into black-box attack scripts (BOLA, injection, rate limiting, prompt injection) | 0.2 |
🚀 Key Features
-
AI Harness Commander: schedules the white-box and black-box sub-agents through six linear phases and arbitrates the final Quality Gate.
-
Heterogeneous multi-model de-biasing: the white-box review stage fans out to four LLM families with different training lineages, so their blind spots do not overlap:
- Claude: complex boundary conditions, non-null-pointer assumptions, and state-machine flaws.
- DeepSeek: input sanitization, injection, defensive guards, resource-exhaustion limits (DoS defense), and unsafe concurrency.
- GLM: API contract compliance, specification consistency, and authorization checks (BOLA/BFLA).
- Gemini: AI-integration and supply-chain security per OWASP Top 10 for LLM Applications 2026 — untrusted input concatenated into prompts (LLM01), secrets in prompts and logs (LLM02/LLM08), tools or agents acting without allowlists (LLM03), unpinned or unsafely loaded models and packages (LLM04/LLM05), uncapped tokens and spend (LLM06), model output trusted unvalidated (LLM07) or reaching
eval/exec/shell/SQL/HTML (LLM10); LLM09 (vector weaknesses) is not in its focus. Reviewers may tag findings with anowasp_mappingsuch asLLM01:2026. - Fewer than two distinct reviewer model families is an evidence gap (the verdict can never be
READY_TO_SHIP; "consensus" from one model is consistency, not corroboration; OWASP LLM07:2026). Verifiers never rule on findings from their own model, so with a single family there are no votes at all, no findings are filtered, and the report shows N/A for both the agreement ratio and the confirmed count (the findings are listed as unverified hypotheses). - Cross-model verification: no single arbiter model. Each finding is ruled on independently by every model that did not report it. Each model alias votes once, however many verifier entries point at it. If another model also reported the same defect (recorded when duplicates are merged; the merge is complete-linkage, so a generic finding cannot chain two distinct defects together), that counts as one confirm vote, but corroborations alone never count as verification: without at least one real confirm / reject from a verifier the findings stay unverified hypotheses. More confirms than rejects → confirmed; more rejects → disputed (no test is generated); a tie → kept and left to the test to decide. The agreement ratio is computed from the actual votes.
-
Mandatory threat-modeling first: no blind brute-force testing. A STRIDE (Spoofing, Tampering, Repudiation, Information disclosure, Denial of service, Elevation of privilege) and OWASP API Security Top 10 attack matrix is built first, then attack tests are generated against it. The threat modeler's inputs are
--desc(or a default description), the routes and symbols from Phase 1, and theai_componentsdetected by the deterministic rules — it does not depend on white-box results. Like the white-box reviewers, target symbols are sent in batches (split by token budget) so every symbol is visible to the threat model; threats from all batches are merged and deduplicated. -
An execution-based quality gate: model findings and threats are only hypotheses; only tests that actually ran (or the deterministic rules below) can fail or block the gate. A black-box test counts as a confirmed breach only when it fails with
AssertionError, a bareassert, orpytest.fail. Assertion-free empty tests, unusable model output (non-JSON / missing fields, or a review, verification or threat-modeling batch or an attack synthesis that crashed outside the gateway — all recorded asinvalid), untested threats, and black-box attack tests that failed without being classified as an exploit are all listed as evidence gaps. -
Deterministic AI-security rules (no model involved): before any model sees the code, AST/text rules scan the target; their results cannot be prompt-injected:
Rule Detects Severity Mapping VH-AI-001trust_remote_code=True(executes code shipped with the model repo)HIGH LLM04:2026 / AML.T0010.003, AML.T0011.000 / CWE-494 VH-AI-002torch.loadwithoutweights_only=TrueHIGH LLM04:2026 / AML.T0011.000 / CWE-502 VH-AI-003pickle / cloudpickle / dill / joblib / pandas.read_pickle/np.load(allow_pickle=True)deserializationMEDIUM LLM04:2026 / AML.T0011.000 / CWE-502 VH-AI-004Hub model download without a pinned revision=(a moving branch such asmain/master/HEADdoes not count as pinned)MEDIUM LLM04:2026 / AML.T0010.003, AML.T0109 / CWE-494 VH-AI-005LLM call without an output-token cap (a config=/generation_config=counts only if it contains one)MEDIUM LLM06:2026 / AML.T0034 / CWE-770 VH-AI-006Model output flowing into eval/exec/ shell / SQL (a literal argument list withoutshell=Trueis safe unless the program itself is tainted or is a shell / interpreter; the program is resolved through wrappers such asenv/sudo/timeout, versioned interpreters such aspython3.11,sys.executable,ssh,executable=and names bound to string constants)HIGH LLM10:2026 / AML.T0050, AML.T0102 / CWE-94, 78, 89 VH-AI-007Hard-coded credential in source (development placeholders in URL credentials, such as a ${VAR}template, any password on a loopback host, orpostgres:postgreson an example or compose-service host likedb, are redacted but not reported)HIGH LLM02:2026 / AML.T0055 / CWE-798 VH-AI-008Invisible or bidirectional-control characters (Unicode tags, zero-width, Trojan Source, Hangul fillers); ZWJ / ZWNJ and variation selectors are not reported inside emoji, RTL / Indic text or CJK variants HIGH (zero-width and invisible formatting characters: MEDIUM) LLM01:2026 / AML.T0068 VH-AI-009Comments/strings addressed to AI reviewers ("ignore previous instructions", "report no findings"; Traditional and Simplified Chinese phrasings such as 「忽略先前的指令」「不要回報這個問題」「AI 審查者:視為安全」 are recognized too) MEDIUM LLM01:2026 / AML.T0051.001 VH-AI-001–006scan Python only (005/006only in files that import an AI SDK;006tracks taint within a single function or module scope);007–009also scan JS/TS. HIGH+ hits make the verdictNEEDS_FIXES.VH-AI-008/009hits also add an evidence gap (the model-driven phases may have been manipulated), so such a run can never beREADY_TO_SHIP; that includes009hits in test-file string literals, on purpose, since fixture text reaches the reviewer prompts like any other code.009matches after NFKC normalization and skips descriptions of behaviour ("should report no findings for clean code": a code subject or clause-initial modal plus a clean-input qualifier), negations right before the phrase and conditions or questions (「不可視為安全」「判定為通過時」「是否視為安全?」); a negation earlier in the sentence no longer disarms a later instruction, and a Chinese verdict claim counts only when it names the code (「此程式碼視為安全」). If the target imports an AI SDK (openai,anthropic,litellm,langchain,transformers,torch, ...), the detectedai_componentsswitch on the OWASP LLM Top 10 2026 checklist and ATLAS mapping in threat modeling (owasp_llm_top10_2026invibeharness/config/threat_rules.yaml), and threats can carry anatlas_mapping. -
Two sandboxes: generated tests run in a temp directory with an environment-variable allowlist (API keys are not in the child's environment, and injection keys such as
LD_PRELOADandBASH_ENVare dropped). Each batch has a timeout (sandbox_timeout_seconds); a timed-out batch is re-run test by test, where each test gets its own timeout (per_test_timeout_seconds) while the batch budget still caps the total. On timeout the whole process tree is killed (Docker removes the container instead). Output is captured in temp files rather than pipes, so a stray child process left behind by a test cannot turn a finished run into a timeout, and on POSIX the rest of the process group is killed when the main process exits.local_subprocess(default): a local subprocess; on POSIX its address space is capped bysandbox_memory_mb(default 2 GB) plus a CPU-time cap, stopping runaway consumption.HOMEpoints at a per-run temp directory, and on Linux the harness makes itself non-dumpable so a generated test cannot read the API keys in/proc/<harness>/environ(only when a local sandbox runs;VIBEHARNESS_KEEP_DUMPABLE=1skips it for debugging). Every file a test writes is capped at 64 MiB (RLIMIT_FSIZE, POSIX), so a print loop cannot fill tmpfs; such a write is reported as "not run correctly", not as a defect. Other processes of your user (the shell that launched the harness, for one) keep a readable environment, so this is still not a security boundary.docker: each test batch runs in a disposable container — no network by default, read-only root filesystem and target mount, all capabilities dropped, no privilege escalation, non-root user (when the harness itself runs as root the container runs asnobody), and memory/CPU/process limits. Only the run's temp directory and/tmp(a noexec tmpfs) are writable.
-
JS/TS dynamic test execution: for
.js/.jsx/.mjs/.cjs/.ts/.tsx/.mts/.ctstargets, white-box tests and in-process black-box attack tests are generated asnode:test+node:assertES modules (node_binis detected the first time one is needed; Node >= 20, and >= 22.6 for.tsthrough type stripping). Each test module runs in its ownnode --testprocess; modules are grouped 10 at a time, each group with its ownsandbox_timeout_secondsbudget (each module runs on what is left of it; after a timeout the rest are re-run onper_test_timeout_seconds, capped at the budget left), so the worst case is ceil(modules / 10) × 2 ×sandbox_timeout_seconds. Each module is judged from the JUnit report: only anAssertionError/ERR_ASSERTIONfailure is a confirmed breach, aTypeErrorthrown inside the target is a defect but not an exploit, and a test-sideReferenceErroror an unresolvable import is a harness error. Tests import the target through itsfile://URL (the prompt states the ESM / CommonJS import form and the detected exports); bare specifiers resolve from the target'snode_modules(linked into the run directory); the static gate refuses package installs (npm/pnpm/yarn/bun/npx) as it does for Python. The Node sandbox sets noRLIMIT_AS(V8 reserves a large address space at start-up) and caps the heap with--max-old-space-sizeinstead, keeping the CPU cap;NODE_OPTIONS/NODE_PATHare never forwarded from the host. Black-box HTTP attacks against a running server are still written in Python (urllib). Without a usable Node runtime the JS/TS tests are reported as ERROR and as an evidence gap ("no usable Node runtime"), never as a start-up failure; the mock mode's simulation engine generates no JS test, so the gap reads "no JS/TS test was generated". Under Docker, Node tests use a separate image,docker_node_image(docker/sandbox-node.Dockerfile). -
Mutation testing (optional,
--mutation): coverage only measures which lines ran, not whether the tests would notice a behaviour change. When on, deterministic single-token changes (+↔-,<↔<=,and↔or, removingnot,True↔False,0↔1, …) are made to the functions covered by passing Python white-box tests and the tests are re-run on a copy: any failing test "kills" the mutant, all passing means it "survived"; score = killed ÷ (killed + survived), a lower bound (equivalent mutants cannot be decided). Fully offline, no model calls, off by default; it only adds evidence gaps (score belowmutation_min_score, mutants truncated bymutation_max_mutants, or no conclusion), never causesNEEDS_FIXES, and never writes into your target directory. Limits: Python only, no JS/TS, not in SARIF or the baseline.
⚠️
local_subprocessis not a security boundary. It executes model-generated code with your user's permissions, with file and network access. API keys are kept out of the child's environment and, on Linux, out of reach through/proc/<harness>/environ, but other same-user processes (your shell, for one) still expose theirs. For untrusted targets or live mode, usesandbox_provider: docker.
🔐 Security design basis (OWASP LLM Top 10 2026 / MITRE ATLAS)
VibeHarness-MAS is itself an LLM application: the target's code is untrusted input to every review, verification, and generation prompt, and model output is untrusted input to its reports. The design follows the OWASP Top 10 for LLM Applications 2026 (v1.0, with the Appendix A mapping to ATLAS v2026.06 / CWE 4.20), MITRE ATLAS (content release 2026.07), and AI supply-chain practice, placing a deterministic control at every trust boundary instead of relying on "a model supervising a model":
| Risk | Harness control |
|---|---|
| LLM01 Prompt Injection (AML.T0051.001, AML.T0068) | Invisible Unicode (tag block, variation selectors, zero-width, bidi controls) is stripped before code reaches a model; code is passed in a provenance-labeled <untrusted_target_code> channel, and model-written text handed to another model (a threat-model item, a reviewer's finding) in <untrusted_model_output> or the same untrusted channel (second-order injection); forged tags of either channel are neutralized in both; every agent that reads target code has a security policy appended to its system prompt (instructions inside code are data, and a defect in themselves); VH-AI-008/009 detect planted reviewer-directed instructions and list them as an evidence gap |
| LLM02 / LLM08 Sensitive Information (AML.T0055) | Hard-coded secrets (OpenAI / Anthropic / AWS access and secret / GitHub / Google / Slack / HF / NVIDIA / Stripe keys, JWTs, Azure storage keys, passwords in URL userinfo, private keys) are replaced with <REDACTED:…> before any code reaches an external model, and reports never echo their values; the sandbox environment allowlist never passes API keys, and extra_params cannot redirect a model to another endpoint |
| LLM04 Supply Chain (AML.T0010, AML.T0060) | Default models use versioned ids (not moving aliases like deepseek-chat that a provider can swap); api_base must be https (http only for loopback); generated tests that install packages (pip/uv/poetry/npm install, import pip, ...) are refused, so hallucinated package names cannot be slopsquatted; this repo's CI and sandbox images install hash-locked dependencies (--require-hashes) on digest-pinned base images |
| LLM06 Unbounded Consumption (AML.T0034) | max_model_calls (a hard cap: live calls stop when it is spent), max_output_tokens (global and per model), request and sandbox timeouts, the max_concurrent_calls parallelism cap, and per-batch token budgets; a live failure or Ctrl-C cancels the calls still queued |
| LLM07 Misinformation | Model findings and threats are hypotheses only; only executed tests or deterministic rules can fail the gate; models never judge their own findings; fewer than two distinct reviewer model families is an evidence gap |
| LLM10 Improper Output Handling (AML.T0077) | Model text has images and raw HTML disabled before it enters Markdown (no image-URL exfiltration), ANSI/OSC control sequences removed and line breaks collapsed (so it cannot open a spoofed heading or table); model-chosen ids go into code spans with pipes escaped; console output escapes rich markup; SARIF messages have control characters stripped |
| Audit trail | Every successfully sent model call records a prompt_sha256, so a verdict can be traced back to its actual input without the report having to store (possibly sensitive) code; a batch that crashes outside the gateway (review, verification, threat modeling, attack synthesis) is booked through record_failure as invalid (with the exception, no hash) |
🧭 Not yet implemented (planned)
| Feature | Status |
|---|---|
| Mutation testing extensions | v1 is Python only (core feature 8); JS/TS mutants, baseline and SARIF integration and per-operator statistics are not implemented |
| Dynamic DAST / fuzzing agent | Not implemented; black-box tests are generated per threat as pytest attack scripts |
| Hypothesis property-based testing | Not built in; models may generate it, but the target environment needs hypothesis installed |
e2b sandbox |
Not implemented; sandbox_provider supports local_subprocess and docker, anything else fails fast |
| JS/TS coverage and third-party packages | JS/TS tests are generated and run with Node (see "JS/TS dynamic test execution" below), but JS coverage is not collected; TypeScript is type-stripped only (enums, namespaces, parameter properties, decorators and tsconfig path aliases are unsupported); inside the Docker sandbox bare ESM specifiers do not resolve (only CommonJS honours NODE_PATH) |
| Scope of JSX lexing | The lexer treats < as JSX only in expression position in .js / .jsx / .mjs / .cjs / .tsx files (in .ts, <T>x is a type assertion); a < that is not valid JSX falls back to plain lexing, a known TS generic opening (<T,>, <T extends U>) is never tried, and an element nested deeper than 120 levels is lexed as plain code; failed attempts share a scan budget of 4 × the file size (at least 256 000 characters), after which every remaining < is lexed as plain code. A file the lexer cannot finish within its caps is skipped and listed as a parse failure (an evidence gap) |
| VH-AI-006 argv coverage | os.exec* / os.spawn* are not treated as sinks; env -S "bash -c …" (a split-string env) is not parsed; a name bound only to string constants is resolved file-wide, not per scope |
| Docker sandbox output | The 64 MiB per-file cap applies to local sandboxes only; the docker CLI's own output on the host is not capped (only 16 MiB per stream is read back) |
📦 Repository Layout
MultiAgentGama/
├── .github/
│ ├── CODEOWNERS # default reviewers
│ ├── dependabot.yml # weekly update proposals for Actions, requirements*.txt / requirements.lock, the sandbox base images and docker/sandbox-requirements.*
│ ├── ISSUE_TEMPLATE/ # bug report / feature request templates
│ ├── PULL_REQUEST_TEMPLATE.md
│ └── workflows/
│ ├── tests.yml # CI: lint (ruff / mypy), audit (pip-audit), package (sdist + wheel: wheel smoke test, suite run from the sdist), test (Ubuntu / Windows × Python 3.10–3.14 plus macOS × 3.12, 93% coverage floor)
│ ├── publish.yml # Release: on tag push, run tests.yml first, verify the version matches, build sdist / wheel, smoke-test the wheel in a clean venv, upload to the GitHub Release and publish to PyPI via Trusted Publishing
│ └── live-smoke.yml # Manual: one live / hybrid run of examples with the repo's provider secrets, reports uploaded
├── docker/
│ ├── sandbox.Dockerfile # Base image for the docker sandbox (digest-pinned python:3.12-slim, USER 65534; pytest / pytest-cov / requests)
│ ├── sandbox-node.Dockerfile # Base image for the docker Node sandbox (digest-pinned node:22-slim, USER 65534; runs JS/TS tests)
│ ├── sandbox-requirements.in # Direct inputs of the sandbox image's lock (constrained by requirements.lock)
│ └── sandbox-requirements.txt # Hash-locked sandbox dependencies (installed with --require-hashes)
├── vibeharness/ # the only top-level package (the sole directory a pip install adds to site-packages)
│ ├── __init__.py # package version (__version__, pyproject's single source)
│ ├── __main__.py # makes `python -m vibeharness …` the same as the vibeharness command
│ ├── py.typed # PEP 561 marker: the package ships type information
│ ├── cli.py # CLI entry point (the vibeharness command)
│ ├── config/
│ │ ├── agents_config.yaml # Global settings (execution mode, sandbox, limits), multi-model routing (Claude, GLM, DeepSeek, Gemini), per-agent prompts
│ │ └── threat_rules.yaml # STRIDE, OWASP API Top 10 and OWASP LLM Top 10 2026 (with ATLAS / CWE mapping)
│ ├── core/
│ │ ├── __init__.py
│ │ ├── harness.py # Commander: six linear phases and the quality gate
│ │ ├── llm_gateway.py # Unified LiteLLM gateway: routing, fallback, simulation fallback, call budget and call_log
│ │ ├── models.py # Pydantic data models
│ │ ├── baseline.py # --baseline: run-independent fingerprints, known problems not counted
│ │ ├── codecheck.py # Static checks on generated tests (every test_ function / method asserts; refuses package installs)
│ │ ├── codecontext.py # Source code and import guidance handed to the models (through the untrusted-code channel)
│ │ ├── estimate.py # --estimate: calls, tokens and cost without calling a model
│ │ ├── i18n.py # bilingual "中文 / English" output helpers (bi / en_only)
│ │ ├── mutation.py # Optional mutation testing (Python)
│ │ ├── sanitize.py # Trust-boundary hygiene: invisible chars, secret redaction, untrusted-code / model-output channels, output sanitizing
│ │ └── sandbox.py # Sandboxes: local subprocess (not a security boundary) and Docker containers; a Node flavour of each (node --test)
│ ├── agents/
│ │ ├── __init__.py
│ │ ├── whitebox/
│ │ │ ├── __init__.py
│ │ │ ├── ast_extractor.py # Symbol parsing and route extraction (AST for Python; dependency-free lexer for JS/TS)
│ │ │ ├── ai_static_scanner.py # Deterministic AI-security rules VH-AI-001..009 (no model involved)
│ │ │ ├── common.py # Shared helper: cancels queued model calls on a live failure or Ctrl-C
│ │ │ ├── multi_reviewer.py # Heterogeneous model ensemble review (built from the whitebox_reviewer_* entries)
│ │ │ ├── deduplicator.py # Merges the same defect reported by multiple models (complete-linkage)
│ │ │ ├── debater.py # Cross-verification: ruled on by non-author models
│ │ │ └── test_generator.py # Automatic unit/boundary test generation (self-repair, WB-PROBE probes)
│ │ └── blackbox/
│ │ ├── __init__.py
│ │ ├── threat_modeler.py # STRIDE / OWASP API / LLM threat-modeling engine (with ATLAS mapping)
│ │ └── scenario_builder.py # Turns threats into executable attack scripts
│ ├── runners/
│ │ ├── __init__.py
│ │ └── test_executor.py # Sandboxed pytest executor with PoC capture (shared by white-box and black-box)
│ ├── reports/
│ │ ├── __init__.py
│ │ └── report_generator.py # Markdown / JSON / SARIF audit report generator
├── examples/
│ └── vibe_sample_app.py # A representative sample of typical Vibe-Coding defects
├── docs/
│ ├── sample_report.md # Sample report produced in mock mode
│ ├── configuration.md # Configuration reference (Traditional Chinese)
│ ├── configuration.en.md # Configuration reference (English)
│ ├── design.md # Design document: architecture, data flow, gate verdicts, sandbox security model (Traditional Chinese)
│ └── design.en.md # Design document (English)
├── tests/ # The toolkit's own tests (offline, no API keys needed)
│ ├── conftest.py # Shared fixtures: a fake litellm so the live paths run offline
│ ├── test_ai_security.py # OWASP LLM Top 10 2026 / ATLAS controls and the deterministic AI-security rules
│ ├── test_baseline.py # --baseline: run-independent keys, known problems not counted, action.yml structure
│ ├── test_cross_verification.py # Cross-verification: models never judge their own findings, vote tallies
│ ├── test_dedup_and_budget.py # Duplicate-finding merge and prompt context within the token budget
│ ├── test_docker_sandbox.py # Docker sandbox: path mapping, hardening flags, real execution (when Docker is present)
│ ├── test_estimate.py # --estimate: exact and scenario figures, prices, no model call
│ ├── test_harness_components.py # Component integration tests in mock mode
│ ├── test_i18n.py # Bilingual helpers, report / console / CLI strings, bilingual docstrings and comments
│ ├── test_i18n_output.py # Prompts stay English, bilingual rule titles, executor reasons and PoCs
│ ├── test_js_execution.py # JS/TS execution: code gate, ESM/CJS detection, Node JUnit classification, Node sandboxes (real runs when node exists), end to end
│ ├── test_js_ts_extractor.py # JS/TS extractor: brace matching, literal/comment masking, arrow functions, routes
│ ├── test_mutation.py # Mutation testing: sites and splicing, determinism and cap, isolation and path rewriting, strong vs weak tests end to end, gate semantics, switch precedence, CLI / Action
│ ├── test_model_output_robustness.py # Graceful degradation on malformed model output, probe-fallback disclosure
│ ├── test_p0_regressions.py # Executor correctness, self-repair and result-provenance honesty regressions
│ ├── test_parallel_and_budgets.py # Parallel test generation, model-call budget, file / worker caps
│ ├── test_prompt_context.py # Live mode: prompts carry real code and import paths, gateway call robustness
│ ├── test_quality_fixes.py # Packaging, model-output robustness, execution and SARIF fix regressions
│ ├── test_r2_cli_gateway.py # Second review round: null confirmed count without verification, exit 5 on a broken install, reserved keys, call-log order, --target default, --init-config, config discovery
│ ├── test_r2_executor_sandbox.py # Target-load attribution, per-file size cap, non-dumpable switch, JS batch budgets, assertions in same-module helpers, Windows fake home
│ ├── test_r2_extractor_scanner.py # JS/TS lexer recursion and time bounds, JSX scan budget, VH-AI-006 argv program resolution, root / wildcard route matching
│ ├── test_r2_packaging.py # PyPI metadata, absolute README links, version references, hashed build backend, action inputs, CHANGELOG links
│ ├── test_r2_sanitize.py # VH-AI-009 negations / conditions / descriptions, URL credentials and placeholder exemptions, linear-time sanitizers
│ ├── test_review_blackbox.py # Threat ids, STRIDE spellings, crashed black-box batches booked
│ ├── test_review_build.py # Pinned actions, workflow inputs via env, image digests, hashed locks, sdist contents
│ ├── test_review_cache_only.py # --cache-only: hits never call litellm, a miss is an evidence gap (hybrid) or exit 4 (live), no re-ask, exit 5 without a cache
│ ├── test_review_cli_config.py # Crash exit codes, config validation, re-ask / cache key, unverified count
│ ├── test_review_dedup_votes.py # Complete-linkage dedup, one vote per verifier model
│ ├── test_review_extractor_scanner.py # JSX / TS lexing, route detection, VH-AI-004/005/006 precision
│ ├── test_review_known_limits.py # JSX text, hapi array and regex routes, nested Flask routes reviewed once, named argv lists
│ ├── test_review_reports.py # Code spans, newline spoofing, SARIF paths, threat symbols, unverified hypotheses
│ ├── test_review_sandbox_executor.py # Fake home, non-dumpable harness, stray children, NameError attribution, JS budgets
│ ├── test_review_sanitize.py # Credential formats, invisible characters, injection markers, channel tags
│ ├── test_review_sanitize_blackbox_fixes.py # Placeholder credentials, VH-AI-008/009 precision, port check, aborting calls
│ └── test_review_whitebox_fixes.py # Routes, nested Python definitions, scanner rules, severity, untrusted finding text
├── .pre-commit-config.yaml # pre-commit: ruff and mypy before every commit
├── CHANGELOG.md # Changelog (Traditional Chinese); CHANGELOG.en.md is the English version
├── CONTRIBUTING.md # Contributing guide: setup, pre-push checks, principles, release steps
├── LICENSE # MIT license
├── SECURITY.md # Security policy: private vulnerability reporting
├── MANIFEST.in # sdist contents: docs, config, tests, examples, main.py, action.yml, requirements and locks, docker/, policy files
├── action.yml # composite GitHub Action: one `uses: chinchiang/MultiAgentGama@vX.Y.Z` line runs the audit in a target repo
├── pyproject.toml # Package metadata, the vibeharness entry point, pytest / ruff / mypy settings
├── requirements.txt # Runtime dependencies
├── requirements-dev.txt # Development dependencies (ruff, mypy, pytest-cov, build, pip-audit; pinned)
├── requirements-build.in # Build backend (setuptools, the same range as pyproject's [build-system])
├── requirements-build.txt # Hash-locked build backend (the action and CI install it with --require-hashes and build without isolation)
├── requirements.lock # Fully locked, hashed versions (CI installs with --require-hashes; pip-audit)
├── README.md # Documentation (Traditional Chinese); README.en.md is the English version
└── main.py # thin launcher: `python main.py …` calls vibeharness.cli
🛠️ Getting Started
1. Install
From PyPI (registers the vibeharness command; Python 3.10+):
pip install vibeharness-mas
vibeharness --version
As an isolated command-line tool: pipx install vibeharness-mas or uv tool install vibeharness-mas. For reproducible CI, pin the version: pip install vibeharness-mas==0.7.1. python -m vibeharness … is the same as vibeharness …; an incomplete install (a missing dependency) exits 5 with reinstall instructions instead of a traceback.
From source (for development; python main.py works from the checkout as well):
git clone https://github.com/chinchiang/MultiAgentGama.git && cd MultiAgentGama
pip install -e .
For exactly the dependency set CI tests with (hash-verified): pip install --require-hashes -r requirements.lock.
2. Configure API keys (optional)
The toolkit supports real API calls and an offline simulation mode (Hybrid Mode):
For a complete reference to every field in
vibeharness/config/agents_config.yaml— defaults, model reassignment, and Docker sandbox examples — see docs/configuration.en.md (正體中文).For why each component is designed the way it is, how data flows between phases, and how the gate and evidence gaps are decided, see docs/design.en.md (design document; 正體中文).
# To use live models:
export ANTHROPIC_API_KEY="sk-ant-..." # Claude
export ZAI_API_KEY="..." # GLM (Zhipu Z.AI)
export DEEPSEEK_API_KEY="sk-..." # DeepSeek
export GEMINI_API_KEY="AIza..." # Gemini
Default models are in vibeharness/config/agents_config.yaml (Claude claude-sonnet-5-5, GLM zai/glm-5.3, DeepSeek deepseek/deepseek-v4-pro, Gemini gemini/gemini-3.8-flash). Each model can set a fallback_model (defaults: claude-haiku-4-5-20251001, zai/glm-5.3-flash, deepseek/deepseek-v4-flash, gemini/gemini-3.5-flash); after the primary model fails max_retries times the fallback is tried, and the report records which model was actually used. Each model's key is read from the environment variable named in api_key_env; providers that authenticate with AWS credentials (e.g. Bedrock) can set aws_profile_name / aws_region_name instead. Target code is sent to every configured provider (with hard-coded credentials redacted first), so only configure providers whose data-retention terms you accept.
Other limits in the global section (effective unless overridden from the CLI): execution_mode (overridden by --mode), timeout_seconds (per model request, default 180 s), max_retries (2), max_output_tokens (16384, reasoning tokens count against it; models.<alias>.max_output_tokens overrides it per model), max_concurrent_calls (parallel model calls per phase, 6), max_model_calls, max_files, max_file_bytes, include / exclude, retry_invalid_output, sandbox_timeout_seconds (120) and per_test_timeout_seconds (30). null for timeout_seconds, max_retries, max_output_tokens or retry_invalid_output means the same as leaving it out (the default; a per-model max_output_tokens: null keeps the global value).
Any agent can point at any defined model via model_alias; an agent without a model_alias resolves to unconfigured (simulated in hybrid mode, a hard failure in live mode). A single API key is enough to run live mode — for example, with only a Gemini key, change every model_alias in the agents section to gemini (note: gemini-2.5-flash is no longer available to newly issued Google API keys; use a 3.x version). With a single model family, though, verifiers never rule on their own model's findings, so there are no verification votes, no findings are filtered, and fewer than two reviewer families is an evidence gap: the best possible verdict is INCONCLUSIVE (exit code 3), so treat the result as indicative only.
Note: without API keys, hybrid mode falls back to the simulation engine so you can experience the workflow offline. The simulation returns fixed templates unrelated to the target code; the report marks them with [SIMULATED] and an "執行來源 / Execution Provenance" section, and the quality gate will never yield READY_TO_SHIP because of them.
--mode |
Behavior |
|---|---|
live |
Real models only; a missing key or failed call aborts (exit code 4). Unusable model responses are recorded as invalid and the verdict becomes INCONCLUSIVE |
hybrid (default) |
Calls live models when possible; otherwise falls back to simulation, disclosed in the report |
mock |
Simulation engine only |
In simulation mode, white-box tests use deterministic None-input robustness probes derived automatically from the target functions (WB-PROBE-*); in any mode, a confirmed finding that ends up without a usable model-generated test also gets a probe and is listed as an evidence gap. Black-box threats without an executable test are listed under "Threats Not Tested" instead of fabricating results.
3. Use the Docker sandbox (recommended for untrusted targets)
docker build -t vibeharness-sandbox:latest -f docker/sandbox.Dockerfile .
docker build -t vibeharness-sandbox-node:latest -f docker/sandbox-node.Dockerfile . # when the target has JS/TS
Set sandbox_provider: "docker" in vibeharness/config/agents_config.yaml. Tunables: docker_image, docker_network, docker_memory, docker_cpus, docker_pids_limit. JS/TS tests use the Node container docker_node_image (sandbox_node_provider follows sandbox_provider by default and can be set to local_subprocess alone to run JS tests with the local Node).
-
The target directory is mounted read-only into the container at
/work/<dir name>; host absolute paths inside generated tests are rewritten to container paths automatically. -
Containers have no network by default, so packages cannot be installed at run time. Pre-install the target's dependencies into a project-specific image and point
docker_imageat it. The base image runs asnobody(USER 65534), so switch to root for the install and back afterwards:FROM vibeharness-sandbox:latest USER root COPY requirements.txt /tmp/requirements.txt RUN pip install --no-cache-dir -r /tmp/requirements.txt USER 65534:65534
-
Both base images are pinned by digest, and the Python image installs its test tools from the hash-locked
docker/sandbox-requirements.txt. -
Black-box tests that hit a running service need network access (e.g.
docker_network: "host"on Linux), which reduces isolation. -
A missing docker command or image fails loudly with the build command — it never silently falls back to the local subprocess.
4. Run the full pipeline against a target project
python main.py --target examples --output vibe_harness_report.md
--target (-t, default ., the current directory) accepts a directory or a single file; --output (-o) defaults to vibe_harness_report.md. --desc (-d) gives the threat modeler a high-level description of the system; --config (-c) points at another config file (without it, ./vibeharness.yaml, then ./.vibeharness.yaml in the current directory is used, else the bundled default; the file used is printed on stderr), and vibeharness --init-config [PATH] writes an editable copy of the bundled default (default ./vibeharness.yaml; an existing file is kept unless --force); --verbose (-v) enables debug logging. Black-box tests import the target in-process by default; they send real HTTP requests only when you pass --base-url (-u, e.g. --base-url http://localhost:8000; only http:///https:// URLs with a host name are accepted, any other scheme is a usage error) explicitly and a service answers there. No default address is ever probed, so attack scripts cannot hit an unrelated local service, and in live HTTP mode attack code aimed at any host other than --base-url, or at another port of that host, is rejected statically. See docs/sample_report.md for example output.
⚠️ Auditing an untrusted checkout? Pass
--configexplicitly. Without it, avibeharness.yaml/.vibeharness.yamlin the current directory is loaded automatically, and a hostile repository can use it to redirectapi_base/api_key_env(sending its code, or your key, to another endpoint). The GitHub Action always passes--config(the bundled config unlessextra-argsnames one), so a workspace config is never auto-loaded there.
Before a run, --estimate shows its size: it runs Phase 1 only (symbol extraction and the deterministic rules, no model call) and prints the expected calls per agent and per model (reviewer and threat-modeling counts are exact; verification, test generation and attack synthesis are scenario figures from stated assumptions such as findings per batch, with a maximum that includes re-asks and heals), the estimated prompt / output tokens and, for models with price_per_million_* configured, the estimated cost; with --format json it is written to --output; the exit code is 0 (5 when a --changed-since ref cannot be resolved or --output cannot be written). Scale and cost limits can be overridden from the CLI: --max-files (number of source files analyzed per run, default 50, must be ≥ 1; the rest are reported as skipped and listed as an evidence gap) and --max-model-calls (cap on model invocations per run, unlimited by default; 0 also means unlimited and a negative value is a usage error; retries and the fallback are not counted separately; when exhausted, hybrid falls back to simulation and lists an evidence gap, live fails). File selection: --include GLOB / --exclude GLOB (repeatable; matched against the path relative to --target or the file name, e.g. --exclude 'tests/*' --exclude '*.min.js'; they replace the config's include / exclude), and --changed-since <git ref> (only files that changed relative to that ref, including unstaged and untracked ones; made for PR checks: --changed-since origin/main). Files above max_file_bytes (default 1 MB) are not analyzed and are listed as an evidence gap. An unusable model answer is re-asked once (retry_invalid_output) before it is booked invalid. An unwritable --output / --sarif location exits 5 before any model call; when a run crashes or is interrupted, the completed phases (call log, findings, test results, threats) are written to <output>.partial.json. --cache-dir DIR enables the response cache: usable live answers are stored in DIR and an identical prompt later is served from it without a provider call or budget spend (handy for re-runs or switching the report format); an answer that only became usable after a re-ask is also stored under the original prompt's key, so a re-run hits on the first attempt. --cache-only serves every call from the response cache only (needs --cache-dir or global.cache_dir, else exit 5): no provider call, no budget spend and no re-ask; a call the cache cannot answer is an evidence gap, simulated in hybrid (the verdict cannot be READY_TO_SHIP) and exit 4 in live. --estimate with --cache-only still lists the calls and notes that no provider is called and nothing is spent. See docs/configuration.en.md.
Report formats: --format md (default) or --format json (machine-readable, with the complete findings, threat model, and test results — suited to CI/CD; schema_version is "2", where summary.confirmed_whitebox_bugs is null when cross-verification was not applied, and older "1" reports still work as a --baseline). Both keep every generated test's code: the Markdown appendix "產生的測試程式碼(稽核軌跡) / Generated Test Code (audit trail)" lists each test's code next to its outcome in collapsible blocks (control characters removed), and the JSON generated_tests / blackbox_generated_tests keep it verbatim for audits or re-runs. Model-written tests are hypotheses too — a "passing" test may be asserting the defect itself (e.g. pytest.raises(KeyError)), so check that the assertions describe the correct behaviour:
python main.py --target examples --output report.json --format json
You can also pass --sarif <path> to emit a SARIF 2.1.0 report alongside: confirmed white-box findings carry file/line locations (resolved from line_number or the symbol table; when cross-verification was not applied they are marked properties.verified: false and their message starts with "Unverified hypothesis"), confirmed black-box breaches are error-level results (located at the attacked symbol when the attack vector names it as a whole identifier or route, the longest match winning), and deterministic AI-security rule hits use their own rules vibeharness/VH-AI-001 … vibeharness/VH-AI-009, tagged with their OWASP, MITRE ATLAS, and CWE ids. Rule ids are stable (vibeharness/whitebox-finding / vibeharness/blackbox-breach / vibeharness/VH-AI-00x) and the run-specific finding/test ids live in properties, so renumbering between runs does not close and reopen Code Scanning alerts. Upload the SARIF to GitHub Code Scanning to surface defects in the repo's Security tab:
# Example .github/workflows/security.yml for the target project: this repo ships a composite GitHub Action (action.yml)
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 0 # changed-since needs the history to diff against the ref
persist-credentials: false
- uses: chinchiang/MultiAgentGama@v0.7.1 # installs the package from the action's own checkout (hash-verified dependencies) and runs the audit
with:
target: src
mode: hybrid
max-model-calls: "80"
changed-since: origin/main # PR checks: only files that changed (optional)
baseline: .vibeharness/baseline.json # optional: only problems that are new fail the job
mutation: "true" # optional: true / false switches mutation testing for this run, empty follows the config
sandbox: docker # optional: local_subprocess (default) | docker (builds the sandbox images; needs a Docker runner)
artifact-name: vibeharness-report # optional: must be unique per run, so give each matrix cell its own
fail-on: needs-fixes # blocked | needs-fixes | inconclusive | never
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
ZAI_API_KEY: ${{ secrets.ZAI_API_KEY }}
DEEPSEEK_API_KEY: ${{ secrets.DEEPSEEK_API_KEY }}
GEMINI_API_KEY: ${{ secrets.GEMINI_API_KEY }}
- uses: github/codeql-action/upload-sarif@9f759ee644a3e7c15c1390abf49868036c00067b # v3.38.3
if: always()
with:
sarif_file: vibeharness.sarif
The action's outputs are verdict (READY_TO_SHIP / NEEDS_FIXES / BLOCKED_CRITICAL_RISK / INCONCLUSIVE, or ERROR when the harness itself failed) and exit-code; the report, the SARIF and a crash's .partial.json are uploaded as an artifact named by artifact-name (default vibeharness-report; artifact names must be unique within a workflow run, so matrix jobs or repeated uses of the action need distinct names). changed-since diffs against a git ref, so check out with fetch-depth: 0 (the default shallow clone has no history to compare). sandbox: docker builds the sandbox images from the action's own docker/*.Dockerfile and runs with a copy of the config (the bundled one, or the --config passed in extra-args) set to sandbox_provider: docker; it needs a runner with Docker such as ubuntu-latest. Without the action, pip install vibeharness-mas==0.7.1 and call vibeharness directly.
Mutation testing switch: --mutation turns it on, --no-mutation off; precedence is CLI flag > config global.mutation_testing > off, so a config that enables it can be switched off for one run with --no-mutation, and a config that disables it can be switched on with --mutation. --mutation-max-mutants N (≥ 1) overrides the mutant cap (default 20); each mutant costs one sandbox batch and no model call, and --estimate lists the ceiling when it is on. The report has a Mutation Score row in the metrics table and lists surviving mutants; the other keys are in docs/configuration.en.md.
Gating on new problems only (--baseline): keep one --format json report (for example committed as .vibeharness/baseline.json) and pass it with --baseline: failed white-box tests (keyed by their source finding), probe failures, black-box breaches and HIGH+ rule hits it already contains are still listed in the report's "基準中的已知問題(不計入) / Known from baseline (not counted)" section and in the recommendations, but no longer count toward the verdict or the exit code, so CI fails only on problems that are new. The keys do not depend on run numbering (a finding is keyed by target-relative path + symbol + normalized title, a breach by the threat's STRIDE category + title + attack vector, a rule hit by rule + file + evidence without the line number, a probe by its symbol), so re-numbered model output or unrelated line shifts never turn an old problem into a new one. Every item in the JSON report carries a fingerprint field; reports from before 0.3.0 without it are recomputed. A crash's .partial.json cannot serve as a baseline.
Note: file paths in the SARIF are relativized against the root of the git repository containing
--target(falling back to the--targetdirectory outside a repository), so--target srcstill yields Code-Scanning-aligned paths such assrc/app.py. A file outside that root gets an absolutefile://URI (nouriBaseId), and a model-reported relative path that would escape it with..gets no location;..is never emitted.
Exit codes (evaluated in order):
| Code | Meaning | Condition |
|---|---|---|
2 |
BLOCKED_CRITICAL_RISK | At least one black-box attack test confirmed a defense was breached |
1 |
NEEDS_FIXES | At least one white-box test failed (exposed a defect), or a HIGH+ deterministic AI-security rule hit; this code is exclusive to that verdict — harness errors never use it |
3 |
INCONCLUSIVE | No failures, but evidence gaps exist: simulation engine used, unusable model output or a crashed batch, max_model_calls exhausted, no reviewers or fewer than two reviewer model families, files skipped by max_files, files above max_file_bytes, non-Python code reviewed statically only, test errors / not run, confirmed findings without a model-generated test, untested threats, black-box tests that failed without being classified as an exploit, file parse failures, planted reviewer-directed instructions or invisible characters (VH-AI-008/009), etc.; an unknown verdict string also fails closed with this code |
0 |
READY_TO_SHIP | None of the above |
4 |
— | Live-mode model invocation failed, or the response cache had no answer for a live-mode --cache-only run |
5 |
— | Harness configuration or usage error: --target missing, config file absent, invalid YAML, a directory, unreadable or binary, or with a top level / section / entry that is not a mapping, an invalid max_files / max_model_calls / max_concurrent_calls / timeout or similar value, a reserved extra_params key (output-token caps max_tokens / max_completion_tokens / max_output_tokens included), --cache-only without a response cache (neither --cache-dir nor global.cache_dir), an --init-config target that already exists (without --force) or cannot be written, an incomplete install (a missing dependency or a module that fails to import), sandbox unavailable, unwritable --output / --sarif, an unresolvable --changed-since ref, CLI usage error, or any unexpected internal error, whether while starting up, during --estimate or during the run (a crash is never misreported as a verdict) |
130 |
— | Interrupted with Ctrl-C; no full report written (completed phases are in <output>.partial.json) |
5. Run the toolkit's own tests and linters
pip install -r requirements-dev.txt
python -m pytest
python -m ruff check .
python -m mypy vibeharness/ main.py
A manual live-smoke.yml (Actions tab → live-smoke → Run workflow) runs the target once against real models with the provider keys stored as repository secrets (live or hybrid, with a call cap), then re-runs for the JSON report with --cache-only --mode hybrid from the same response cache (no provider call, so the whole workflow makes at most the cap in real calls; calls the first run could not cache are simulated and listed as an evidence gap), and uploads the reports and SARIF as an artifact; it is the only check that reaches the providers and does not run on PRs. CI (.github/workflows/tests.yml) has four jobs: lint (ruff and mypy), audit (pip-audit), package (builds the wheel and runs it once in mock mode outside the checkout) and test (Ubuntu / Windows × Python 3.10–3.14 plus one macOS cell on Python 3.12, installed with --require-hashes from requirements.lock; the Linux job also builds the Docker sandbox images and runs the real-container tests). package builds the sdist and the wheel, runs the wheel once in mock mode outside the checkout and runs the full suite from the extracted sdist; publish.yml runs this whole workflow on the tagged commit before it builds a release. The test job enforces a coverage floor; to run the same check locally:
python -m pytest --cov=vibeharness --cov-fail-under=93
Supply-chain check (the CI audit job runs the same command against the locked set; a dependency with a known CVE fails before merge):
pip install "$(grep '^pip-audit==' requirements-dev.txt)"
python -m pip_audit -r requirements.lock --disable-pip --require-hashes
Optional: install pre-commit to run ruff and mypy automatically before every git commit (see CONTRIBUTING.md for the contribution workflow and SECURITY.md for vulnerability reporting):
pip install pre-commit
pre-commit install
Metadata
Release files for vibeharness-mas 0.7.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vibeharness_mas-0.7.1.tar.gz | 898.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vibeharness_mas-0.7.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.2 MB
Release files / vibeharness_mas-0.7.1.tar.gz
| Download URL | vibeharness_mas-0.7.1.tar.gz |
|---|---|
| Size | 898.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
731d7dceeeda072e1c57238594c33664a0d22ada2e89a0c9e87de0f984550b8f
|
|
BLAKE2b-256 checksum How to use checksums |
ba0e90568ac750153aee1e1dfb56526c8248c231f85ec3f5819e260cf28b08ec
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.
Transparency logRelease files / vibeharness_mas-0.7.1-py3-none-any.whl
| Download URL | vibeharness_mas-0.7.1-py3-none-any.whl |
|---|---|
| Size | 319.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
19bda50ad738b5b91ac948a54581a7194938b023f1d4f4c3fca94bcf36f491bb
|
|
BLAKE2b-256 checksum How to use checksums |
77b5fecade81e9817c159d637906aa3f1ffcbca34e864ca15446421f4374d292
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.
Transparency log