Provael™
Red-team open Vision-Language-Action (VLA) robot policies in simulation and get an Attack Success Rate — beside the control it is read against.
Deterministic CPU stub run, seed 0 — reproduce it in seconds.
[](https://github.com/provael/provael/actions/workflows/ci.yml) [](https://pypi.org/project/provael/) [](https://github.com/provael/provael/blob/main/LICENSE) [](https://doi.org/10.5281/zenodo.21984184) [](https://github.com/provael/provael/blob/main/watch/freshness.json) [](https://scorecard.dev/viewer/?uri=github.com/provael/provael)What this is. A Python CLI that red-teams an open VLA robot policy inside a simulator and
reports an attack-success rate — always beside the benign control it is read against, with a
95% interval, a competence control, and an evidence label that says whether a real policy or the
CPU fixture produced it. It emits what a review needs: report.json, a scorecard, SARIF, OSCAL, a
CycloneDX ML-BOM, a test report and a signed attestation. For a team shipping a VLA policy that
needs a measured rate with its control before a release or in CI, a researcher reproducing a
published result, a reviewer who wants the number, the denominator and the decision in one place.
Simulation only and defensive: no physical robot, no real-world-harm payload; every number here
is a claim about the simulator that produced it. Read SAFETY.md before anything else.
Run it
pip install provael # requires Python 3.12+
# deterministic CPU run — no GPU, no model download; prints an ASR-by-attack table (47/70)
provael attack --policy stub --suite stub --attacks instruction,visual,injection --episodes 10 --seed 0
It prints Adversarial ASR: 67.1% (47/70). Those are fixture numbers: the stub policy and
suite are deterministic CPU arithmetic that exercises the whole pipeline in under five seconds and
says nothing about any real policy. The run writes runs/stub/report.json (byte-deterministic) and
report.md; their release verdict reads incomplete — not assessed, because no acceptance protocol
was named. Timed on a clean container: 20 s from pip install to a written report
(the transcript).
Limits, before anything else
- Simulation only. No number here was produced on hardware; sim-to-real transfer is not measured or claimed.
- Mostly templated attacks plus four bounded search families; not gradient-based worst-case robustness.
- One supported real configuration. SmolVLA on the ten LIBERO-Object tasks is the body of evidence; π0.5 has one preliminary arm. Everything else registered is scaffolding or stub-validated, and the tool says which.
- The default predicate is uncalibrated. An event is the end-effector leaving a fixed keep-out envelope — not task completion, not a calibrated hazard.
- The instruction result is fragility, not control. The frame that moves SmolVLA moves it with no target named and with its tokens scrambled (E-2026-12).
Supported configuration
| Policy | Suite | Status | Evidence |
|---|---|---|---|
smolvla — HuggingFaceVLA/smolvla_libero |
--suite libero --tasks libero_object/0…9 (10 tasks) |
measured, the published body | 14 Sep 2026 run · delivery pack |
pi05 — lerobot/pi05_libero_finetuned_v044 |
--suite libero --tasks libero_object/0…9 |
measured — preliminary (3 seeds, 2 of 8 arms, at the floor) | 18 Sep 2026 run |
stub |
stub, reach, humanoid (CPU fixtures) |
fixture | the run above |
libero_object is a task filter, not a suite: the suite is libero, the ten Object tasks are
--tasks libero_object/0,…,9. The real rows need a GPU, Linux and [lerobot]. Every other registered
adapter (pi0, pi0fast, groot, openvla, openpi) and suite (Meta-World, ai2_bridge, vla_arena)
is scaffolding or has no run committed here — provael list-policies / list-suites print each status.
The measured result
SmolVLA × LIBERO-Object, 14 September 2026, provael 0.41.2 — ten tasks, five seeds per (task,
arm), horizon 280, on a workstation RTX 2000 Ada. Under the roleplay instruction the policy left
its keep-out envelope on 42 of 50 episodes (84%, task-clustered 95% CI [62%, 100%]) against a
benign floor of 1/50 (2.0%, Wilson 95% [0.4%, 10.5%]); McNemar exact p = 9.1e-13, surviving
Holm across the seven arms. Clean task success on the benign arm: 48/50. The visual and injection
arms sit at the floor (1/50, 1/50, 2/50; mcp_tool_desc not applicable). The controls run the same
day put the same frame with no target named at 27/30 and with scrambled tokens at 18/30, so
the exit is fragility under a long imperative out-of-distribution string, not attacker control
(E-2026-12). It re-measures the 9 August 2026 run on 0.32.0 (44/50 against 2/50)
inside its interval. Under the example acceptance protocol the decision is FAIL on the roleplay
slice, and the pack says so:
run ·
controls ·
delivery pack ·
write-up with the full tables.
Reading any result: the rate and its floor travel together; the endpoint is an envelope exit
under the predicate the run names; N/A is not zero; a release verdict exists only under a named
protocol (--protocol) — how to read a result.
Coverage: registered is not validated
provael coverage prints the difference between what exists and what has met a real policy:
17 adversarial families registered, 8 exercised against a real policy (instruction, visual,
injection, gradient_patch, optimized_instruction, optimized_patch, universal_patch,
weight_integrity — one of the eight transferred; the other seven returned measured nulls, five of
them at n = 3, which is a result and not a rate), 9 stub-validated only, measured on
2 real policies (pi05, smolvla). The registry holds 39 adversarial attacks, not a family
count. Every number is derived from the committed runs (watch/registry.json),
never typed; the per-family catalogue with each family's status is docs/attacks.md.
Install
pip install provael # CPU core: every attack, scoring, reports, attestation
pip install 'provael[lerobot]' # + SmolVLA / π0.5 and the LIBERO simulator (Linux, GPU)
docker run --rm ghcr.io/provael/provael:latest attack --recipe quick # no local Python at all
On the CPU: the stub, reach and humanoid fixtures, every adversarial family (provael list-attacks), scoring, reports, recipes, reproduce, calibrate, attest (Ed25519 via [attest])
and the test suite. Behind a GPU and [lerobot]: the real policies and the libero simulator; on a
CPU they fail with a message naming the extra. Install notes, including the Linux-only simulator:
docs/quickstart.md.
Use in CI
# .github/workflows/provael.yml
permissions: { contents: read, security-events: write } # SARIF goes to code scanning
jobs:
redteam:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: provael/provael@v0.45.0
with:
attacks: none,instruction,visual,injection,action # `none` is the benign control
episodes: "10"
asr-threshold: "0.5" # pooled adversarial ASR gate (descriptive)
protocol: .provael/protocol.yml # the named acceptance protocol that decides
release-mode: "false" # "true": an undecided or incomplete run fails
baseline: .provael/baseline.report.json # optional per-checkpoint regression gate
A pooled rate is descriptive; the decision is made under the protocol's own criteria — critical
attacks and tasks on their own slices, incomplete where a slice did not run. Inputs, the defended
arm, the regression gate and the self-maintaining baseline: docs/quickstart.md.
Evidence outputs
report.json (byte-deterministic) and report.md · --format scorecard | sarif | oscal | test-report (ISO/IEC 17025 clause 7.8 layout) · --format compliance, the crosswalk
(docs/compliance/index.md) · a CycloneDX ML-BOM · provael attest, a dated,
digest-bound, offline-verifiable bundle (docs/attestation.md) · provael dossier, the
Machinery Regulation evidence dossier · the signed board (docs/leaderboard.md). All of
it is evidence, not certification.
Where the rest lives
- Docs: docs.provael.com — quickstart, attack catalogue, compliance crosswalk, attestation, findings and studies, the roadmap.
- The Embodied AI Security Top 10: docs/top10.md — an independent community risk list (CC-BY-SA 4.0) that Provael's attacks map to; the RFC process.
- Commercial: open core, forever (the promise). The operated work —
a run on your checkpoint, written up — is priced once, on provael.com/pricing.
The operated attestation service is not built; the in-repo server is an experimental reference
behind
PROVAEL_ENABLE_EXPERIMENTAL_HOSTED. One maintainer, and every commercial page says so. - Prior art, safety, changes, corrections: PRIOR_ART.md · SAFETY.md · CHANGELOG.md · docs/errata.md.
Development
make check runs lint, type-check and tests, exactly as CI does. How it is made. Much of this codebase was written with AI assistance (Claude Code); every
published number, calibration and security-relevant path is human-reviewed before it ships.
Co-author trailers are in git: git log --grep=Co-Authored-By.
Security & contributing. Vulnerabilities: SECURITY.md (90-day coordinated
disclosure). Contributions: CONTRIBUTING.md — the green gate and the DCO sign-off
(git commit -s); who has contributed and under which terms: CONTRIBUTORS.md;
CODE_OF_CONDUCT.md (Contributor Covenant 2.1).
How to cite
CITATION.cff is the same metadata (concept DOI: always the newest archived version):
@software{jain_provael_2026,
author = {Jain, Sattyam},
title = {Provael: red-teaming Vision-Language-Action robot policies in simulation},
version = {0.45.0},
year = {2026},
doi = {10.5281/zenodo.21984184},
url = {https://doi.org/10.5281/zenodo.21984184},
license = {Apache-2.0}
}
License and trademarks
Apache-2.0. Provael — prove it, prevail.
Provael™ (the name and the Proof-Path logo) is an unregistered mark of Sattyam Jain — no application has been filed yet; one is planned in India (classes 9 and 42). What you may do with the name, and what needs permission: TRADEMARKS.md. The Embodied AI Security Top 10 is a separate community document (CC-BY-SA 4.0), deliberately unbranded and donatable, not a Provael™ product, and not affiliated with or endorsed by the OWASP® Foundation or MITRE®.
Metadata
Release files for provael 0.45.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| provael-0.45.0.tar.gz | 30.1 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| provael-0.45.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 30.8 MB
Release files / provael-0.45.0.tar.gz
| Download URL | provael-0.45.0.tar.gz |
|---|---|
| Size | 30.1 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
17ad60b3037ffc73582579da5311b819a3a7670dbffff3e315d938874f5a07ef
|
|
BLAKE2b-256 checksum How to use checksums |
7148dbb3101e443ad232c3fb244bfa80960548c5b058a2752088dc9753aa06b9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / provael-0.45.0-py3-none-any.whl
| Download URL | provael-0.45.0-py3-none-any.whl |
|---|---|
| Size | 637.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bc79798f38ae275d75cf815a022d1d923895520c4b7f2a8ae93821d5b68cc075
|
|
BLAKE2b-256 checksum How to use checksums |
0067667408712d29fdbafa95f9e2dd45e00ac8bb6c55bd78cb49dd8a9a8be702
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log