AI SAFE² CLI
Agent-facing assessment, evidence, decision support, and enforcement for AI SAFE² v3.1
Start | Give this to your agent | Install | Choose a workflow | Harness compatibility | Command reference | Limits
The CLI helps a person or an agent answer four practical questions:
- What agent systems and safety-relevant assets are here?
- What evidence do we actually have, and what is missing?
- What should a human review, block, remediate, or test next?
- Can another reviewer reproduce the result without trusting the chat?
It produces bounded JSON evidence for machines and readable Decision Cards for people. It does not silently connect to an agent account, approve deployment, or turn a self-assessment into a compliance claim.
Start with the outcome
| If you need to... | Start here | What you receive |
|---|---|---|
| Understand a workstation or repository that runs several agents | safe2 doctor . --assess --card-format markdown --card-output environment-card.md |
Harness inventory, coverage gaps, findings, and a human-readable Decision Card |
| Assess a project before adoption or release | safe2 assess . --scan-content --inspect-config |
A sealed assessment bundle with canonical JSON and a Decision Card |
| Screen a downloaded or newly created agent skill | safe2 gate skill PATH --strict |
An approve or reject decision with findings; the skill is never executed |
| Check whether an agent's completion claim has receipts | safe2 evidence claims ... |
Evidence-consistent, contradicted, unverifiable, or limited claim results |
| Determine which model, harness, tools, memory, and policy were assessed | safe2 evidence system ... |
A versioned system identity graph rather than only a model name |
| Diagnose where an agent-system failure likely occurred | safe2 evidence diagnose ... |
Evidence-backed investigation priorities, conflicts, assumptions, and next actions |
| Turn an AISM assessment into an actionable plan | safe2 aism plan ... |
A remediation card with owners, dependencies, alternatives, exit criteria, and residual risk |
| Test a governance claim rather than merely document it | safe2 challenge quickstart 001 ... |
Replayable Challenge Lab evidence with explicit provenance and limitations |
| Give CI a deterministic gate | safe2 gate project . |
Stable exit behavior and a human-readable findings report |
If you are evaluating the CLI itself, begin with the offline stranger-acceptance workflow. It proves that the installed package can reproduce its fixed acceptance controls; it is not an independent security certification.
Give this to your agent
Paste this into Codex, Claude Code, Hermes, OpenClaw, Antigravity, or another agent that can run local commands and read files:
Install AI SAFE² CLI 1.0 in an isolated Python environment. Do not use elevated
privileges, do not expose secrets, do not enable a daemon, and do not connect to
remote systems unless I explicitly authorize a named target. Run `safe2
self-check --strict`, initialize this project without overwriting existing
configuration, then run the bounded project assessment. Explain: (1) facts,
(2) assumptions, (3) missing coverage, (4) conflicts, (5) high-priority risks,
(6) recommended next actions, and (7) the exact evidence files created. Stop
before any remediation, deployment, policy change, network access, or risk
acceptance and ask me to decide.
The agent should follow the complete CLI 1.0 operator workflow. That workflow deliberately separates machine evidence from human authority.
Choose a workflow
| Situation | Workflow | Guide |
|---|---|---|
| First installation or upgrade | Self-check and offline acceptance | Installation · Acceptance |
| New repository or unfamiliar agent environment | Initialize, discover, then assess | Operator workflow · Unified assessment |
| New skill, plugin, or copied agent instruction | Quarantine, native gate, optional SkillSpector evidence, human decision | Skill-screening demo · Lessons learned |
| Agent says work is complete | Capture or import receipts, then audit explicit claims | Task receipts · Claim audit |
| Several harnesses contributed to one task | Normalize attributed exports and evaluate operational truth | Operational truth · Adapter SDK |
| Release or deployment decision | Bind system identity, scope, changes, evidence, and residual risks | Release readiness · Change attribution |
| Organization-level maturity decision | Ingest evidence conservatively, confirm mappings, score, and plan remediation | AISM guide · AISM remediation |
| Independent falsification or Challenge Lab participation | Run or import a bounded challenge and retain raw provider evidence | Challenge CLI |
Harness compatibility
The CLI is harness-neutral by default. Any agent or automation that can call a local executable, inspect JSON, and preserve files can use the core workflow. That includes coding agents, research agents, browser agents, bots, orchestrators, CI workers, and custom runtimes.
| Integration level | Current support | Examples |
|---|---|---|
| Direct CLI use | Supported | Codex, Claude Code, Hermes, OpenClaw, Antigravity, MiroFish, TinyFish, bots, Muse/Dot-style assistants, CI, and custom agents that can run commands |
| Metadata discovery | Named indicators for selected harnesses | Codex, Claude Code, Antigravity, Hermes, OpenClaw, and Grok indicators through safe2 doctor |
| Native evidence translation | Supported where an explicit adapter exists | Codex JSONL and OpenTelemetry JSON exports |
| Provider-neutral evidence | Supported | Versioned adapter, harness, system-identity, receipt, usage, scanner, and challenge contracts |
| Automatic account/session access | Not supported | The CLI does not log into or scrape an agent service |
| Universal pre-install or pre-action interception | Not supported in 1.0 | Requires a harness-specific hook or an external admission layer |
Do not confuse “works when invoked by this harness” with “native integration.” Named products change quickly; adapters must preserve provider version, scope, raw evidence, missing coverage, and the distinction between evidence and an AI SAFE² decision.
What it does not do
- It does not read private prompts, credentials, environment-variable values, or arbitrary configuration contents by default.
- It does not scan a network, discover cloud accounts, or connect to remote hosts unless an operator supplies an explicit supported target.
- It does not prove that a detected harness is active, current, securely configured, or using the files found on disk.
- It does not certify AI SAFE² conformance, organizational AISM maturity, successful task completion, verified billing, or absence of vulnerabilities.
- It does not authorize remediation, installation, release, deployment, exceptions, or risk acceptance.
Reference library
- Security advisories · Python compatibility · Stability policy
- System identity · Assessment scope · Failure localization
- Continuous evidence · OpenTelemetry · Codex JSONL
- Development method · Post-1.0 priorities
- Framework Home · AISM · Examples · NEXUS
The safe2 package turns repository controls, assessment logic, and evidence
contracts into a single command surface for agents, engineers, governance
teams, and CI systems. JSON is the canonical agent exchange format. Human
decisions remain human-owned and can be rendered as Markdown or HTML Decision
Cards.
Technology profiles and current CLI compatibility
The Technology Contribution Profile
and Technology Card are separate Markdown
review instruments. No TCP command, schema, or automatic score translation is
implemented. Existing safe2 aism ingest preserves attributed evidence with
unscored organizational cells for review. CLI E0-E5 weights, verification caps,
and category completeness remain unchanged; their input/summary grades are not
automatically Evidence Assurance v1.0 ratings.
Static, MCP, skill, and gateway scores retain their native meanings.
Where This Capability Lives
| Location | Role |
|---|---|
AISM/ |
Normative maturity model, architecture, methodology, assessment, and crosswalk |
safe2/aism/ |
Executable AISM validation, scoring, ingestion, comparison, and Decision Card rendering |
safe2/evidence/ |
Provider-neutral harness and system identity intake, evidence manifests, NEXUS, and NVIDIA SkillSpector adapters |
safe2/challenge/ |
Offline Challenge Lab runner, independent graders, provider translation, comparison, and artifact verification |
safe2/commands/ |
Unified CLI command groups and stable error handling |
examples/aism-decision-card/ |
Runnable reference assessment and human acceptance example |
This is a repository and packaging boundary, not a new framework pillar. AI SAFE² remains 161 core controls with CP.1 through CP.10. UAS is a separate 27-requirement regulatory profile extension.
Install
From a repository clone:
pip install -e ".[all]"
safe2 --help
Initialize a Project
Create the versioned project configuration before the first assessment:
safe2 init . --profile local
safe2 config show
safe2 config validate .safe2/config.toml
safe2 assess . --scan-content --inspect-config
Initialization creates .safe2/config.toml exclusively and refuses to replace
an existing file. The secure defaults collect no prompts, file contents,
environment-variable values, or network telemetry. Configuration precedence is
explicit command input, SAFE2_CONFIG, the nearest project configuration, then
built-in defaults. See the configuration contract
for profiles, limits, trust boundaries, and recovery.
safe2 assess is the bounded golden path. Without content consent it produces
an honestly INCOMPLETE metadata assessment rather than presenting uninspected
content as clean. --scan-content opts into local static content analysis;
--inspect-config additionally emits only allowlisted structural configuration
facts. The command writes assessment.json, environment.json,
project-scan.json, manifest.json, and a human decision-card.md atomically
to a new directory. It refuses to overwrite an existing bundle.
For contributor checks:
pip install -e ".[all,dev]"
pytest tests/ scanner/tests/
Command Map
| Command | Purpose | Decision behavior |
|---|---|---|
safe2 init PATH |
Create a secure-default, versioned project configuration | Refuses overwrite and symbolic-link configuration paths |
safe2 config show |
Show the normalized effective configuration and its precedence source | Inspection only; does not run an assessment |
safe2 config validate FILE |
Validate bounded safe2.config.v1 TOML |
Rejects unknown keys, unsafe paths, symlinks, and malformed values |
safe2 assess PATH [--scan-content] [--inspect-config] |
Produce one sealed environment and project assessment bundle | Missing or unrequested evidence remains explicit; no deployment authorization or conformance claim |
safe2 adapter validate/conformance ... |
Validate attributed external evidence adapters and specimens | Contract evidence only; never executes or endorses a provider |
safe2 scan project PATH |
Informational 161-control project scan | Reports findings; does not gate |
safe2 score project PATH |
Compact project score | Reports score only |
safe2 gate project PATH |
CI/CD project decision | Enforces tier or --fail-under threshold |
safe2 gate skill PATH --strict |
Skill trust decision | Approve, reject, or hold for review |
safe2 scan mcp PATH |
Static MCP source analysis | Informational |
safe2 score mcp URL |
Remote MCP assessment | Reports evidence-backed score |
safe2 report ... |
JSON, SARIF, Markdown, or HTML artifacts | Preserves native engine semantics |
safe2 evidence nexus PATH |
Collect NEXUS implementation/runtime evidence | Does not infer maturity |
safe2 evidence skillspector PATH |
Run optional NVIDIA SkillSpector adapter | Preserves upstream output and attribution |
safe2 evidence manifest FILE... |
Bind heterogeneous evidence into one run record | Hashes and validates artifacts without claiming conformance |
safe2 evidence harness SOURCE --output FILE |
Import a provider-neutral harness evidence export and expose coverage gaps | Source-attributed inventory only; does not verify execution, completion, billing, or conformance |
safe2 evidence system SOURCE --output FILE [--strict] |
Normalize model, harness, tool, skill, memory, environment, policy and evaluator identity | Declared/observed inventory only; does not verify deployed configuration or conformance |
safe2 evidence diagnose SOURCE --system-identity FILE --output FILE [--card FILE] [--strict] |
Rank evidence-backed failure locations across the complete agent system | Diagnostic support only; never claims a verified root cause, probability, or conformance |
safe2 evidence scope SOURCE --project-root PATH --system-identity FILE --output FILE [--strict] |
Inventory the declared deployment relationship of repository paths | Metadata-only attributed scope; does not verify build inclusion, execution, or conformance |
safe2 evidence attribute SOURCE --system-identity FILE --baseline-scope FILE --current-scope FILE --output FILE [--strict] |
Attribute normalized findings across a trusted baseline and current revision | Comparison support only; does not prove causation, deployment state, or conformance |
safe2 evidence readiness SOURCE --system-identity FILE --assessment-scope FILE --change-attribution FILE --output FILE --card FILE [--strict] |
Create agent JSON and a human technical release-readiness card | Human decision support only; never authorizes release or claims conformance |
safe2 evidence truth POLICY EVIDENCE... --output FILE --card FILE [--strict] |
Correlate multi-harness task evidence, receipts, usage, coverage and completion claims | Readiness for human decision only; never verifies completion, billing or conformance |
safe2 evidence claims SOURCE RECEIPTS... --output FILE --card FILE [--strict] |
Audit explicit agent outcome claims against receipt criteria | Detects contradictions and disclosures; never infers deception or verifies completion |
safe2 evidence changes ROOT [--baseline FILE] --output FILE [--strict] |
One-shot local/CI detection of changed skills and agent configuration | No daemon, telemetry or content export; changed configuration requires review |
safe2 evidence watch ROOT --state FILE --evidence-dir DIR [--continuous] [--strict] |
Preserve repeated local change reports and rescan changed skills | Polling detection only; does not intercept installs, pasted context, or harness execution |
safe2 aism remediation-init ASSESSMENT --system-identity FILE --assessment-scope FILE --decision-owner NAME --output FILE |
Create a source template bound to the exact AISM, system-identity, and deployment-scope artifacts | Produces no recommendation and authorizes no action |
safe2 aism plan SOURCE ASSESSMENT --system-identity FILE --assessment-scope FILE --output FILE --card FILE [--previous FILE] [--strict] |
Validate evidence-bound remediation actions, dependencies, alternatives, residual risk, completion evidence, and history | Keeps normative AISM scoring and human authorization separate |
safe2 aism init FILE |
Create a 30-cell unscored assessment | Missing evidence remains unscored |
safe2 aism ingest BUNDLE... |
Import evidence conservatively | Suggests mappings; requires human confirmation |
safe2 aism score FILE |
Validate and score AISM assessment | Produces agent JSON or human Decision Card |
safe2 aism compare OLD NEW |
Compare score, decision, and coverage history | Rejects malformed inputs cleanly |
safe2 example list |
Discover executable examples | Works in a clone and installed wheel |
safe2 feedback receipt FILE --artifact-root DIR |
Compare artifact hashes and bound test/tool reports; JSON/Markdown receipt | Artifact/report consistency only, not authenticated execution, completion, or billing |
safe2 feedback sign-report FILE --private-key KEY --signer-id ID --output FILE |
Create an expiring detached signature | Signs report bytes, not their truth |
safe2 feedback verify-report FILE SIGNATURE --trusted-public-key KEY |
Check report origin, integrity, and freshness | Does not establish execution or conformance |
safe2 feedback capture-process --execute ... -- ABSOLUTE_EXECUTABLE ARGS |
Directly observe an operator-authorized local process | Not a sandbox, test-success certificate, or task-completion gate |
safe2 feedback import-junit FILE ... --output FILE |
Normalize supported JUnit XML and check testcase totals | Import success is not passing tests or authenticated execution |
safe2 feedback capture-pytest --execute ... TARGET... |
Bind a local pytest process to its fresh JUnit report | Not isolated execution, tamper-proof runner evidence, or task acceptance |
safe2 feedback verify-pytest CAPTURE.json |
Recompute capture hashes, bindings, and test assessment without executing code | Internal consistency only; exit 0 does not mean tests passed or execution is authentic |
safe2 feedback usage INPUT... |
Correlate per-task usage declarations and detect duplicate ownership | Estimates remain separate; no billing verification |
safe2 example verify NAME |
Verify declared example outcomes | Fails on expectation drift |
safe2 mcp wrap ... |
Consumer-side MCP inspection and policy proxy | Applies runtime policy and audit behavior |
safe2 doctor PATH |
Metadata-only harness, shell, host, and WSL discovery | Inventory evidence only; does not claim assessment or conformance |
safe2 decision evaluate REQUEST --output FILE [--ledger FILE] |
Apply the deterministic decision firewall and optionally collect shadow System One evidence | Advisory review routing only; all merge, release, deploy, exception, and policy authority remains false |
safe2 decision replay CORPUS --output FILE |
Replay labeled routing cases after policy, rubric, threshold, or provider changes | Deterministic regression evidence; does not claim model calibration unless a provider evaluation is separately supplied |
safe2 dev plan SOURCE --output FILE |
Derive delivery-shape and risk-adjusted development requirements | Exit 0 means planning prerequisites are represented; grants no implementation or action authority |
safe2 dev receipt PLAN SOURCE --artifact-root DIR --output FILE |
Bind required development evidence, final revision, test cycle, reviews, findings, and rollback | Evidence consistency only; never authorizes completion, merge, release, deployment, or risk acceptance |
safe2 dev verify ARTIFACT |
Verify a development plan or receipt contract and integrity seal | Structural and byte-integrity check only |
safe2 dev replay CORPUS --output FILE |
Replay deterministic development-policy cases | Policy regression only; no live model or calibration claim |
safe2 feedback record ... |
Capture sanitized operational friction | Records typed outcome and verification state in local JSONL |
safe2 self-check [--format json] [--output FILE] [--strict] |
Verify installed runtime, dependencies, entry point, and packaged contracts | Offline installation evidence only; not signature, vulnerability, project, or conformance validation |
safe2 acceptance run DIR [--strict] / safe2 acceptance verify DIR |
Create and replay an offline first-user control bundle | Self-produced reproducibility evidence; explicitly not independent validation |
safe2 feedback summary FILE |
Measure recurring friction and completion-verification gap | Aggregates local evidence without sending telemetry |
safe2 schema list |
Discover packaged machine-readable contracts | Returns stable schema identifiers as JSON |
safe2 schema export NAME |
Export one versioned JSON Schema | Writes to stdout or an integration-owned file |
safe2 schema validate NAME FILE |
Validate an evidence artifact | Exit 0 valid, 1 contract violation, 2 unreadable input |
safe2 challenge ... |
Run inert fixtures or explicitly authorized bounded evaluators; import, compare, verify, sign, and report evidence | Challenge CLI guide; bounded execution is not sandboxing or independent replication |
Challenge Lab Evidence Workflow
The Challenge CLI guide provides an offline fixture workflow and an opt-in controlled evaluator seam for the same six Challenge 001 scenarios. Plans bind the named executable and limits before execution; receipts bind requests, responses, source evidence, normalized results, and optional system identity. The process runs with current-user authority and is not sandboxed. Raw provider verdicts, observation gaps, provenance, and incompatible conditions stay visible. Matching translated results do not establish independent replication.
safe2 challenge list
safe2 challenge quickstart 001 --output-dir my-first-run
safe2 challenge verify-bundle my-first-run
safe2 challenge run 001 --output challenge-run.json
safe2 challenge report challenge-run.json --format markdown
JSON runs can enter safe2 evidence manifest and safe2 aism ingest; the latter
preserves evidence with all 30 maturity cells unscored for human review. Optional
artifact signing uses pip install 'ai-safe2[challenge]' and an explicitly trusted
public key. An artifact signature is not human action approval or proof of state.
Multi-Harness Environment Discovery
safe2 doctor is the first local-first discovery surface for environments that
run more than one agent harness. It detects known command/configuration
indicators for Codex, Claude Code, Antigravity, Hermes, OpenClaw, and Grok,
along with available shells, the host operating system, CI markers, and WSL
availability.
safe2 doctor .
safe2 doctor . --format json --output environment-inventory.json
safe2 doctor . --assess
safe2 doctor . --no-wsl
safe2 doctor . --wsl-distro Ubuntu-24.04
safe2 doctor . --ssh-host audit@devbox.example --ssh-port 22
The v1 collector is deliberately metadata-only: it does not read configuration
contents, environment-variable values, prompts, tool output, or credentials.
It inventories WSL distribution names when the host exposes them. A named WSL
distribution can be inspected with --wsl-distro; an SSH-accessible Linux host
or cloud VM can be inspected with --ssh-host. Both execute a fixed,
metadata-only POSIX probe. SSH uses batch mode, requires an already trusted host
key, and will not prompt for passwords or accept a new host key. Targets must be
provided explicitly: safe2 doctor never scans a network. A discovery result
is not proof that a harness is active, current, or securely configured.
Cloud control-plane inventory, Windows remoting, containers, and deep configuration assessment are not implemented in the v1 collector. Failed or unreachable explicit targets are reported as incomplete rather than being treated as clean.
Add --assess to derive a first metadata-bounded posture. It identifies
project-policy review needs, stale or non-PATH installation indicators,
unreachable target coverage gaps, and multi-harness consistency needs. It does
not convert metadata into a security score: missing policy indicators retain
explicit alternative explanations, mappings are labeled as candidate controls,
and runtime/configuration/cloud coverage remains false until directly tested.
By default, the doctor also performs a bounded, filename-and-metadata-only inventory of security-relevant project assets:
- agent instruction and definition files;
- agent skills;
- MCP configuration candidates;
- persistent agent state and heartbeat indicators;
- CI/CD workflows;
- container definitions; and
- Terraform, Bicep, and Pulumi infrastructure definitions.
safe2 doctor . --assess --max-files 50000
safe2 doctor . --no-assets
safe2 doctor . --assess --hash-assets
safe2 doctor . --assess --inspect-config
Dependency, VCS, build, cache, and local evidence directories are excluded;
symbolic links are not followed. File contents and hashes are not collected by
default. --hash-assets opts into bounded local reads of recognized
security-relevant assets and emits only SHA-256 digests, enabling stronger
change detection for agent instructions, skills, CI, containers, persistent
state, and infrastructure definitions. Files above
--max-asset-hash-bytes remain explicit hash coverage gaps.
If the traversal limit is reached, the posture receives a high-severity
coverage-gap finding instead of treating the partial inventory as complete.
--inspect-config is a separate opt-in boundary. It reads only discovered JSON
or TOML harness/MCP configuration files and emits an allowlisted structural
summary: top-level key names, permission-rule counts, hook event names, MCP
server names and transport classes, selected sandbox/approval modes, file size,
and a content hash for later drift comparison. It does not emit URLs, commands,
arguments, header names or values, environment key names or values, prompts, or
raw configuration. Files that cannot be parsed or exceed the configured limit
remain explicit coverage gaps.
safe2 doctor . --inspect-config --max-config-bytes 1048576 --format json
safe2 doctor . --assess --inspect-config \
--output environment-inventory.json \
--card-format markdown --card-output environment-card.md
safe2 doctor . --assess --inspect-config \
--card-format html --card-output environment-card.html
The posture flags fully permissive sandbox/approval combinations for human review and reports only the count of secret-like key names. A matching key does not prove that a plaintext secret is present; it may contain a placeholder or environment reference.
The optional environment Decision Card is a concise human briefing derived
from the same sealed JSON inventory. Markdown is optimized for repositories and
review workflows; HTML is self-contained, responsive, and printable. Both show
the disposition, evidence confidence, scope, coverage, harness and asset counts,
drift history, integrity, deduplicated facts and assumptions, evidence
conflicts, persona impacts, prioritized actions, alternatives with pros and
cons, a recommended path, ownership gap, and exit criteria. Outcome probability
is explicitly NOT ESTIMABLE when only metadata evidence exists. Card creation
requires --assess and a separate --card-output, preserving the canonical JSON
artifact instead of replacing it.
Agent and CI Policy Decisions
An environment policy turns the evidence-bounded posture into deterministic agent and CI behavior. For example:
{
"schema_version": "safe2.environment-policy.v1",
"id": "production-agent-default",
"allowed_dispositions": ["BASELINE", "REVIEW"],
"max_findings": {"critical": 0, "high": 0},
"require_baseline": true,
"require_baseline_integrity": true,
"require_config_inspection": true,
"require_all_targets_completed": true,
"max_drift_changes": 0
}
safe2 doctor . --assess --inspect-config \
--baseline trusted-inventory.json \
--policy environment-policy.json --enforce-policy \
--output environment-decision.json \
--card-format markdown --card-output environment-card.md
Policy decisions are ALLOW (exit 0), DENY (exit 1), and HOLD (exit
2). DENY represents an observed threshold or allowed-disposition breach.
HOLD represents missing or incomplete evidence needed to decide safely. An
INCOMPLETE posture always holds even if a policy attempts to allow it.
Without --enforce-policy, evaluation is advisory and the command exits 0;
with enforcement, all requested JSON and card artifacts are written before the
decision exit code is returned.
ALLOW means only that supplied evidence met the named local policy. It is not
a universal safety determination, authorization, certification, AISM maturity
rating, or AI SAFE² conformance claim.
Trusted Baselines and Drift
Save a reviewed inventory, then compare later runs against it:
safe2 doctor . --no-wsl --assess --inspect-config \
--output .safe2/evidence/environment-baseline.json
safe2 doctor . --no-wsl --assess --inspect-config \
--baseline .safe2/evidence/environment-baseline.json \
--output .safe2/evidence/environment-current.json
The comparison reports added, removed, or modified harness indicators and
security-relevant assets, changed hashes for configurations inspected in both runs, lost target
coverage, and comparison-scope changes. Configuration hashes are available only
when both inventories used --inspect-config. Asset content comparison uses
hashes when both runs use --hash-assets; otherwise size or modification-time
changes are reported as weaker metadata drift. A change is a review signal, not
proof of unauthorized activity or elevated risk. The baseline should be retained
as trusted evidence only after an authorized human or policy workflow reviews
its scope, coverage, and known exceptions.
Use the same explicit --wsl-distro and --ssh-host targets in both runs. If a
baseline target is omitted or can no longer be inspected, the result carries a
high-severity coverage finding rather than treating missing evidence as no
change. Comparing a different root or target scope is also a high-severity
coverage finding.
Every newly written discovery inventory includes a deterministic SHA-256
integrity block covering the complete JSON result except the integrity block
itself. A modified sealed baseline is rejected before comparison. Older
safe2.discovery.v1 inventories remain usable but are labeled
baseline_integrity: not_present. Integrity proves that bytes represented by
the canonical JSON have not changed since sealing; because the v1 seal is
unsigned, it does not prove who created or approved the baseline. Store and
approve baselines through the repository's trusted evidence workflow.
Operational Friction Evidence
User and agent frustrations can be recorded as evaluation evidence instead of being lost in chat history. The initial taxonomy covers false completion, missing evidence, silent tool failure, wrong conclusions from missing data, stuck loops, sycophancy, context loss, permission friction, and integration failure.
safe2 feedback record \
--category false_completion \
--outcome unverified_done \
--severity high \
--harness codex \
--summary "Agent claimed completion without a resulting diff."
safe2 feedback summary .safe2/evidence/friction.jsonl \
--output .safe2/evidence/friction-summary.json
Outcome states are verified_done, unverified_done, failed, blocked, and
stuck. An evidence reference changes the verification label from
self_reported to external_reference_supplied; it does not independently
prove that the referenced evidence is valid. The summary exposes the gap
between claimed completion and verified completion so future evaluations can
optimize for truth rather than confident status language.
Each newly recorded event has a deterministic SHA-256 integrity seal. Summary
generation verifies every available seal and fails closed on a modified event,
preventing corrupted evidence from silently changing completion or frustration
metrics. Legacy unsigned events remain readable, contribute to the metrics, and
are counted explicitly through sealed_events, unsigned_events, and
integrity.coverage. These unsigned digests detect modification only; they do
not authenticate the person or agent that created or approved an event.
Run safe2 COMMAND --help for complete options.
Machine-Readable Contracts
Agent harnesses and CI integrations can discover and export the exact JSON contracts shipped with their installed CLI version:
safe2 schema list
safe2 schema export discovery-v1 --output discovery-v1.schema.json
safe2 schema export environment-posture-v1
safe2 schema validate discovery-v1 environment-inventory.json
The catalog includes AISM assessments, environment discovery, discovery drift, environment posture, friction events, and friction summaries. Integrations should select schemas by their versioned identifier and reject unknown major contracts rather than guessing from fields. Exporting a schema does not perform an assessment or validate an evidence artifact; it provides the contract for the harness's native validator.
schema validate is suitable for agent and CI branching: exit 0 means the
artifact satisfies the selected structural contract, exit 1 means contract
violations were found, and exit 2 means the input could not be safely read or
parsed. Validation output never includes instance values or verbose validator
messages. Structural validation does not verify evidence integrity, factual
accuracy, authorization, control effectiveness, or conformance.
Provider-Neutral Adapters
Validate third-party adapter descriptors and translate explicit Codex CLI JSONL exports without retaining prompts, commands, or output content. See the adapter SDK and Codex reference adapter.
safe2 adapter codex-jsonl codex-trace.jsonl \
--codex-version YOUR_CODEX_VERSION --output codex-evidence.json
Adapter records remain attributed evidence_only inputs. They are not provider
endorsements, independent truth verification, or AI SAFE² conformance claims.
OpenTelemetry users can also import offline OTLP/JSON trace files and export non-content SAFE² metadata. See the OpenTelemetry adapter.
For opt-in repeated change evidence, see Continuous Local Evidence.
For evidence-bounded completion and failure disclosure review, see the Agent Claim Audit.
Unified Evidence Run Manifest
After collectors produce their JSON artifacts, bind them into one portable run record:
safe2 evidence manifest \
environment-inventory.json \
friction-summary.json \
assessment.json \
--subject-id governed-workstation-01 \
--output run-manifest.json \
--strict
Each artifact record preserves its path, byte size, SHA-256 digest, declared
schema version, selected packaged contract, structural-validation state, and
available integrity-verification state. Unsupported, malformed, oversized,
symlinked, structurally invalid, or integrity-invalid artifacts remain visible
as invalid evidence instead of disappearing. --strict exits 1 after writing
the complete manifest if any artifact is invalid, allowing agents and CI to
retain diagnostic evidence while stopping promotion.
The manifest receives its own deterministic SHA-256 seal and records the CLI version, run ID, timestamp, subject, and coverage summary. It is an evidence inventory—not a maturity score, authorization decision, control-effectiveness test, or conformance claim. Its unsigned hashes establish change detection, not author identity or approval.
AISM Decision Workflow
safe2 evidence nexus ./NEXUS --output nexus-evidence.json
safe2 evidence skillspector ./candidate-skill --output skillspector-evidence.json
safe2 aism ingest nexus-evidence.json skillspector-evidence.json \
--subject-id governed-agent --subject-name "Governed Agent" \
--output assessment.json
safe2 aism score assessment.json --format json --output decision.json
safe2 aism score assessment.json --format markdown --output decision-card.md
The Decision Card exposes scores, maturity, evidence trust, facts, assumptions, conflicts, impacts, history, alternatives, pros and cons, why/why-not reasoning, outcome and complementary non-outcome estimates, recommendation ownership, review timing, and exit criteria. The tool never invents probability or treats scanner availability as proof of conformance.
Evidence and Trust Rules
- Unverified evidence is capped in the supplemental evidence-adjusted score.
- Digest-verified and independently verified artifacts require provenance.
- The raw AISM Sovereignty Score remains the normative maturity score.
- NEXUS is a reference implementation, not a mandatory conformance dependency.
- SkillSpector is optional; its upstream identity, version, license, target digest, timestamp, limitations, and non-endorsement are retained.
- Critical evidence conflicts force
HOLDfor accountable human resolution.
Output Contracts
- JSON is intended for agents and governance automation.
- Markdown and HTML provide human-readable Decision Cards.
- SARIF supports code-scanning and review systems.
- Exit codes distinguish successful execution, failed gates, human review, and invalid input where the command is decision-bearing.
Limitations
Static scans identify evidence and risk signals; they do not certify an organization or prove complete runtime behavior. AISM results depend on the scope, freshness, provenance, and independence of supplied evidence. Probability ranges are attributed assessment inputs, not predictions invented by the CLI.
Continue from your result
- If installation evidence is incomplete, run the stranger-acceptance workflow.
- If the assessment has missing coverage, authorize only the specific content, configuration, target, or provider evidence needed for the decision.
- If risks are actionable, use the AISM remediation workflow and retain the human decision owner.
- If a claim needs falsification, move the bounded question into the Challenge Lab.
- If the core is sufficient but integration is manual, use the post-1.0 priorities rather than inventing an undocumented native integration.
Navigation
| Previous | Current | Next |
|---|---|---|
| AISM | AI SAFE² CLI | Examples |
Framework Home | AISM | Cross-Pillar Governance | NEXUS | Dashboard
AI SAFE² v3.1 · Cyber Strategy Institute
Metadata
Release files for ai-safe2 1.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ai_safe2-1.0.1.tar.gz | 458.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ai_safe2-1.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.1 MB
Release files / ai_safe2-1.0.1.tar.gz
| Download URL | ai_safe2-1.0.1.tar.gz |
|---|---|
| Size | 458.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bdabac56abfdbd87833816267321579075545afa3bd2dfaa501d898884d75ad0
|
|
BLAKE2b-256 checksum How to use checksums |
af2ea8e53ec4773db72d0e68f482d09909648aaebe069c0d810e30cc151a906a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / ai_safe2-1.0.1-py3-none-any.whl
| Download URL | ai_safe2-1.0.1-py3-none-any.whl |
|---|---|
| Size | 607.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
48692ee48e02336ecae3686591459dfc12b369e839c12de14fd6a25af8976113
|
|
BLAKE2b-256 checksum How to use checksums |
8302d2b71187cc132b2cb6d8e472dc0d995d36cb25145a17cee2667a3e5200a0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log