oi-starship-helm
Install with pip install oi-starship-helm; import as starship_helm.
A lightweight governance SDK for AI agents: drop it into any agent to enforce identity, policy, kill-switch and reliability checks on every action, without that agent having to know it's talking to the Microsoft Agent Governance Toolkit (AGT) underneath. AGT itself lives behind a stable, layer-modular interface.
This is the user-facing doc — installing the published package, govern()'s
call signature, enabling/disabling layers. For the contributor-facing doc —
how the SDK is built, the layer pattern, the core/ facade — see
CONTRIBUTING.md in the source distribution. If you're
integrating this SDK into your own backend (especially with an AI coding
assistant's help), start with
docs/INTEGRATION_GUIDE.md (also in the source distribution) instead — it's a
step-by-step playbook, including a recipe for running the integration as a
guided conversation with Claude Code.
Why this exists
Wrapping AGT packages (agentmesh, hypervisor, agent_sre, ...) behind
one governed entry point keeps application code from ever touching AGT
directly. Hand-built per app, that wrapping tends to scatter AGT imports
across many files, so every AGT release becomes a hunt through the codebase.
This SDK packages that pattern once. Every AGT concern lives behind a port (a
stable interface) with exactly one adapter file that imports the real
AGT package. When AGT changes — a class renames, a package gets deprecated
in favor of another (this already happened three times during this SDK's
own design: agent-runtime → hypervisor, agent-mcp-governance →
deprecated stub with the real logic living in agentmesh instead,
agentmesh.marketplace → standalone agent-marketplace) — only that one
adapter file needs to change. Nothing that depends on the port notices.
Every layer is also independently enable/disable-able via config, with no conditional logic anywhere else: a disabled layer routes to a "noop" adapter implementing the same port with safe defaults, instead of the real one.
Install
Base install plus one extra per layer you want backed by real AGT:
pip install oi-starship-helm
pip install "oi-starship-helm[identity-trust,policy,runtime,reliability]" # the 4 core layers
pip install "oi-starship-helm[full]" # every layer
reliability is a no-op extra (an empty dependency list) kept only so
the command above still parses — its real dependency, a subset of
agent_sre, ships vendored inside this distribution (see
agent_sre/VENDORED.md), so the reliability layer's real adapter
works from the base install alone, with no extra to add.
Requirements: AGT-backed extras target agent-governance-toolkit-core
5.x. A workflow that ships a policy.rego also needs the
OPA CLI on
PATH — AGT 5.x has no built-in Rego evaluator, so without it every
governed call for that workflow is denied (rule="rego_unavailable")
rather than enforced partially.
A layer left disabled in config needs no extra installed at all — that's
the point of the noop path. See pyproject.toml's
[project.optional-dependencies] for the full extras list, one per layer.
Quick start
Author one YAML file:
# sdk.yaml
storage:
backend: local
root_dir: ./telemetry
policies_dir: ./policies/usecases # <workflow_id>/mesh.yaml + policy.yaml, shared below
identity_trust: {}
policy: {}
runtime: {}
reliability: {}
from starship_helm import GovernanceSDK, GovernanceDenied
sdk = GovernanceSDK.from_yaml("sdk.yaml")
def issue_refund(action: dict) -> dict:
return {"status": "refunded", "order_id": action["order_id"]}
try:
result = sdk.govern(
"orders", # workflow_id
"support-agent", # role
"issue_refund", # action_type
{"order_id": "O-1", "amount": 250},
issue_refund,
session_id="session-1",
)
except GovernanceDenied as exc:
print(f"denied by {exc.rule}: {exc.reason}")
GovernanceSDK.from_yaml(path) is cached per resolved path, so calling it
at every call site (rather than once at import time) is free after the
first call. SDKConfig/GovernanceSDK(config) — building the config tree
by hand in Python — still works exactly as before and is the only option
if you need a custom trust_store/backup_store object or a non-HTTP
mirror target; see docs/INTEGRATION_GUIDE.md steps 4 and 6.
Every govern() call runs the same 5 steps in order: identity resolve →
kill-switch check → guardrails + policy → trust/SLO update → audit.
PolicyConfig.policies_dir above is per-usecase and still required. Every
govern() call is also checked against a second, non-usecase-dependent
tier — compliance/harness/governance policy categories, cached
in-memory and read from this SDK's own bundled defaults unless
PolicyConfig.global_policies_dir points somewhere else, so no path is
needed to get global policy coverage. See CONTRIBUTING.md's "global
(non-usecase) policy cache" section for how it's cached and reloaded.
Enabling and disabling layers
Every layer's config has an enabled flag. Flipping it is the entire
mechanism — no other code changes:
config.reliability.enabled = False # SLO/incident recording becomes a no-op
sdk = GovernanceSDK(config) # govern() still works identically
The four layers above (identity_trust, policy, runtime,
reliability) default to enabled=True and participate in govern().
The other seven are peripheral — disabled by default, invoked on their own
schedule rather than gated per action:
from starship_helm.layers.assurance.config import AssuranceConfig
config.assurance = AssuranceConfig(enabled=True)
sdk = GovernanceSDK(config)
report = sdk.assurance.scan_prompt(system_prompt_text)
if report.is_blocking:
raise RuntimeError(f"prompt failed assurance scan: grade {report.grade}")
scan = sdk.discovery.scan() # find agents running in this environment
shadow_agents = sdk.discovery.reconcile()
A peripheral layer left disabled either returns a permissive empty result
(assurance, discovery, mcp_governance, rag_governance — these gate nothing
destructive) or raises LayerDisabledError when there's no safe
pass-through (sandbox code execution, plugin installation, wrapping an RL
training runner — these perform an action, so "disabled" can't silently
pretend to succeed).
Storage: where governance state gets exported
The audit trail, identity_trust's trust scores, runtime's kill-switch
state, and (optionally) reliability's SLO/incident snapshot are the SDK's
own runtime state — separate from policies_dir (mesh/policy/sre YAML),
which is versioned deployment config a host app ships alongside its code,
not something the SDK exports.
Every audit entry (sdk.audit.entries) has a stable, versioned shape
(starship_helm.audit.AuditEvent/SCHEMA_VERSION): fixed fields
session_id, agent_id (the role), use_case (the workflow id), action,
tool, ring, result ("success", "failure" or "denied") and
target_table ("session_executions" for governed actions,
"kill_switch_events" for kill-switch trips/disarms), plus a free-form
attributes bag for everything else. starship_helm.audit.to_platform_row(entry)
and to_otel_log_record(entry) translate one stored entry into a flat
database row, or into an OTel-style log record (name/timestamp/
attributes/resource — no opentelemetry package required) respectively —
whichever a host's own export pipeline needs. See CONTRIBUTING.md's
"audit/ — AuditTrail, and the canonical event schema it stores"
section for the full field mapping.
All of it is written through one StoragePort, selected by a single
SDKConfig.storage field:
from starship_helm.core.storage.config import StorageConfig
# Local disk (default) — files under root_dir.
StorageConfig(backend="local", root_dir="./telemetry")
# Export to any HTTP endpoint instead — an Azure Function, an AWS API
# Gateway + Lambda, a GCP Cloud Run service, or a plain REST app. No cloud
# SDK is imported, so the SDK itself stays platform-agnostic: the same
# adapter works unmodified regardless of which cloud is actually listening
# behind endpoint_url.
StorageConfig(
backend="http",
endpoint_url="https://my-api.example.com/governance",
api_key="...",
)
# Process-local, non-persisted — tests and short-lived processes.
StorageConfig(backend="memory")
An http backend must implement, keyed on {endpoint_url}/{key}:
| Method | Behavior |
|---|---|
GET |
200 with the raw bytes previously written, or 404 if the key was never written (or was deleted) |
PUT |
store the request body verbatim at that key; any 2xx |
DELETE |
remove the key if present; any 2xx or 404 both count as success |
That's the whole contract — an Azure Function backed by Blob Storage, a
Lambda backed by S3 or DynamoDB, or a Cloud Run service backed by GCS or
Firestore all satisfy it identically, so switching where the SDK is
hosted is a matter of pointing endpoint_url at that cloud's function,
not a code change here.
Each layer's own build() also accepts an optional storage argument
directly (identity_trust.build(config, storage)), so a layer built
standalone outside a full GovernanceSDK still works — it falls back to
an in-memory store if none is given.
Adding a fourth backend (e.g. a native Azure Blob / S3 / GCS SDK adapter
instead of going through HTTP) means adding one file implementing
StoragePort (see core/storage/local_adapter.py for the shape) and one
branch in core/storage/__init__.py's build() — nothing else in the
SDK references a storage backend directly.
Mirroring state into your own store: the three hooks
StoragePort covers "where does the SDK's own state live." Separately,
three generic, best-effort hooks exist for a host app that wants to mirror
that state into a second, app-owned store (e.g. a database your own
dashboard queries) without the SDK taking on a dependency on it. Every hook
follows the same contract: called synchronously right after the state it
mirrors is finalized, and any exception it raises is caught and logged —
a mirror failure can never break, deny, or delay a governed call.
| Hook | Set on | Fires | Payload |
|---|---|---|---|
audit_on_record |
SDKConfig |
once per sdk.audit.record(...) call |
the finished audit entry (dict) |
ReliabilityConfig.on_error_budget |
ReliabilityConfig |
once per SLO, after every record_outcome/emit_*_signal |
{slo_id, agent_id, use_case, budget_total, budget_consumed, period_start, period_end, updated_at} |
IdentityTrustConfig.trust_store |
IdentityTrustConfig |
on every trust resolve/update | not a callback — a full TrustStorePort you implement (load/seed/record_event) |
If your mirror target for the first two rows is a plain HTTP endpoint,
GovernanceSDK.from_yaml's http: shorthand (see
docs/INTEGRATION_GUIDE.md step 4) builds the callable for you — no Python
needed. trust_store (and anything that isn't a plain HTTP POST) still
means hand-building SDKConfig as below.
config = SDKConfig(
audit_on_record=lambda entry: my_db.insert("audit_events", entry),
reliability=ReliabilityConfig(
policies_dir="./policies/usecases",
on_error_budget=lambda row: my_db.upsert("error_budgets", row, key="slo_id"),
),
identity_trust=IdentityTrustConfig(
policies_dir="./policies/usecases",
trust_store=MyTrustStore(), # implements load/seed/record_event -- see
# starship_helm/layers/identity_trust/trust_store.py
),
...
)
Leave any of these unset and behavior is unchanged from before the field
existed: audit entries still get written (just not mirrored), reliability
still persists a snapshot if persist_snapshot=True, and trust scores
persist through the default StorageBackedTrustStore (current-value-only,
via the shared StoragePort).
All layers
| Layer | Config class | Wraps | In govern()? |
|---|---|---|---|
identity_trust |
IdentityTrustConfig |
agentmesh.client.AgentMeshClient |
Yes |
policy |
PolicyConfig |
agentmesh.governance.govern() + Policy + approval flow |
Yes |
runtime |
RuntimeConfig |
hypervisor.security.kill_switch.KillSwitch |
Yes |
reliability |
ReliabilityConfig |
agent_sre SLO + incidents |
Yes |
sandbox |
SandboxConfig |
agent_sandbox.docker_provider |
No — standalone |
assurance |
AssuranceConfig |
agent_compliance verify/prompt-defense/lint |
No — standalone |
supply_chain |
SupplyChainConfig |
agent_marketplace installer/signing/trust tiers |
No — standalone |
rag_governance |
RagGovernanceConfig |
agent_rag_governance.RAGGovernor |
No — standalone |
mcp_governance |
McpGovernanceConfig |
agentmesh.services.behavior_monitor.AgentBehaviorMonitor |
No — standalone |
training_governance |
TrainingGovernanceConfig |
agent_lightning_gov runner/reward |
No — standalone |
discovery |
DiscoveryConfig |
agent_discovery scan/reconcile/risk |
No — standalone |
a2a |
A2AConfig |
agentmesh.integrations.a2a (A2AAgentCard + A2ATrustProvider) + agentmesh.identity.AgentIdentity |
No — standalone |
Every layer, real AGT package installed or not, is reachable on the SDK
instance: sdk.identity, sdk.policy, sdk.runtime, sdk.reliability,
sdk.sandbox, sdk.assurance, sdk.supply_chain, sdk.rag_governance,
sdk.mcp_governance, sdk.training_governance, sdk.discovery, sdk.a2a.
a2a: Agent2Agent protocol support
Mints this agent's signed A2A card, verifies a peer before delegating a task to it, and creates that task:
config.a2a = A2AConfig(enabled=True, sponsor="ops@example.com")
sdk = GovernanceSDK(config)
card = sdk.a2a.issue_agent_card(
"orders", "support-agent", url="https://support-agent.example.com",
capabilities=["issue_refund"],
)
if sdk.a2a.verify_peer("orders", "support-agent", peer_did="did:mesh:..."):
task = sdk.a2a.create_task("orders", "support-agent", peer_did="did:mesh:...", message={"order_id": "O-1"})
Peer verification here covers the trust-bridge/cache path only — it does
not fetch and verify a remote peer's AI Card from its
.well-known/ai-card.json endpoint (what the A2A wire format expects
instead of an embedded card). Wiring in a real
agentmesh.trust.bridge.TrustBridge is the natural next step; not
implemented yet. AgentIdentity's DID also isn't stable across process
restarts — only its public key round-trips (to_jwk/from_jwk); the
private signing key lives in an AGT keystore this layer doesn't manage.
Adding support for a new AGT version
- Check
starship_helm/compat/versions.py—ADAPTER_TARGETSnames which AGT package and import root each layer'sagt_adapter.pytargets. Start there to find the right file. - Edit that layer's
agt_adapter.pyto match the new API. Theport.pyin the same folder should not need to change — if it does, that's a breaking change for every caller, not routine AGT-version drift. - Update the corresponding entry in
ADAPTER_TARGETS.
Nothing outside that one layer folder references the AGT package directly, so nothing else needs touching.
Authoring a new layer
Every layer is a folder under starship_helm/layers/<name>/
with four files, following layers/identity_trust/ as the reference
example:
config.py— aLayerConfigsubclass (fromlayers/base.py) withenabled: boolplus whatever the real adapter needs.port.py— aProtocoldefining the stable interface. Never imports AGT.agt_adapter.py— the only file that imports the real AGT package. Translates AGT exceptions/types intocore.errors/core.typesshapes.noop_adapter.py— implements the same port with permissive defaults (or raisesLayerDisabledErrorif there's no safe default — seelayers/sandbox/noop_adapter.pyfor that case).__init__.py— exposesbuild(config) -> Port, lazily importingagt_adapteronly whenconfig.enabledis true (so importing the package never requires the AGT dependency to be installed).
Register the new layer's config in core/config.py's SDKConfig, wire it
into core/sdk.py's GovernanceSDK.__init__ (and govern() if it's a
core, per-action layer), and add it to core/registry.py's
LAYER_MODULES so the generic test suite picks it up automatically.
Running the tests
pip install -e ".[dev]"
pytest tests
The suite passes with zero AGT packages installed — every test either
exercises a noop adapter directly, or is skipped (not failed) via
pytest.importorskip when the real AGT package it needs isn't present.
Install any layer's extra (e.g. pip install -e ".[identity-trust]") to
also exercise that layer's real adapter import. reliability is the one
exception: its real adapter (agent_sre, vendored — see
agent_sre/VENDORED.md) is always importable, extra or not, so its
test_real_agt_smoke.py case never skips.
Building a wheel
pip install build
python -m build --wheel
This directory (pyproject.toml's own location) is the build root. Its
[tool.hatch.build.targets.wheel]'s only-include = ["starship_helm", "agent_sre"]
ships exactly those two top-level packages — starship_helm and the vendored
agent_sre — and nothing else from this directory (tests/, docs,
pyproject.toml itself) ends up in the wheel. A new top-level package added
later needs a one-line addition to only-include, not a new remap.
Metadata
Release files for oi-starship-helm 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| oi_starship_helm-0.1.0.tar.gz | 148.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| oi_starship_helm-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 316.3 kB
Release files / oi_starship_helm-0.1.0.tar.gz
| Download URL | oi_starship_helm-0.1.0.tar.gz |
|---|---|
| Size | 148.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cd93da04c76eeb1d9bf720fbc24bb7b9fd377543bdf3542e85b9b126d69ee55f
|
|
BLAKE2b-256 checksum How to use checksums |
d1912b8f0acad2d0298fe99801d396b7077b9c20639fa0dbb553dd0e316054ad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|
Release files / oi_starship_helm-0.1.0-py3-none-any.whl
| Download URL | oi_starship_helm-0.1.0-py3-none-any.whl |
|---|---|
| Size | 168.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2d1682e4e53347c9c7af4e6c24dfee22375099adcb7748fe76a8468accf817da
|
|
BLAKE2b-256 checksum How to use checksums |
e0c9882f807d1bebf5690c65042dadac866b19c3c7ead7a75f4db37274a2ced3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|