Skip to main content

sembl-stack

License: Apache 2.0 Built around Sembl

A swappable, spec-driven software factory. A task becomes declared bounds, an agent writes the change inside a disposable sandbox, the Sembl gate judges the real diff against those bounds, a PASS merges and deploys, and a post-deploy gate confirms it's healthy — or rolls it back. Every stage is an interchangeable adapter behind one typed artifact contract, and every run is recorded.

We sell process correctness — the change did what the spec declared, stayed in bounds, is honestly evidenced, and reached production accountably — never "the model writes better code." The stack takes no side in the agent wars: swap the executor, the sandbox, or the deploy target with one line of config and the rest of the pipeline doesn't notice.

Website · The gate (Sembl) · Architecture & plan

task ─▶ bounds ─▶ execute ─▶ sandbox ─▶ SEMBL GATE ─▶ merge ─▶ deploy ─▶ verify-in-prod
        (L2)       (L3)        (L4)       (L5)         (L6.5)    (L7)       (L8)
                                          every arrow is a typed artifact on disk

Quickstart

pip install sembl-stack
cd your-repo
sembl-stack

That's the whole quickstart. Bare sembl-stack launches the guided run — an OpenCode-style TUI that walks you through, one screen at a time:

  1. repo — confirms the repo it detected (a fresh non-git directory is offered a safe demo scaffold instead)
  2. agent & keys — every way to run AI work, with live status: is Claude Code installed and logged in, which API-key env vars are set, is OpenCode on PATH — and exactly what to do for any option you're missing (mock always works, zero keys)
  3. task — describe the change in plain English and pick which paths the agent may touch (suggested from your repo); the tool writes task.yaml/bounds.json, you never do
  4. run — the live stage rail (BLOCK → feedback → retry → PASS), then the verdict with its receipt and the one next command

Everything the guide does is also a scriptable command (same machinery, byte-identical runs — that's the point):

sembl-stack init                  # scaffold config + starter task + demo repo
sembl-stack doctor                # config-aware preflight
sembl-stack loop task.yaml        # plan → execute → gate → retry-on-BLOCK
sembl-stack runs [<id>]           # list / inspect runs
sembl-stack apply <id>            # apply the accepted patch (a BLOCK is never applied)

Presets (sembl-stack init --preset …):

Preset What runs Needs
just-gate gate any diff, nothing else only sembl
gate+sandbox the whole loop with a mock executor no API keys
full-loop real agent + sandbox + gate an executor key

Swap any layer in sembl.stack.yaml — e.g. execute: opencode, execute: aider — with no code change.

The stage map (L0–L8)

Each stage consumes and produces typed artifacts; that hand-off is the whole interface, which is what makes every stage swappable.

Stage Does Artifact flow Who owns it
L0 Protocol & hub one wire between stages we own (the contract)
L1 Repo intel code-graph context Task → Context adapter
L2 Spec → bounds scope the change Task → Bounds we own (sembl)
L3 Execute write the change Task + Bounds → Change adapter (claude / aider / opencode)
L4 Sandbox contain a bad diff Change → Change adapter (disposable clone)
L5 Verify gate the diff Change + Bounds → Verdict the gate (sembl)
L5.5 Review (advisory) code-quality signal diff → findings adapter (llm — BYO agent-CLI reviewer; CodeRabbit optional)
L6 Orchestrate loop, retry, trace wiring + * → Trace we own (LangGraph)
L6.5 Merge gated merge Verdict(PASS) → MergeRecord we own
L7 Deploy ship Verdict(PASS) → Delivery adapter
L8 Verify-in-prod gate production Delivery → Verdict the gate (health + rollback)

We own exactly three things: the artifact contract + stage Protocol, the gate (L5 + the post-deploy L8), and the glue + layer-replacement protocol. Everything else is deliberately a best-in-class tool behind an interface.

The accountable spine

A verdict is bound to the change it judged — most agent pipelines stop at "the check passed"; this one guarantees a verdict can only ship the exact change it was issued for:

  • Verdicts carry their subject. Every verdict is stamped with the SHA-256 and file set of the diff it judged. apply recomputes the patch hash and refuses a verdict issued for a different patch; merge refuses if the merge would ship files the verdict never saw.
  • BLOCK means blocked. A BLOCK verdict is never applied and never merged — the loop retries the executor instead. Overrides (--skip-binding-check) exist but are recorded permanently in the MergeRecord.
  • Production is gated too. After deploy, the L8 gate checks the live delivery (health + payload, deterministically) and triggers a rollback when it fails.

Every run leaves a complete paper trail in .sembl/runs/<id>/:

.sembl/runs/2ca41f/
├─ task.json          # what was asked
├─ bounds.json        # the declared contract
├─ change.json        # the actual diff
├─ verdict.json       # the gate's judgement + subject binding
├─ merge-record.json  # what shipped, and under whose PASS
└─ trace.json         # the timeline

The guided TUI

The guide is the default surface (see Quickstart) and a core dependency — no extra install. It streams the real loop live under your configured profile and is byte-identical to a headless sembl-stack loop, because it drives the same adapters. Re-run the agent & keys step anytime with sembl-stack --reconfigure.

The full picture

docs/PROCESS-ACTION-PLAN.md is the single source of truth: architecture, the L0–L8 stage map with build status, the eval metric, locked decisions, the guided-TUI vision, and the remaining-work plan.

Reference: process-self-improvement.md (north-star theory) · eval-metric-O3.md (the metric) · memory-plane-hypothesis.md.

Local development

uv sync --extra all
uv pip install -e ../sembl          # or: pip install sembl
.venv/Scripts/python -m pytest -q   # run from the repo root (corpus paths are cwd-relative)

Releasing

Publishing uses GitHub Actions + PyPI Trusted Publishing (OIDC); no tokens are stored. .github/workflows/release.yml builds and publishes when you publish a GitHub Release whose tag (vX.Y.Z) matches pyproject.toml and sembl_stack/__init__.py.


Agents write the code. sembl-stack makes the whole pipeline accountable.

Release files for sembl-stack 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sembl-stack 0.2.0
File Size Uploaded
sembl_stack-0.2.0.tar.gz 124.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sembl-stack 0.2.0
File Interpreter ABI Platform
sembl_stack-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 238.1 kB

Release files / sembl_stack-0.2.0.tar.gz

Download URL sembl_stack-0.2.0.tar.gz
Size 124.1 kB
Tags Source
SHA-256 checksum
How to use checksums
914502545ed41af0e9c7fa071f009d3a6b97e31eb8a706fcbc1133176662d15b
BLAKE2b-256 checksum
How to use checksums
90fa0541b442a1c4da52b6ac84a78248a162712598c77269d28161cf0fbb52e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release files / sembl_stack-0.2.0-py3-none-any.whl

Download URL sembl_stack-0.2.0-py3-none-any.whl
Size 114.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c4aded85a1f409e12f4cdd32759bf2935c8d28a5133a1bab391d10a06c37faf9
BLAKE2b-256 checksum
How to use checksums
3a782579c5825437e231e97c3d1c54e7f35a57fa2d9bd462d9a690fc3476b046
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release history Release notifications | RSS feed

0.3.0

2 release files

This release

0.2.0 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page