Skip to main content

Spine — governed, provenance-grounded autonomous delivery

Spine

Governed, provenance-grounded autonomous delivery — turn requirements into reviewed, tested pull requests, with a human in control.

Naming. Spine is the product. It's distributed as the synaptixs-spine package and its command is orchestrator — those names stay in install lines and commands throughout the docs.

Spine reads a requirement (from Confluence, Notion, a Markdown file, or an OpenSpec spec-driven change), understands your target repo, generates code grounded in that repo's own conventions, writes and runs tests, and opens a pull request for you to review. It pauses for your approval before it starts and before anything merges. Nothing is pushed, merged, or written to your tracker unless you say so.

It's built for teams who want agents that are inspectable, reproducible, and safe to run on real code — not demos.

How it fits together. Everything starts from one deterministic graph of your repo — built from the code and its docs — and every surface is a read of that graph:

flowchart LR
    repo["Your repo<br/>(code + docs)"]
    pkg["Product Knowledge Graph<br/>(deterministic, file:line)"]
    know["understand / state<br/>(episteme + health)"]
    ask["What breaks if I<br/>change X? What's<br/>untested?"]
    ev["Evidence<br/>(where it lands, root cause,<br/>blast radius — no model)"]
    build["sdlc feature<br/>(grounded codegen)"]
    pr["Reviewed PR"]
    repo --> pkg
    pkg --> know
    pkg --> ask
    pkg --> ev
    ev --> build
    build --> pr

Try it on your own repo in under a minute. No API key, no configuration, and it writes nothing — state reads your code and prints what it found:

pip install synaptixs-spine
orchestrator state /path/to/your/repo

Measured cold on a clean machine: 25s to install, 0.8s to answer. On pallets/click it opens with

This is a python library / service — 173 types and 1722 functions across 22 components (78 files). Top priority: Refactor 3 god-classes (>40 members), e.g. Context (60), ProgressBar (48).

then the architecture, the areas, documentation coverage and drift. Deterministic — no model runs, so the same commit always gives the same report.

When you want it to write code, that's when configuration starts to matter:

orchestrator init && orchestrator doctor                     # scaffold .env, check readiness
orchestrator sdlc feature --source file://./spec.md --safe   # build locally — no pushes, no PRs

What Spine does that other tools don't

Plenty of tools read your codebase. The difference is what they'll let themselves say about it.

1 · The graph is built by parsers, not by a model — and its accuracy is published. Eight language front-ends, every fact carrying file:line. Scored against a hand-labelled corpus in CI: precision 1.00 on every node and edge kind. Where it's weaker, that's published too — CALLS recall runs 1.00 on C and SQL down to 0.50 on TypeScript, reported separately rather than averaged into something flattering.

2 · The failure mode is silence, not fiction. Everything the graph asserts exists; what it can't resolve, it drops. A missing edge sends you looking. A fabricated one sends you confidently into a function nobody wrote. We hold precision at 1.00 and let recall be imperfect because that trade is the right way round for whoever reads it — and the invented-edge count is gated at zero, per language, on every commit.

3 · We measured whether it finds the right file, on code we don't control. Given a real bug report — the title alone — the file that actually fixed it is in Spine's top 10 27 times out of 38, and its first guess is right 12 times. Picking ten files at random from the same repositories would score 0.085. The corpus is five open-source projects pinned by commit, the answer key comes from each bug's own fixing commit, and the whole thing is reproducible with one command. The limits are published beside it — n=38, so top-1 sits in a 0.17–0.47 interval, and the bugs are cleaner than average. See BENCHMARK.md.

4 · We measured whether any of it helps, with a control. Across 260 ticket-runs on two frontier models: 47 of 68 new modules integrated correctly with the graph in context, against 3 of 68 without. The control — tickets that already named their target file — scored 122 of 124 either way, which is what rules out "more context just helps". Every published benchmark we could find in this category measures efficiency ("70% fewer tokens"). That answers how cheap, not is it right.

5 · Comprehension is deterministic, so it can be gated. Same commit in, same bytes out — no model, no cost, no variance. That's why understand --check can prove a knowledge base is current rather than hoping, why extraction can be cached per commit, and why a run's evidence can be replayed and diffed.

The honest summary: Spine is slower to claim things than the alternatives, and that is the product. Where it can't know something, it says so and stops.


What's new

3.26.0 (current)Spine measures whether it finds the right file, and publishes the number. Given a real bug report, the file that actually fixed it is in the top 10 for 27 of 38 issues and the first guess is right for 12; picking ten files at random from the same repositories scores 0.085. The corpus is five open-source projects pinned by commit, the answer key is each bug's own fixing commit, and one command reproduces it — BENCHMARK.md, which also states the limits. Building it turned up three checks that were passing while measuring nothing: fact freshness parsed every language as Python, the graph-grounded review layer never ran on a pull request at all, and a drift finding was rendered that nothing called.

3.25.1the issue type finally reaches the run. Spine has been issue-type shaped since 3.21.0 — the profile selector, the localization check, the bug/enhancement profiles — and nothing ever supplied the type, so every run took the default profile. A Bug now gets root-cause analysis and must localize; an enhancement gets a churn reading over its landing sites instead, and is no longer refused for naming the module it is about to create.

3.24.0 — the Claude Code and Codex plugins catch up with the product: multi-repo tools, pkg_joins, and an [all] install that actually extracts all eight languages.

3.23.0multi-repo comprehension: several repositories merge into one graph, and a ticket landing in one reports what depends on it in another. Declare them in .spine/repos.yaml, let orchestrator pkg joins --propose derive the topology from evidence, and investigate --repos reads it.

3.22.0 — four front-ends were fabricating a CALLS edge when a parameter shadowed a resolvable name. Measured at 47 fabricated edges across 11 public repositories, fixed, and gated at zero so it can't come back.

Full history, including the features we measured and didn't ship: CHANGELOG.


👉 See it work end to end — one ticket, start to finish

A real bug, in a real public codebase you can clone yourself (pallets/click) — from "where does this even live?" to a reviewed PR. Every command is one you can run, and the output is real. It finds the four functions in the blast radius that no test covers, in about a minute, with no API key.


🔒 Security

Spine runs on real code, clones untrusted repositories, and executes generated code — so we hold its own source to the same bar.

  • Checks run in CI on every pull request: CodeQL (Python + JavaScript), pip-audit over the locked dependency set, bandit-class static analysis, and Dependabot.
  • We security-reviewed our own source with a multi-model adversarial pass — 7 confirmed issues fixed, each with a regression test (path traversal, an SSRF backstop gap, prompt-injection hardening in the review pipeline, and a web-UI XSS). Details in the changelog.
  • All patchable dependency CVEs are resolved, and the audit fails CI on any new one.

Found something? Please follow our coordinated-disclosure policy in SECURITY.md — don't open a public issue.


Documentation

Guide Read it for
Worked example Start here. One ticket, start to finish, on a public repo you can clone — with real output you can reproduce command for command.
Setup & Install Installing the CLI, the .env, and standing up the full stack (Temporal + Postgres) for the autonomous pipeline.
User Guide A step-by-step walkthrough: from your first local build to a real PR, local models, the web dashboard, and connecting tools (MCP).
Using Spine from Codex Drive Spine from the Codex app — install (plugin or MCP server), credentials, the tool reference, and end-to-end greenfield + brownfield walkthroughs.
Using Spine from Claude Code Drive Spine from Claude Code — install (plugin or MCP server), credentials, the tool reference, and end-to-end greenfield + brownfield walkthroughs.
Features & Capabilities The capability catalog — everything Spine can do today, its status, the command/flag to use it, and a link to each deep dive.
Architecture How the whole platform fits together — the six layers, all components, the two human gates, and the knowledge graph they all read from. Includes an animated diagram.
Knowledge Graph (PKG) How Spine understands your codebase — the code-native graph, its model, the CLI, and how it powers brownfield and greenfield work.
Benchmark How well it actually works, measured — top-k localization on 38 real bugs across five languages, the corpus with its commit SHAs, what the numbers do not show, and the command to reproduce them yourself.
CLI Reference Every orchestrator command across all 7 areas — arguments, options, and defaults. Run orchestrator <command> --help for the live version.
Operations & Developer Guide How to operate it: deployment modes, the full environment-variable reference, and standing up each advanced capability — including the semantic spine (ontomesh × infodrift).
Community brief A one-page overview to share — what it does, lifecycle coverage, how to try it, and the feedback we're looking for.

New here? Install → User Guide Steps 1–4. That's the whole everyday workflow in about ten minutes.


Features & capabilities

Requirements → reviewed PR. Point it at a requirements source and a code repo. It extracts a backlog of intents, writes a spec, generates the implementation and tests, gets them green, and opens a PR — with two human gates (before building, before merging). A safe mode builds entirely locally (branch + diff, no external writes) so you can inspect everything first. Already written the spec yourself? Hand it straight to orchestrator sdlc autorun --spec <file> instead of deriving one from a source.

Plan before code. Before a run spends anything, orchestrator sdlc plan assembles a build document for the ticket — the requirement, the root cause, what the graph knows, the blast radius, the files, the acceptance criteria reconciled against code that already satisfies them, and what the codegen prompt will carry. Twelve sections, always the same, each labelled with where it came from: quoted, computed, inferred, or decided by a person. No model call, so the same commit produces the same document. sdlc approve records the decision against a digest of what you read, and a run refuses if the plan has changed since.

Code-grounded understanding. Before generating, it builds a Product Knowledge Graph of your repo — modules, types, functions, call sites, blast radius — and grounds new code in what already exists, so output reads like your team wrote it. The measurement behind that is above; the method and its bounds are published in full (on this repo · replicated on an unrelated codebase), and the harness ships with the package so you can get your own number.

Works across Python, Java, TypeScript, C#, C, C++ and Go, plus SQL data-layer comprehension (schema, queries, stored procedures, migration folding). It reads your documentation too — Markdown, reST, plain text and PDF — folding it in as Doc nodes linked to the code they describe, so you can ask which docs cover this symbol and where they've drifted. orchestrator understand writes a committed, code-true episteme/ your team and any AI tool can read — epistēmē, knowledge grounded in evidence, because every word of it is derived from the code rather than written by hand.

Across repositories, not just inside one. Declare your services in .spine/repos.yaml and they merge into a single graph, so "what breaks if I change this?" can answer with a caller in a different repo — an HTTP client, a shared table, an imported library. Spine derives the topology from evidence (pkg joins --propose) rather than asking you to draw it, and reports what it could not place, because a missing cross-repo edge looks exactly like two services that aren't coupled.

Governed autonomy. The workflow itself is a typed, validated artifact. A planner decomposes the objective, a runtime executes it, and per-edge verifiers check every step against schemas, evidence, and policy. Failures trigger replan, a human approval, or a clean stop. Every tool call, approval, and decision lands in an append-only audit log, and each run is capped by a spend budget.

Learns across runs. Cross-run semantic memory lets the agent recall conventions, pitfalls, and decisions from past runs — each memory cites the run it came from.

You can see inside it. Live OpenTelemetry tracing covers every LLM call, loop step, and tool call, joined to the audit log — so you can debug a run, not just read its result.

Use it your way. A CLI for scripting and CI, a web dashboard (delegate runs, watch them live, approve gates inline), a terminal UI, and MCP in both directions — consume external MCP tools, or expose the whole pipeline as an MCP server to Claude Code, Codex, or your IDE.

Bring your own model. Multi-provider via LiteLLM (Anthropic, OpenAI, Bedrock), or run fully offline on a local model (Ollama). Mix models per stage. Run orchestrator models to see which models are available — each id with its context window, price, and whether it supports the tool calling codegen and the judge need. The default is claude-opus-5.

Durable. Long-running pipelines are checkpointed (Temporal + Postgres) — they survive restarts and resume across human approval pauses.


How it works

A request flows top to bottom — through comprehension and planning, into a governed execution loop that pauses at two human gates — and out as a reviewed PR. The full architecture, with an animated diagram, is in ARCHITECTURE.md.

Spine platform architecture: surfaces → comprehension → planning → governed execution loop with two human gates → reviewed PR, over the Product Knowledge Graph

Every number on that diagram is read from the source, not typed into it. The version, the command count, the node and edge kinds and the language front-ends are computed at render time by scripts/render_architecture_svg.py, and CI fails if the checked-in image no longer matches. The SVG is the source; the PNG above is a rendering of it.

The version it replaced was stamped 3.8.4 and claimed 7 node kinds · 9 edge kinds — two releases after ARCHITECTURE.md had corrected them to 8 and 11. Nothing noticed, because a picture is the one artefact no test reads. Now one does.

  requirement (Confluence / Notion / Markdown)
        │
        ▼
   plan ──► validate ──► generate code ──► run tests ──► review ──► open PR
        │        (grounded in your repo's knowledge graph)        │
        └──────────── per-edge verifiers + audit ────────────────┘
                 human gate 1 ▲                    ▲ human gate 2
                 (before build)                    (before merge)
Concept What it is
Planner → GraphIR Turns an objective into a typed, validated execution graph (nodes, edges, budgets, approval points).
Registry Versioned agent templates + tool contracts the planner assembles from.
Runtime LangGraph-based executor with Postgres checkpointing and typed state.
Verifier chain Per-edge schema / confidence / evidence / policy checks that gate every handoff.
Approval gates First-class nodes that pause for human review and resume on your decision.
Audit log Append-only record of every tool call, approval, and policy decision.

FAQ

Does it merge code on its own? No. It opens a PR; a human reviews and merges. There are two approval gates — before building and before merging — and safe mode makes no external writes at all.

Where does my code/data go? To whichever LLM provider you configure — or nowhere external, if you run a local model (Ollama). Generated code stays in a local branch until you choose --live.

Do I need Docker or a database? Not for the everyday path (sdlc feature --safe builds one requirement locally). The autonomous multi-feature pipeline + web dashboard needs Temporal + Postgres — see the Setup guide.

Which languages and models? Comprehension and codegen cover Python, Java, TypeScript, C#, C, C++ and Go — each front-end going beyond structure into what that stack actually does (Java and C# REST endpoints, EF Core entities, C's #include graph, C++ templates and namespaces, Go interface satisfaction by method-set matching). SQL adds data-layer comprehension plus greenfield migration codegen validated against an ephemeral database. Docs fold in automatically; media (diagrams, screenshots, recorded reviews) via the opt-in media extract. Any LiteLLM provider — Anthropic, OpenAI, Bedrock — or a local Ollama model, and you can set a different model per stage. Extras and details: FEATURES.md.

How is it safe to run on real repos? Write guards on generated files, allow-listed + write-gated external tools, a per-run spend budget, an append-only audit trail, and human approval before any push or merge.

CLI or web UI? Either — they drive the same engine and the same API. Use the CLI for scripting/CI, the web UI (or terminal UI) for watching runs and approving gates by hand.

Can other tools call it? Yes. It speaks MCP both ways: it can use external MCP servers, and it can run as an MCP server so Claude Code / Codex / your IDE can call the pipeline (with the same gates).


Contributing

We'd genuinely like the help, and the codebase is unusually easy to be useful in.

It's plain Python. pip install -e ".[dev]", and the test suite runs in about three minutes with no services, no API key and no network. There's no build step anywhere — the web UI is vanilla JS on purpose. Most of the interesting work is a pure function over a graph, which means you can hold a change in your head and prove it with a fixture.

Good places to start

If you want to… Look at
Add a language pkg/*_extractor.py. Eight front-ends today; each is one file plus a labelled corpus case. Rust, Kotlin and Ruby are the obvious next three.
Improve accuracy corpus/ — hand-written fixtures with expected facts. Adding a case that fails is a real contribution; it's how the last four front-end bugs were found.
Fix something we've written down STATE-OF-SPINE §8 is a standing list of what's broken or missing, kept honest at each release.
Work on a bigger idea docs/specs/ — every design record, including the ones we closed unshipped and why.

How we work, in three points

  1. A fixture that fails first. New behaviour lands with a test written before it works and seen to fail. A test that passes before the code is a test that measures nothing — and a green check over an unexamined case is the mistake this project has made most often.
  2. Say what you didn't measure. A 0 that means "not checked" must not read like a 0 that means "clean". Bound your output honestly — "top N of M", never a clipped list implying completeness.
  3. Never guess in the graph. If a fact can't be resolved from a real parse tree, drop it. CLAUDE.md has the full set of invariants and the scars behind each one.

Before pushing: mypy src tests (not just src), ruff format --check ., and the suite. Work off develop. Details in CONTRIBUTING.md.

Or just tell us what you found

You don't have to write code to be useful — running it on a codebase we've never seen is genuinely valuable, especially if it gets something wrong.

See also the CODE_OF_CONDUCT.md.

License

MIT License. See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

synaptixs_spine-3.26.0.tar.gz (4.4 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

synaptixs_spine-3.26.0-py3-none-any.whl (1.2 MB view details)

Uploaded Python 3

File details

Details for the file synaptixs_spine-3.26.0.tar.gz.

File metadata

  • Download URL: synaptixs_spine-3.26.0.tar.gz
  • Upload date:
  • Size: 4.4 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for synaptixs_spine-3.26.0.tar.gz
Algorithm Hash digest
SHA256 d798b3246f6c2feb1cbe6fbfc61f34d5403c0af6f278a75f10b0588cff0c8294
MD5 0589d98593e82dbbeafbdeda316f99bd
BLAKE2b-256 74a99e96592a5fdb1fd48fd471fc2f629da29cd6e9b0e43c031b234ab3668ff6

See more details on using hashes here.

Provenance

The following attestation bundles were made for synaptixs_spine-3.26.0.tar.gz:

Publisher: publish-pypi.yml on synaptixs/spine

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file synaptixs_spine-3.26.0-py3-none-any.whl.

File metadata

File hashes

Hashes for synaptixs_spine-3.26.0-py3-none-any.whl
Algorithm Hash digest
SHA256 2a3aa25d2677e2860c61e2571947e1f349e8a602654fbb8a493b1ce304e26dc2
MD5 cd4d9e85c748087ba75b784b242eb52f
BLAKE2b-256 c58b18ae36f6bce402092895cdc546f6ae1d6c7881063ac30009cb0c5a667913

See more details on using hashes here.

Provenance

The following attestation bundles were made for synaptixs_spine-3.26.0-py3-none-any.whl:

Publisher: publish-pypi.yml on synaptixs/spine

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

3.32.0

2 files

3.31.0

2 files

3.30.0

2 files

3.29.1

2 files

3.29.0

2 files

3.28.0

2 files

3.27.0

2 files

3.26.1

2 files

This release

3.26.0 This release

2 files

3.25.1

2 files

3.25.0

2 files

3.24.0

2 files

3.23.0

2 files

3.22.0

2 files

3.21.0

2 files

3.20.0

2 files

3.19.0

2 files

3.18.1

2 files

3.18.0

2 files

3.17.0

2 files

3.16.2

2 files

3.16.1

2 files

3.16.0

2 files

3.15.0

2 files

3.14.0

2 files

3.13.0

2 files

3.12.0

2 files

3.11.1

2 files

3.10.0

2 files

3.9.3

2 files

3.9.1

2 files

3.9.0

2 files

3.8.4

2 files

3.8.3

2 files

3.8.2

2 files

3.8.1

2 files

3.8.0

2 files

3.7.0

2 files

3.6.1

2 files

3.6.0

2 files

3.5.0

2 files

3.4.0

2 files

3.3.0

2 files

3.2.0

2 files

3.1.0

2 files

3.0.0

2 files

2.8.1

2 files

2.8.0

2 files

2.7.0

2 files

2.6.2

2 files

2.6.1

2 files

2.6.0

2 files

2.5.0

2 files

2.4.0

2 files

2.3.0

2 files

2.2.0

2 files

2.0.1

2 files

1.24.0

2 files

1.23.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page