Skip to main content

Cookiecutter-style scaffolder for an autonomous Claude Code SDLC configuration (no app code, no Docker). Asks ordered questions and installs CLAUDE.md + .claude/ (rules, the chosen profile's agents/skills, hooks, artifact templates) + optional .mcp.json; run /sdlc to drive spec → review → build → test → security → ship through profile-aware quality gates, working memory, and a self-improving learnings loop.

Project description

claude-kit

A Cookiecutter-style scaffolder for an autonomous SDLC (software-delivery lifecycle) inside Claude Code.

One command turns a one-line request into reviewed, tested, secured, shippable code — with a quality gate between every phase. No application code. No Docker. Configuration only.

PyPI Python License: MIT Built for Claude Code CI Changelog

🚀 Quick start · 🧭 How it works · 🔁 The pipeline · 🌱 What we adopted · 🤖 Agents · 🧩 Catalog · 🛠️ CLI · 📖 Agent guide


What is this?

claude-kit installs a complete software-delivery lifecycle into Claude Code. Instead of one assistant doing everything in a single pass, your work flows through a pipeline of focused agents — a spec writer, a story planner, reviewers, a developer, code reviewers, testers, security scanners, and a PR raiser — coordinated by an Orchestrator that runs independent work in parallel and refuses to advance a phase until its quality gate passes. You drive it all with one command: /sdlc <your task>.

At a glance:

  • 🧱 Stack-agnostic — the pipeline assumes no language or framework. Pick a stack at init and it installs matching overlay rules (React · FastAPI · PostgreSQL · MongoDB) and your exact lint/test/build commands. It never writes your app code and never needs Docker.
  • 🎚️ Dial the rigor with profileslean ⊊ standard ⊊ enterprise decide how many agents, skills, hooks, and gates are active, from "fast track" to "full audit".
  • 👥 Scope to your teamindividual / team (default) / organization. Org scope adds a vibe-coding layer so PMs, designers, QA, support, and founders can drive work safely too.
  • 🧠 Remembers across sessions — working memory (CONTINUITY.md) survives context compaction, and a learnings loop (agent-memory/) means the same mistake isn't made twice.
  • 📦 Two channels, one source — a first-class Claude Code plugin and a pip scaffolder.

Inspired by the autonomous-SDLC idea, rebuilt from the ground up for Claude Code, and kept small by a reuse-first policy — see what we adopted and from where.


Quick start

A) As a Claude Code plugin  (recommended)

Makes all agents, skills, commands, and hooks available inside Claude Code:

/plugin marketplace add ajyadav013/claude-kit
/plugin install claude-kit

Then, inside any project you want the pipeline to manage:

/claude-kit:init        # asks the ordered questions, lays down CLAUDE.md + .claude/
# ↻ restart Claude Code so the project's agents, skills & hooks load
/sdlc Add a CSV export button to the reports page

/sdlc is a project skill installed by init, so it becomes available after the restart. The plugin also exposes /claude-kit:sdlc <task>, which works immediately (no restart needed).

B) As a pip package  (CI, onboarding, non-plugin workflows)

A CLI (claude-kit, aliases ckit / claude-sdlc) that scaffolds the same config into any repo:

# Until the first PyPI release, install straight from the repo:
pip install "git+https://github.com/ajyadav013/claude-kit.git"
# Once published to PyPI this becomes:  pip install claude-code-kit

claude-kit init                 # interactive: prompts for stack, profile, MCP
claude-kit init --defaults      # non-interactive: React + Python/FastAPI + Postgres + standard

Prerequisites: Claude Code; Python ≥ 3.9 for the CLI; jq to enable the shell hooks (they no-op without it); Node / npx only if you enable an MCP (Model Context Protocol) server.

What the init flow asks & what lands on disk

claude-kit init asks an ordered set of questions (all with sensible defaults), then writes the config — nothing else:

  1. Target path (default: current dir; if .claude/ exists → merge / overwrite / backup / abort)
  2. Frontend framework (default: React) → frontend language (default: TypeScript)
  3. Backend language (default: Python) → backend framework (default: FastAPI)
  4. Database (PostgreSQL · MongoDB)
  5. SDLC profile (lean · standard · enterprise)
  6. Optional MCP integrations (GitHub · Jira/Linear · Postgres/Mongo · Playwright · Docs) — a project-root .mcp.json is written only if you select any (env placeholders, never secrets)

Non-interactive equivalents: --defaults, or --config init.yaml (flat or nested YAML). What lands:

CLAUDE.md                      # "Project-specific rules" filled from your stack's commands
README.claude-sdlc.md
.claude/
  settings.json                # assembled from the profile's hooks
  rules/                        # stack-agnostic core + selected overlay rules
  agents/                       # the profile's agent subset + DB overlay agents
  skills/  (incl. sdlc/)        # the profile's skill subset; sdlc/ is the /sdlc entrypoint
  hooks/                        # the profile's hook scripts
  templates/                    # artifact templates (spec, ADR, test-plan, …)
  config/                       # init-options.json (checksums) + stack snapshot
  state/  tmp/                  # gitignored runtime
.mcp.json                       # only if MCP servers were selected

How it works

flowchart LR
    subgraph SRC["claude-kit — single source of truth"]
        direction TB
        A["agents/"]
        S["skills/"]
        C["commands/"]
        H["hooks/"]
        R["rules/"]
        T["templates/"]
        K["catalog/<br/>(stacks · profiles · mcp · org)"]
    end
    SRC -->|"pip install + claude-kit init"| PROJ["Your project<br/>CLAUDE.md + .claude/"]
    SRC -->|"/plugin install"| CC["Claude Code<br/>(agents · skills · commands · hooks)"]
    CC -->|"/claude-kit:init"| PROJ
    PROJ --> RUN(["/sdlc — autonomous SDLC active"])

Three ideas do the heavy lifting:

  1. Quality gates with a shared severity model. Every finding is classified Critical / High / Medium / Low / Cosmetic. A gate passes only with zero Critical/High/Medium open. No silent advancement.
  2. RARV self-check. Every agent runs Reason → Act → Reflect → Verify and must show a green Verify (real commands run, not imagined) before handing off.
  3. Blind review + Devil's Advocate. Parallel reviewers judge independently; a unanimous PASS is treated as suspicious and triggers an adversarial devils-advocate pass before the gate may close — an explicit guard against agents rubber-stamping each other.

See docs/architecture.md for the full diagrams.


The pipeline

/sdlc reads the profile you chose and runs only that profile's gates:

flowchart TD
    REQ(["/sdlc request"]) --> CLS{"Classify"}
    CLS -->|"feature"| SPEC["Spec & Dev Docs"]
    SPEC --> EM{{"Gate: EM approved"}}
    EM -->|"pass"| STORY["Story breakdown + coverage gate<br/>story-planner"]
    STORY --> LANES["Parallel lanes:<br/>Senior Dev → Architect → Developer → Code Review"]
    LANES --> MR{{"Gate: Merge Reviewer"}}
    MR --> TEST["Unit · E2E · Integration + Senior verification"]
    TEST --> TCG{{"Gate: Test coverage<br/>+ Devil's Advocate"}}
    TCG --> SEC{{"Gate: Security Clear"}}
    SEC --> OPS{{"Gates: Pipeline Green ·<br/>Observability Ready · Acceptance"}}
    OPS --> PR["PR Raiser"] --> HUMAN(["Human review + deploy"])
    CLS -->|"fast-track (<5 files)"| FT["Developer → Review → Test → PR"] --> HUMAN
Profile Gates that run
lean code-review · build-green
standard spec-complete · em-approved · code-review · build-green · test-coverage · security-clear
enterprise standard + pipeline-green · observability-ready · acceptance

A fast-track mode collapses small changes (< 5 files) to Developer → Code Reviewer → Tester → PR.


Influences & what we adopted

claude-kit is built reuse-first. We periodically review excellent open-source projects and adopt only the genuinely-new ideas — never duplicating what the kit already does (near-duplicates would dilute Claude's ability to auto-select the right skill). Each adoption follows the same method: fetch the real source → adversarially map it against the kit's existing files → ship only the non-duplicative gaps, minimally and catalog-wired.

Source What we learned What we shipped Since
Agentic Design Patterns — A. Gulli (coverage map) Reasoning, guardrails, resilience, human-in-the-loop, evals, and tool design as first-class agent disciplines 8 agent-operation rules + docs/agentic-patterns.md 0.4.0
ponytail YAGNI / anti-over-engineering as an explicit recurring pass; deferral-debt tracking; surfacing the active autonomy level over-engineering-review & simplification-debt skills, the load-autonomy hook, median-of-N in evals 0.8.0
GitHub spec-kit Spec → tasks → analyze coverage gate; tasks → tracker issues; stable requirement IDs + assumptions in specs Wired the (previously orphaned) story-planner as the coverage gate (1f), a tracker-agnostic task-tracker-sync skill, and enriched the feature-spec template 0.9.0
protectai/llm-guard Input→model→output guardrails for LLM features — prompt injection, PII vault, treating model output as untrusted Opt-in "LLM / AI Feature Security" guidance in security-and-hardening + the advisory warn-llm-io hook (warns, never blocks) 0.10.0

Each adoption is detailed in the CHANGELOG — including, for every review, what we deliberately did not add because the kit already covered it.

The latest three reviews, in a bit more depth

🪶 ponytail → minimalism layer (0.8.0). Most of ponytail's philosophy (YAGNI, stdlib-first, surgical diffs) was already enforced by CLAUDE.md "Simplicity First" + code-simplification, so we added only the missing mechanisms: over-engineering-review (a complexity-only, report-only delete-list), simplification-debt (harvests TODO/FIXME/inline shortcut: markers into a ledger and flags ones that name no upgrade path), and the load-autonomy SessionStart hook (surfaces the active autonomy level each session).

🧭 GitHub spec-kit → coverage gate + tracker sync (0.9.0). The headline was reuse: the kit already had a story-planner agent that verifies every acceptance criterion maps to a story, but it was never wired into the pipeline. We made it stage 1f (between EM approval and the developer), so implementation can't start until coverage is proven. We also added task-tracker-sync (mirrors a plan into GitHub / Linear / Jira issues, dependencies preserved) and gave the feature-spec template stable requirement IDs + an Assumptions section. We skipped spec-kit's /constitution, /clarify, and /checklist — all already covered.

🛡️ protectai/llm-guard → opt-in LLM security (0.10.0). The kit secured the agent itself and traditional web appsec, but nothing covered the LLM features you build into your product (the OWASP LLM Top 10). Per request, the new layer is opt-in, bypassable, and states the risk of bypassing: an "LLM / AI Feature Security" section in security-and-hardening (input/output guardrails, PII vault, untrusted-output handling, a security-implications-of-bypassing table) plus a non-blocking warn-llm-io hook. We deliberately did not add a new rule or fold it into the mandatory security gate — that would have made it mandatory.


The agents

28 specialized roles in agents/, each tagged with a tier (orchestrator · stage-lead · specialist · review) and installed per profile — plus per-database overlay agents and, in organization scope, persona agents. The agent guide explains how to drive them.

See the full roster (28 + overlays + personas)
Agent Role
orchestrator Pipeline controller — decomposes, delegates, runs lanes in parallel, gates progression (never writes code)
spec-doc-writer Turns requirements into a spec + developer documentation in one pass
story-planner Decomposes an approved spec into ordered, parallelizable stories; verifies every acceptance criterion maps to a story (workflow gate 1f)
ui-designer Drafts and self-reviews UI/UX design specs
senior-backend-dev · senior-frontend-dev Senior review of a work stream's spec (the two-lane example)
technical-architect Cross-system architecture, scalability, integration review
em-reviewer Engineering-manager strategic & completeness review
merge-reviewer Verifies consistency between parallel lanes at join points
developer Writes production code from an approved spec, in an isolated worktree
sdlc-code-reviewer Reviews code for bugs, security, performance, spec compliance
unit-tester · e2e-tester Author unit and end-to-end test suites
tester · senior-tester Integration testing and independent verification of coverage
auditor Read-only audit for accessibility, performance, responsiveness, console errors
devils-advocate Anti-sycophancy adversarial reviewer (runs on a unanimous PASS)
acceptance-reviewer Verifies delivery against acceptance criteria before the human gate
risk-classifier Read-only — classifies work low/medium/high/restricted and names the required gates (enterprise + org)
security-reviewer Security stage coordinator — owns the Security Clear gate
secret-scanner · dependency-scanner · owasp-reviewer · policy-validator The four parallel security sub-scanners
devops-engineer CI/build/release, env, migrations, runbook — container-optional; owns Pipeline Green
observability-engineer SLOs, health/readiness, structured logging, alerts — owns Observability Ready
incident-responder Production-incident triage, mitigation, and postmortem (enterprise scope)
pr-raiser Final checks, commit hygiene, and PR creation
DB overlays installed for the selected database — PostgreSQL → postgres-specialist · migration-specialist · db-performance-reviewer; MongoDB → mongodb-specialist · migration-specialist
Org personas pm-copilot · founder-prototype-agent · support-ticket-engineer · data-workflow-agent · internal-tools-builder (organization scope only)

Rules & skills

Rules (rules/) are the 23 stack-agnostic contracts every agent obeys — the mandatory-workflow pipeline, quality-gates, rarv-cycle, continuity, documentation, testing, the eight agent-operation rules (reasoning-techniques, agent-guardrails, agent-resilience, goal-setting-and-monitoring, human-in-the-loop, model-tiers, evals, tool-design — see docs/agentic-patterns.md), and autonomy-levels + risk-classification (see docs/org-capabilities.md). Stack overlay rules (fastapi-patterns, react-patterns, postgres-patterns, …) and, in organization scope, org policy rules (secrets-policy, pii-policy, compliance-policy, …) layer on top.

Skills (skills/) are on-demand capabilities Claude activates by context — led by the sdlc entrypoint. Highlights, including this session's additions:

Skill What it does
spec-driven-development · planning-and-task-breakdown Spec first, then a verifiable task breakdown
task-tracker-sync Mirror a plan/story breakdown into GitHub / Linear / Jira issues (tracker-agnostic, idempotent)
security-and-hardening Traditional appsec + opt-in LLM / AI Feature Security (OWASP LLM Top 10, with a bypass + implications)
threat-model Design-time STRIDE — now with an LLM/AI branch
over-engineering-review · simplification-debt Keep the code lean: a complexity-only delete-list, and a deferral-debt ledger
test-driven-development · debugging-and-error-recovery · code-review-and-quality The build / fix / review staples
remember The self-improving learnings loop into agent-memory/

Each profile installs a subset (lean ⊂ standard ⊂ enterprise).


Catalog & extensibility

Everything selectable lives in catalog/ as data — adding a stack, framework, database, profile, or MCP server is a YAML edit plus a templates/stacks/<dir>/ folder, never a code change.

The four catalog files
  • catalog/stacks.yaml — frontend frameworks, backend languages → frameworks, and databases. Live today: React · Python/FastAPI · PostgreSQL/MongoDB. Vue/Svelte/Django/Express are listed as planned (offered by list-options, not yet selectable).
  • catalog/profiles.yaml — what each profile activates (inherit: composes; all = everything).
  • catalog/mcp.yaml — ready .mcp.json fragments per server, with ${ENV} placeholders.
  • catalog/org.yaml — the organization layer: scopes, teams, the autonomy model, review strictness, and the 7 capability packs. Scope-gated content under templates/org/ installs only when scope == organization. See docs/org-capabilities.md.

A third install dimension joins profile (a subset) and stack (an overlay): org (scope-gated). resolve() stays branch-free — adding a pack, team, autonomy level, or org rule is a catalog/org.yaml edit plus content under templates/org/, never a code change. Run claude-kit list-options to see everything available.


CLI reference

All commands (claude-kit · aliases ckit · claude-sdlc)
Command Description
init [path] [--defaults] [--config FILE] [--force] Scaffold CLAUDE.md + .claude/ (interactive or non-interactive)
validate [path] Structurally validate an installed config
doctor [path] Validate + environment/health checks with fix hints
diff [path] Preview what an upgrade would change (no writes)
upgrade [path] [--force] Refresh kit/overlay files; protect your edits; prune orphans
list-options List available frontend/backend/database/profile/MCP options
status [path] Show what's installed, the selection, and working memory
version Print the version
package-org-pack · install-org-pack Package / install an organization capability pack (org scope)

Plugin slash commands: /claude-kit:init, /claude-kit:sdlc <task>, /claude-kit:status; and the /sdlc skill inside any scaffolded project.

Safe upgrades — how your edits are protected

Every install records per-file checksums and an owner (kit / overlay / user-editable) in .claude/config/init-options.json. upgrade refreshes kit and overlay files to the latest version, never clobbers your edits (a user-modified file is kept and the new version dropped beside it as a .claude-kit sidecar), backs up anything it changes or removes, and restores files you deleted. Run diff first to preview.

Troubleshooting

Run claude-kit doctor first — it checks your environment (git, jq, hook scripts) and prints fix hints.

Symptom Likely cause Fix
/sdlc, agents, or skills "not found" right after init Claude Code hasn't loaded the new project config yet Restart Claude Code — or use /claude-kit:sdlc <task> (works without a restart)
Guard / quality hooks seem to do nothing jq isn't installed (the hooks parse tool input with it) Install jq; without it the hooks degrade to no-ops by design
A selected MCP server won't start node / npx missing (most MCP servers launch via npx) Install Node.js, or remove the server from .mcp.json
pip install claude-code-kit fails Not yet published to PyPI Use pip install "git+https://github.com/ajyadav013/claude-kit.git"
validate reports missing files Partial or outdated install Re-run claude-kit init (choose merge), or claude-kit upgrade

Project structure

Repository layout
claude-kit/
├── .claude-plugin/        plugin.json + marketplace.json
├── agents/                28 SDLC agents          rules/        23 engineering rules
├── skills/                on-demand skills        templates/    CLAUDE.md, settings, artifacts, memory seeds
├── commands/              /claude-kit:* commands  hooks/        hooks.json + scripts/
├── catalog/         stacks·profiles·mcp·org       templates/stacks/  per-stack overlay rules + agents
│                                                  templates/org/     org packs · personas · policies (scope-gated)
├── scripts/init.sh        thin fallback scaffolder  src/claude_kit/  the pip CLI (Typer + Jinja2 + PyYAML)
├── docs/architecture.md   diagrams                pyproject.toml   packaging

See docs/architecture.md for the full picture and CLAUDE.md for how to develop the kit itself.


Contributing

Issues and PRs welcome — see CONTRIBUTING.md. To dogfood a local checkout:

# As a plugin:  /plugin marketplace add .   then   /plugin install claude-kit
# As the CLI:   pip install -e '.[dev]'   then   claude-kit init /tmp/demo --defaults   &&   pytest

License

MIT © Arjunsingh Yadav

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

claude_code_kit-0.11.2.tar.gz (435.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

claude_code_kit-0.11.2-py3-none-any.whl (532.5 kB view details)

Uploaded Python 3

File details

Details for the file claude_code_kit-0.11.2.tar.gz.

File metadata

  • Download URL: claude_code_kit-0.11.2.tar.gz
  • Upload date:
  • Size: 435.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for claude_code_kit-0.11.2.tar.gz
Algorithm Hash digest
SHA256 594c6e91ae808091ac9f1d02364df5a25e3e810c94a3c03a348eabf959e2414c
MD5 a748c42b6f887617e312d326559cf001
BLAKE2b-256 212459b544e085fc74fb537672edf45a8b4ab084c204f552fbc466f395bffd7d

See more details on using hashes here.

File details

Details for the file claude_code_kit-0.11.2-py3-none-any.whl.

File metadata

File hashes

Hashes for claude_code_kit-0.11.2-py3-none-any.whl
Algorithm Hash digest
SHA256 fd6e669ff841e378cc503956fd888ce644b72a19e6f43c69c6f92ba0d1eabbd1
MD5 50ed80a3807e08d41bcda2617540b08f
BLAKE2b-256 722a17bbfb09147981af74367d450193c3aaf87c2f2c35368445bf7ea62787ed

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page