This release is a pre-release and may not be stable for production use.
AI_Team
CLI-first autonomous software-engineering orchestration platform
From a plain-language idea to a planned, executed, validated, and deployable project — with sandboxed execution, Git automation, human approval, and a live developer control panel.
Quick Start · Documentation · Architecture · M14 · M15 · Frontend
Experimental proof-of-concept. This repository is a research project, not a finished product or enterprise platform. See Release Status and Known Limitations.
Workflow at a glance
flowchart LR
Idea(["Idea"]) --> Brain["Engineering Brain<br/><i>6-stage LLM pipeline</i>"]
Brain --> Plan["Plan<br/><i>requirements · PRD · spec<br/>architecture · tasks</i>"]
Plan --> Prep["Prepare<br/><i>scaffold + init Git</i>"]
Prep --> Exec["Execute<br/><i>adapter in ExecutionSandbox</i>"]
Exec --> Val["Validate<br/><i>Ruff + Pytest</i>"]
Val -->|fail| Rec["Recover<br/><i>self-healing retry</i>"]
Rec --> Exec
Val -->|pass| Git["Git<br/><i>branch + atomic commit</i>"]
Git --> Appr["Approve<br/><i>HITL gate</i>"]
Appr --> Deploy["Deploy<br/><i>manifest · validate · probe</i>"]
Deploy --> Dash["Observe<br/><i>REST · WebSocket · React UI</i>"]
classDef auto fill:#eef2ff,stroke:#6366f1,color:#1e1b4b;
classDef gate fill:#fff7ed,stroke:#f59e0b,color:#7c2d12;
class Brain,Plan,Prep,Exec,Val,Rec auto;
class Appr gate;
What is AI_Team?
AI_Team is a Python platform that turns a high-level idea into a machine-readable engineering plan and then attempts to implement, validate, and package it autonomously. It is organised around three concerns:
- The Engineering Brain — a six-stage LLM pipeline that produces requirements, a PRD, a project specification, an architecture, and a task plan as versioned Markdown/JSON artifacts.
- The Execution Engine — a provider-agnostic framework that dispatches tasks to coding-agent adapters inside a confined local sandbox, tracks workspace diffs, validates the result with Ruff and Pytest, and retries through a self-healing loop.
- The Control Panel — a FastAPI REST API with WebSocket event streaming, plus a React UI for projects, live runs, approvals, deployments, artifacts, and health.
Why it exists
Producing a complete software project from an idea requires many coordinated engineering steps that are usually done manually and are rarely auditable. AI_Team explores whether that chain can be made structured, inspectable, and reproducible: every stage emits typed artifacts, every execution is sandboxed and validated, every run is traced with correlation IDs, and automation is gated by human approval where it matters.
What V1 actually does
V1 is a working, tested proof-of-concept. Concretely, it can:
- Turn an idea into requirements, a PRD, a specification, an architecture, and a task plan.
- Generate a project repository scaffold from that plan.
- Execute tasks through a coding-agent adapter inside a sandboxed workspace.
- Run Ruff and Pytest validation against the result, with a bounded self-healing retry loop.
- Record the work in Git — task branches, atomic commits with structured metadata, and a pull-request summary artifact.
- Gate key stages behind a human-in-the-loop approval lifecycle.
- Generate and validate container deployment manifests (Dockerfile / compose / deployment metadata) and probe service health where an environment is available.
- Observe and drive all of the above through the dashboard API, WebSocket events, and React UI.
See Known Limitations for what is deliberately not included.
Architecture
The system is organised into strict, unidirectional layers. The dependency direction is
enforced at test time by AST-based boundary checks (architecture_rules.py).
CLI (cli/)
└── Application services (app/)
├── Pipeline engine (pipeline/, brain/)
│ └── Providers (providers/)
└── Execution engine (execution/)
├── Adapters (execution/adapters/)
├── Sandbox (execution/sandbox.py)
├── Validation (execution/validation/)
└── Recovery (execution/recovery/)
Dashboard API (dashboard/) ── services ── reliability/state (core/reliability/)
Core infrastructure (core/): config, logging, exceptions, reliability, deployment, benchmark
| Layer | Directory | Responsibility |
|---|---|---|
| CLI | cli/ |
Click entry points (init, analyze, pipeline, run, …) |
| Application | app/ |
Service façade between CLI and engines |
| Brain | brain/, pipeline/ |
LLM stages, orchestration, artifact export |
| Providers | providers/ |
LLM provider implementations + factory |
| Execution | execution/ |
Adapters, sandbox, diff tracking, validation, recovery, Git |
| Deployment | core/deployment/ |
Manifest generation, container validation, health probing |
| Dashboard | dashboard/ |
FastAPI routers, services, WebSocket, React frontend |
| Core | core/ |
Config, logging, exceptions, reliability, benchmark/evaluation |
Architecture decisions are recorded as ADRs in docs/adr/ (layering, provider
registry, adapter pattern, validation pipeline, structured logging, boundary tests).
See the documentation index for the full map.
Core capabilities
Status legend: Implemented (works and is tested) · Partial (works, but with a bounded scope) · Scaffolded (interface + lifecycle only, no live integration) · Environment-dependent (real behavior requires external services) · Not in V1.
| Area | What V1 provides | Status |
|---|---|---|
| Engineering Brain | 6-stage pipeline: idea → requirements → PRD → spec → architecture → tasks | Implemented |
| Artifacts | Markdown + JSON outputs (requirements.md, PRD.md, architecture.json, task_plan.json, …) |
Implemented |
| LLM providers | OpenAI, Google Gemini, NVIDIA NIM, and auto fallback routing |
Implemented |
| Execution adapters | OpenHands adapter; live / live_llm LLM-driven adapter |
Implemented |
| Adapter registry | Pluggable ExecutionAdapter contract + capability validation |
Implemented |
| Execution sandbox | Process-level confinement, credential isolation, timeouts, process-tree termination, redaction | Implemented |
| Validation | Ruff + Pytest, sequential and parallel | Implemented |
| Orchestration | End-to-end autonomous orchestrator with stage events | Implemented |
| Self-healing | Validation-log-driven retry / repair loop | Implemented |
| Git automation | Init, task branches, atomic commits with .ai/commits/ metadata, merge, PR-summary artifact |
Implemented |
| HITL approvals | PENDING gates, approve/reject by ID, project-scoped listing | Implemented |
| Dashboard API | FastAPI REST endpoints for projects, runs, approvals, deployments, health | Implemented |
| Event streaming | WebSocket stage/task/approval/deployment events | Implemented |
| React UI | Projects, project detail, active execution, approvals, health | Implemented |
| Reliability | Checkpointing, state store, schema/planning-artifact locks | Implemented |
| Telemetry | Structured JSONL logging with correlation IDs | Implemented |
| Security | Sandbox isolation, path-traversal protections, CORS pinning, regression tests | Implemented |
| Quality tooling | M9–M11 benchmark, evaluation, and release-engineering modules | Implemented |
| Deployment (M15) | Manifest + Dockerfile/compose generation, validation, health probing, records/logs | Partial — container build/probe are environment-dependent |
| Live LLM execution | Real provider-backed generation and task implementation | Environment-dependent — requires provider credentials |
| Other adapters | Claude, Cursor, Devin, VS Code, Codex, Antigravity | Scaffolded |
| Anthropic provider | AnthropicProvider |
Scaffolded (declared stub) |
| Per-task container isolation | Ephemeral container sandbox per task | Not in V1 |
| GitHub PR API submission | Creating pull requests via the GitHub API | Not in V1 — a local PR-summary artifact is generated |
| Multi-agent peer review | Security Auditor / Code Reviewer agents | Not in V1 |
How it works
- Planning —
PipelineEngineruns the registered brain stages and writes artifacts intoprojects/<slug>/. Each stage is an LLM call through the provider abstraction. - Preparation — the orchestrator scaffolds a repository from the plan and initialises Git.
- Execution — each
ExecutionTaskis dispatched to an adapter inside anExecutionSandbox; the workspace is an isolated copy, and file changes are tracked via diffs. - Validation — Ruff and Pytest run against the workspace through the sandbox.
- Recovery — on failure, the self-healing engine builds a repair prompt and retries up to the configured budget.
- Git — successful tasks are committed atomically on a task branch with structured metadata.
- Approval — completed planning/architecture stages create PENDING approval gates.
- Deployment — deployment runs generate manifests, validate them, probe health, and record lifecycle logs.
- Observability — every transition is emitted as an event over WebSocket; checkpoints
persist to
core/reliabilitystate.
Quick Start
Prerequisites
- Python 3.11 or newer
- Git (on
PATH) - Optional: Node.js/npm to run the React dashboard
- Optional: Docker for real container build/validation in the deployment lifecycle
- Optional: an LLM provider API key for live provider-backed generation
1. Install
From PyPI (Recommended for users): (Note: After the first PyPI release, this will be the standard installation method)
pip install ai-team
From Source (For developers):
git clone https://github.com/Jayesh01323/AI_Team.git
cd AI_Team
python -m venv venv && source venv/bin/activate # Windows: venv\Scripts\activate
pip install -e ".[dev]" # runtime + dev/test dependencies
2. Configure environment (only needed for provider-backed commands)
cp .env.example .env
AI_PROVIDER=gemini # or openai / nvidia / auto
GEMINI_API_KEY=your-key
3. First CLI command (no credentials required)
ai-team --help
ai-team init "A simple test project"
4. Start the backend API
python -m uvicorn dashboard.api:app --host 127.0.0.1 --port 8000
Health check: http://127.0.0.1:8000/api/health
5. Start the dashboard UI (optional, separate terminal)
cd dashboard/frontend
npm install
npm run dev # Vite dev server on http://localhost:3000, proxying /api and /ws to :8000
CLI
| Command | Purpose | Provider credentials? |
|---|---|---|
ai-team --help |
Show all commands and options | No |
ai-team init "<idea>" |
Create a project folder with placeholder templates | No |
ai-team test-provider |
Verify the configured LLM provider connection | Yes |
ai-team analyze "<idea>" |
Run the Idea Analysis stage | Yes |
ai-team generate "<idea>" |
Analyze and generate requirements.md |
Yes |
ai-team pipeline "<idea>" |
Run the full 6-stage Engineering Brain pipeline | Yes |
ai-team scaffold "<idea>" |
Run the pipeline and generate a physical repo structure | Yes |
ai-team run "<idea>" --provider openhands |
Run the autonomous pipeline with execution + validation | Yes for live providers |
Dashboard
The backend is a FastAPI app exported as dashboard.api:app.
Backend (port 8000):
python -m uvicorn dashboard.api:app --host 127.0.0.1 --port 8000
Frontend (port 3000): see docs/frontend.md.
cd dashboard/frontend
npm install
npm run dev
Key REST routes (all under /api): GET /health, GET|POST /projects,
GET /projects/{id}, GET /projects/{id}/artifacts/{artifact_name},
POST /projects/{id}/execute, POST /projects/{id}/deploy,
GET /runs, GET /runs/{id}, GET /approvals, POST /approvals/{id}/approve|reject,
GET /deployments, GET /deployments/{id}, GET /deployments/{id}/logs.
Live events stream over WS /ws/events.
Example workflow
A factual end-to-end run (provider credentials required for the planning step):
# 1. Plan + generate artifacts (requires a provider key)
ai-team pipeline "A SaaS platform for analyzing resumes"
# -> projects/build-a-saas-resume-analyzer/{requirements.md,PRD.md,architecture.json,...}
# 2. Scaffold a repository from the plan
ai-team scaffold "A SaaS platform for analyzing resumes"
# 3. Run the autonomous pipeline (planning -> execution -> validation -> git)
ai-team run "A SaaS platform for analyzing resumes" --provider openhands
# 4. Or drive everything from the dashboard API
python -m uvicorn dashboard.api:app --port 8000
curl -X POST http://127.0.0.1:8000/api/projects \
-H 'Content-Type: application/json' -d '{"name":"demo","raw_idea":"A todo app"}'
Release Status
- Version:
v1.0.0-rc1(release candidate) - Release commit:
bd26c8d - Branch:
master - Nature: experimental research / proof-of-concept
Release notes: docs/releases/v1.0.0-rc1.md.
Verification
The suite is credential-hermetic (live cases skip rather than fail) and organised by milestone.
python -m pytest -q # backend test suite
python -m ruff check core/ dashboard/ tests/ # lint
python run_m15_capstone.py # M15 deployment lifecycle capstone
- Test suite: ~825+ backend tests plus module-level tests under
brain/*/tests. - Ruff: clean on
core/,dashboard/,tests/. - M15 capstone: passes end-to-end (manifest generation → container validation → Git → dashboard lifecycle).
- Security/approval regressions: pass (artifact traversal, deployment-ID traversal, CORS pinning, approval lifecycle).
- CI:
.github/workflows/ci.ymlruns Ruff + Pytest on Windows (Python 3.11).
Known test-harness issue: the M14 capstone asserts
pull_request.mdwhile the implementation writesPULL_REQUEST.md. This passes on case-insensitive filesystems (matching the current Windows CI) and fails on case-sensitive ones. It is a portability/test-harness mismatch, not a product defect.
Known Limitations
- Process-level sandboxing, not per-task container isolation.
- One implemented live execution backend (OpenHands) plus the
liveLLM adapter; the remaining six adapters are scaffolds. - Anthropic provider is a declared stub.
- Live LLM execution requires provider credentials; without them the pipeline falls back.
- Deployment container build and health probing are environment-dependent (Docker/HTTP) and otherwise simulated.
- GitHub API PR submission is not implemented — V1 generates a local pull-request summary artifact.
- Multi-agent role peer review is not implemented.
- CI is Windows-only; frontend Vitest tests are not run by CI.
- No Python dependency lockfile is committed.
- Several
projects/*/context.jsonruntime-state files are tracked and churn when the suite runs.
Documentation
Start at the documentation index.
| Topic | Location |
|---|---|
| Architecture decisions | docs/adr/ |
| Core architecture | docs/core/ARCHITECTURE.md |
| Dashboard / M12 | docs/m12/ |
| Sandbox, Git & autonomous execution (M14) | docs/m14/ |
| Deployment lifecycle (M15) | docs/m15/ |
| Frontend | docs/frontend.md |
| Security policy | SECURITY.md |
| Release notes | docs/releases/ |
| Contributing | CONTRIBUTING.md |
Contributing
See CONTRIBUTING.md and the CODE_OF_CONDUCT.md.
Security
See SECURITY.md for the security policy and vulnerability reporting.
License
MIT — see LICENSE.
AI_Team · experimental V1 release candidate v1.0.0-rc1 (bd26c8d)
Built by Jayesh Patil · LinkedIn
Metadata
Release files for ai-team 1.0.0rc1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ai_team-1.0.0rc1.tar.gz | 276.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ai_team-1.0.0rc1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 644.2 kB
Release files / ai_team-1.0.0rc1.tar.gz
| Download URL | ai_team-1.0.0rc1.tar.gz |
|---|---|
| Size | 276.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
23a9660743f33aca62960f054e38bc8ea0f437dbf817d176c089775e171c50a0
|
|
BLAKE2b-256 checksum How to use checksums |
cb1cde857fbd6b41cd76e619498f0b91b90e3de4a045ac217b8d9d40a3cbe8ef
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.
Transparency logRelease files / ai_team-1.0.0rc1-py3-none-any.whl
| Download URL | ai_team-1.0.0rc1-py3-none-any.whl |
|---|---|
| Size | 367.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e77df59c175c8e68d847580abc2e566b8242932d3a3b03b3ed23259bf4323be7
|
|
BLAKE2b-256 checksum How to use checksums |
8200bed9f66d9cec794099b0df0a810a00b463d5caf06081b7dcb8e8b6c13052
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.
Transparency log