Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

AI_Team

CLI-first autonomous software-engineering orchestration platform

From a plain-language idea to a planned, executed, validated, and deployable project — with sandboxed execution, Git automation, human approval, and a live developer control panel.

Python License: MIT Release Status Code style: Ruff CI

Quick Start · Documentation · Architecture · M14 · M15 · Frontend

Experimental proof-of-concept. This repository is a research project, not a finished product or enterprise platform. See Release Status and Known Limitations.


Workflow at a glance

flowchart LR
    Idea(["Idea"]) --> Brain["Engineering Brain<br/><i>6-stage LLM pipeline</i>"]
    Brain --> Plan["Plan<br/><i>requirements · PRD · spec<br/>architecture · tasks</i>"]
    Plan --> Prep["Prepare<br/><i>scaffold + init Git</i>"]
    Prep --> Exec["Execute<br/><i>adapter in ExecutionSandbox</i>"]
    Exec --> Val["Validate<br/><i>Ruff + Pytest</i>"]
    Val -->|fail| Rec["Recover<br/><i>self-healing retry</i>"]
    Rec --> Exec
    Val -->|pass| Git["Git<br/><i>branch + atomic commit</i>"]
    Git --> Appr["Approve<br/><i>HITL gate</i>"]
    Appr --> Deploy["Deploy<br/><i>manifest · validate · probe</i>"]
    Deploy --> Dash["Observe<br/><i>REST · WebSocket · React UI</i>"]

    classDef auto fill:#eef2ff,stroke:#6366f1,color:#1e1b4b;
    classDef gate fill:#fff7ed,stroke:#f59e0b,color:#7c2d12;
    class Brain,Plan,Prep,Exec,Val,Rec auto;
    class Appr gate;

What is AI_Team?

AI_Team is a Python platform that turns a high-level idea into a machine-readable engineering plan and then attempts to implement, validate, and package it autonomously. It is organised around three concerns:

  • The Engineering Brain — a six-stage LLM pipeline that produces requirements, a PRD, a project specification, an architecture, and a task plan as versioned Markdown/JSON artifacts.
  • The Execution Engine — a provider-agnostic framework that dispatches tasks to coding-agent adapters inside a confined local sandbox, tracks workspace diffs, validates the result with Ruff and Pytest, and retries through a self-healing loop.
  • The Control Panel — a FastAPI REST API with WebSocket event streaming, plus a React UI for projects, live runs, approvals, deployments, artifacts, and health.

Why it exists

Producing a complete software project from an idea requires many coordinated engineering steps that are usually done manually and are rarely auditable. AI_Team explores whether that chain can be made structured, inspectable, and reproducible: every stage emits typed artifacts, every execution is sandboxed and validated, every run is traced with correlation IDs, and automation is gated by human approval where it matters.

What V1 actually does

V1 is a working, tested proof-of-concept. Concretely, it can:

  1. Turn an idea into requirements, a PRD, a specification, an architecture, and a task plan.
  2. Generate a project repository scaffold from that plan.
  3. Execute tasks through a coding-agent adapter inside a sandboxed workspace.
  4. Run Ruff and Pytest validation against the result, with a bounded self-healing retry loop.
  5. Record the work in Git — task branches, atomic commits with structured metadata, and a pull-request summary artifact.
  6. Gate key stages behind a human-in-the-loop approval lifecycle.
  7. Generate and validate container deployment manifests (Dockerfile / compose / deployment metadata) and probe service health where an environment is available.
  8. Observe and drive all of the above through the dashboard API, WebSocket events, and React UI.

See Known Limitations for what is deliberately not included.

Architecture

The system is organised into strict, unidirectional layers. The dependency direction is enforced at test time by AST-based boundary checks (architecture_rules.py).

CLI (cli/)
  └── Application services (app/)
        ├── Pipeline engine (pipeline/, brain/)
        │     └── Providers (providers/)
        └── Execution engine (execution/)
              ├── Adapters (execution/adapters/)
              ├── Sandbox (execution/sandbox.py)
              ├── Validation (execution/validation/)
              └── Recovery (execution/recovery/)
Dashboard API (dashboard/) ── services ── reliability/state (core/reliability/)
Core infrastructure (core/): config, logging, exceptions, reliability, deployment, benchmark
Layer Directory Responsibility
CLI cli/ Click entry points (init, analyze, pipeline, run, …)
Application app/ Service façade between CLI and engines
Brain brain/, pipeline/ LLM stages, orchestration, artifact export
Providers providers/ LLM provider implementations + factory
Execution execution/ Adapters, sandbox, diff tracking, validation, recovery, Git
Deployment core/deployment/ Manifest generation, container validation, health probing
Dashboard dashboard/ FastAPI routers, services, WebSocket, React frontend
Core core/ Config, logging, exceptions, reliability, benchmark/evaluation

Architecture decisions are recorded as ADRs in docs/adr/ (layering, provider registry, adapter pattern, validation pipeline, structured logging, boundary tests). See the documentation index for the full map.

Core capabilities

Status legend: Implemented (works and is tested) · Partial (works, but with a bounded scope) · Scaffolded (interface + lifecycle only, no live integration) · Environment-dependent (real behavior requires external services) · Not in V1.

Area What V1 provides Status
Engineering Brain 6-stage pipeline: idea → requirements → PRD → spec → architecture → tasks Implemented
Artifacts Markdown + JSON outputs (requirements.md, PRD.md, architecture.json, task_plan.json, …) Implemented
LLM providers OpenAI, Google Gemini, NVIDIA NIM, and auto fallback routing Implemented
Execution adapters OpenHands adapter; live / live_llm LLM-driven adapter Implemented
Adapter registry Pluggable ExecutionAdapter contract + capability validation Implemented
Execution sandbox Process-level confinement, credential isolation, timeouts, process-tree termination, redaction Implemented
Validation Ruff + Pytest, sequential and parallel Implemented
Orchestration End-to-end autonomous orchestrator with stage events Implemented
Self-healing Validation-log-driven retry / repair loop Implemented
Git automation Init, task branches, atomic commits with .ai/commits/ metadata, merge, PR-summary artifact Implemented
HITL approvals PENDING gates, approve/reject by ID, project-scoped listing Implemented
Dashboard API FastAPI REST endpoints for projects, runs, approvals, deployments, health Implemented
Event streaming WebSocket stage/task/approval/deployment events Implemented
React UI Projects, project detail, active execution, approvals, health Implemented
Reliability Checkpointing, state store, schema/planning-artifact locks Implemented
Telemetry Structured JSONL logging with correlation IDs Implemented
Security Sandbox isolation, path-traversal protections, CORS pinning, regression tests Implemented
Quality tooling M9–M11 benchmark, evaluation, and release-engineering modules Implemented
Deployment (M15) Manifest + Dockerfile/compose generation, validation, health probing, records/logs Partial — container build/probe are environment-dependent
Live LLM execution Real provider-backed generation and task implementation Environment-dependent — requires provider credentials
Other adapters Claude, Cursor, Devin, VS Code, Codex, Antigravity Scaffolded
Anthropic provider AnthropicProvider Scaffolded (declared stub)
Per-task container isolation Ephemeral container sandbox per task Not in V1
GitHub PR API submission Creating pull requests via the GitHub API Not in V1 — a local PR-summary artifact is generated
Multi-agent peer review Security Auditor / Code Reviewer agents Not in V1

How it works

  1. Planning — PipelineEngine runs the registered brain stages and writes artifacts into projects/<slug>/. Each stage is an LLM call through the provider abstraction.
  2. Preparation — the orchestrator scaffolds a repository from the plan and initialises Git.
  3. Execution — each ExecutionTask is dispatched to an adapter inside an ExecutionSandbox; the workspace is an isolated copy, and file changes are tracked via diffs.
  4. Validation — Ruff and Pytest run against the workspace through the sandbox.
  5. Recovery — on failure, the self-healing engine builds a repair prompt and retries up to the configured budget.
  6. Git — successful tasks are committed atomically on a task branch with structured metadata.
  7. Approval — completed planning/architecture stages create PENDING approval gates.
  8. Deployment — deployment runs generate manifests, validate them, probe health, and record lifecycle logs.
  9. Observability — every transition is emitted as an event over WebSocket; checkpoints persist to core/reliability state.

Quick Start

Prerequisites

  • Python 3.11 or newer
  • Git (on PATH)
  • Optional: Node.js/npm to run the React dashboard
  • Optional: Docker for real container build/validation in the deployment lifecycle
  • Optional: an LLM provider API key for live provider-backed generation

1. Install

From PyPI (Recommended for users): (Note: After the first PyPI release, this will be the standard installation method)

pip install ai-team

From Source (For developers):

git clone https://github.com/Jayesh01323/AI_Team.git
cd AI_Team
python -m venv venv && source venv/bin/activate   # Windows: venv\Scripts\activate
pip install -e ".[dev]"                            # runtime + dev/test dependencies

2. Configure environment (only needed for provider-backed commands)

cp .env.example .env
AI_PROVIDER=gemini        # or openai / nvidia / auto
GEMINI_API_KEY=your-key

3. First CLI command (no credentials required)

ai-team --help
ai-team init "A simple test project"

4. Start the backend API

python -m uvicorn dashboard.api:app --host 127.0.0.1 --port 8000

Health check: http://127.0.0.1:8000/api/health

5. Start the dashboard UI (optional, separate terminal)

cd dashboard/frontend
npm install
npm run dev        # Vite dev server on http://localhost:3000, proxying /api and /ws to :8000

CLI

Command Purpose Provider credentials?
ai-team --help Show all commands and options No
ai-team init "<idea>" Create a project folder with placeholder templates No
ai-team test-provider Verify the configured LLM provider connection Yes
ai-team analyze "<idea>" Run the Idea Analysis stage Yes
ai-team generate "<idea>" Analyze and generate requirements.md Yes
ai-team pipeline "<idea>" Run the full 6-stage Engineering Brain pipeline Yes
ai-team scaffold "<idea>" Run the pipeline and generate a physical repo structure Yes
ai-team run "<idea>" --provider openhands Run the autonomous pipeline with execution + validation Yes for live providers

Dashboard

The backend is a FastAPI app exported as dashboard.api:app.

Backend (port 8000):

python -m uvicorn dashboard.api:app --host 127.0.0.1 --port 8000

Frontend (port 3000): see docs/frontend.md.

cd dashboard/frontend
npm install
npm run dev

Key REST routes (all under /api): GET /health, GET|POST /projects, GET /projects/{id}, GET /projects/{id}/artifacts/{artifact_name}, POST /projects/{id}/execute, POST /projects/{id}/deploy, GET /runs, GET /runs/{id}, GET /approvals, POST /approvals/{id}/approve|reject, GET /deployments, GET /deployments/{id}, GET /deployments/{id}/logs. Live events stream over WS /ws/events.

Example workflow

A factual end-to-end run (provider credentials required for the planning step):

# 1. Plan + generate artifacts (requires a provider key)
ai-team pipeline "A SaaS platform for analyzing resumes"
#    -> projects/build-a-saas-resume-analyzer/{requirements.md,PRD.md,architecture.json,...}

# 2. Scaffold a repository from the plan
ai-team scaffold "A SaaS platform for analyzing resumes"

# 3. Run the autonomous pipeline (planning -> execution -> validation -> git)
ai-team run "A SaaS platform for analyzing resumes" --provider openhands

# 4. Or drive everything from the dashboard API
python -m uvicorn dashboard.api:app --port 8000
curl -X POST http://127.0.0.1:8000/api/projects \
  -H 'Content-Type: application/json' -d '{"name":"demo","raw_idea":"A todo app"}'

Release Status

  • Version: v1.0.0-rc1 (release candidate)
  • Release commit: bd26c8d
  • Branch: master
  • Nature: experimental research / proof-of-concept

Release notes: docs/releases/v1.0.0-rc1.md.

Verification

The suite is credential-hermetic (live cases skip rather than fail) and organised by milestone.

python -m pytest -q                            # backend test suite
python -m ruff check core/ dashboard/ tests/   # lint
python run_m15_capstone.py                     # M15 deployment lifecycle capstone
  • Test suite: ~825+ backend tests plus module-level tests under brain/*/tests.
  • Ruff: clean on core/, dashboard/, tests/.
  • M15 capstone: passes end-to-end (manifest generation → container validation → Git → dashboard lifecycle).
  • Security/approval regressions: pass (artifact traversal, deployment-ID traversal, CORS pinning, approval lifecycle).
  • CI: .github/workflows/ci.yml runs Ruff + Pytest on Windows (Python 3.11).

Known test-harness issue: the M14 capstone asserts pull_request.md while the implementation writes PULL_REQUEST.md. This passes on case-insensitive filesystems (matching the current Windows CI) and fails on case-sensitive ones. It is a portability/test-harness mismatch, not a product defect.

Known Limitations

  • Process-level sandboxing, not per-task container isolation.
  • One implemented live execution backend (OpenHands) plus the live LLM adapter; the remaining six adapters are scaffolds.
  • Anthropic provider is a declared stub.
  • Live LLM execution requires provider credentials; without them the pipeline falls back.
  • Deployment container build and health probing are environment-dependent (Docker/HTTP) and otherwise simulated.
  • GitHub API PR submission is not implemented — V1 generates a local pull-request summary artifact.
  • Multi-agent role peer review is not implemented.
  • CI is Windows-only; frontend Vitest tests are not run by CI.
  • No Python dependency lockfile is committed.
  • Several projects/*/context.json runtime-state files are tracked and churn when the suite runs.

Documentation

Start at the documentation index.

Topic Location
Architecture decisions docs/adr/
Core architecture docs/core/ARCHITECTURE.md
Dashboard / M12 docs/m12/
Sandbox, Git & autonomous execution (M14) docs/m14/
Deployment lifecycle (M15) docs/m15/
Frontend docs/frontend.md
Security policy SECURITY.md
Release notes docs/releases/
Contributing CONTRIBUTING.md

Contributing

See CONTRIBUTING.md and the CODE_OF_CONDUCT.md.

Security

See SECURITY.md for the security policy and vulnerability reporting.

License

MIT — see LICENSE.


AI_Team · experimental V1 release candidate v1.0.0-rc1 (bd26c8d)

Built by Jayesh Patil · LinkedIn

Back to top · Documentation · Release notes

Metadata

Release files for ai-team 1.0.0rc1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ai-team 1.0.0rc1
File Size Uploaded
ai_team-1.0.0rc1.tar.gz 276.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ai-team 1.0.0rc1
File Interpreter ABI Platform
ai_team-1.0.0rc1-py3-none-any.whl Python 3 none any Details

Total release size: 644.2 kB

Release files / ai_team-1.0.0rc1.tar.gz

Download URL ai_team-1.0.0rc1.tar.gz
Size 276.6 kB
Tags Source
SHA-256 checksum
How to use checksums
23a9660743f33aca62960f054e38bc8ea0f437dbf817d176c089775e171c50a0
BLAKE2b-256 checksum
How to use checksums
cb1cde857fbd6b41cd76e619498f0b91b90e3de4a045ac217b8d9d40a3cbe8ef
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.

Transparency log

Release files / ai_team-1.0.0rc1-py3-none-any.whl

Download URL ai_team-1.0.0rc1-py3-none-any.whl
Size 367.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e77df59c175c8e68d847580abc2e566b8242932d3a3b03b3ed23259bf4323be7
BLAKE2b-256 checksum
How to use checksums
8200bed9f66d9cec794099b0df0a810a00b463d5caf06081b7dcb8e8b6c13052
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.0rc1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page