Multi-agent AI incident commander for Kubernetes — correlates deploys, logs, metrics, and suggests rollback with human approval
Project description
Incident Commander
Multi-agent Incident Commander for production incidents. All integrations are live — kubectl, Loki, Prometheus, and Groq (optional). No mock mode.
Agentic AI features
| Feature | Implementation |
|---|---|
| Worker ReAct loops | Deterministic multi-step tools + optional Groq LLM ReAct per worker |
| Post-action verification | After rollback, polls pods/logs/metrics until healthy or escalate |
| Eval / replay | JSON fixtures, incident-commander eval, incident-commander record |
ReAct workers
Each worker runs a Reason → Act → Observe loop (up to MAX_WORKER_ITERATIONS):
- Deploy — recent deploys → expanded lookback → rollout history → optional LLM
- Logs — error patterns → expanded window → optional LLM
- K8s — deployment pods →
app=label fallback → warning events - Metrics — Prometheus snapshot → retry → optional LLM
Verification loop
After approved rollback:
- Execute
kubectl rollout undo - Poll every
VERIFY_INTERVAL_SECONDS(default 15s) - Check pods healthy, errors quiet, error rate (if Prom available)
- Resolved if checks pass, else escalated
Eval / replay
incident-commander eval # built-in scenarios (packaged with pip install)
incident-commander record INC-... # save incident as fixture
Built-in fixtures ship inside the Python package (incident_commander/eval/fixtures/).
Prerequisites
- Python 3.11+
- Docker Desktop (running)
kubectl(installed with Docker Desktop or separately)- Optional: Prometheus, Loki, Groq API key
One-time cluster setup (Docker + kind)
./scripts/setup-k8s.sh
This creates a kind cluster named incident-commander on Docker with host ports wired (localhost:9090 → Prometheus, localhost:3100 → Loki) and deploys a sample payment-api deployment. Config lives in k8s/kind-config.yaml.
If you already have a cluster without those ports:
KIND_RECREATE=1 ./scripts/setup-k8s.sh
Observability (Prometheus + Loki)
./scripts/setup-observability.sh
Installs Prometheus, kube-state-metrics, Loki, and Promtail in the monitoring namespace. After setup, Prometheus and Loki are reachable on your Mac at :9090 and :3100 without kubectl port-forward.
Set in .env:
LOG_BACKEND=loki
PROMETHEUS_URL=http://localhost:9090
LOKI_URL=http://localhost:3100
Optional fallback if host ports are missing: ./scripts/observability-forward.sh
Or manually (new cluster with port wiring):
brew install kind # if needed
kind create cluster --name incident-commander --config k8s/kind-config.yaml
kubectl apply -f k8s/sample-payment-api.yaml
Health check (run before demos or commits)
make check
# or: ./scripts/check.sh
Runs: compile → pytest (10 tests) → eval fixtures (3 scenarios).
Quick start
Install from PyPI
pip install incident-commander
# Optional: configure cluster + LLM
cp .env.example .env # from the GitHub repo, or set env vars directly
incident-commander doctor
incident-commander open payment-api --namespace default
PyPI: pypi.org/project/incident-commander
Clone for local development (kind demo + k8s scripts)
git clone https://github.com/sahilleth/incident-commander.git
cd incident-commander
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
cp .env.example .env
# Optional: add GROQ_API_KEY — heuristic hypotheses work without it
./scripts/setup-k8s.sh
./scripts/setup-observability.sh
make scenario-bad-deploy
incident-commander open payment-api --namespace default
CLI
| Command | Description |
|---|---|
incident-commander open <deployment> |
Open + investigate (deployment name = K8s Deployment) |
incident-commander list |
Recent incidents |
incident-commander show INC-... |
Incident detail |
incident-commander approve INC-... APR-... |
Approve rollback (runs kubectl rollout undo) |
incident-commander serve |
API on http://localhost:8080 |
Options for open
incident-commander open payment-api \
--namespace production \
--trigger manual:health-check \
--severity SEV1
API
incident-commander serve
curl -X POST http://localhost:8080/incidents \
-H 'Content-Type: application/json' \
-d '{"service":"payment-api","namespace":"default","trigger":"manual"}'
Architecture
Trigger (CLI / API) → Incident Commander (supervisor)
↓ parallel workers
Deploy (kubectl rs) | Logs (Loki/kubectl) | K8s (pods/events) | Metrics (Prometheus)
↓
Hypothesis synthesizer (Groq or heuristic)
↓
Approval → Runbook executor (kubectl rollout undo) → Verifier
↓
SQLite persistence
Configuration
| Variable | Purpose |
|---|---|
GROQ_API_KEY |
Primary Groq LLM key (optional) |
GROQ_API_KEY_FALLBACK |
Second Groq key — used when primary is rate-limited |
GROQ_MODEL |
Default llama-3.3-70b-versatile |
KUBECONFIG / KUBE_CONTEXT |
Cluster access |
LOG_BACKEND |
kubectl or loki |
LOKI_URL |
Loki base URL |
PROMETHEUS_URL |
Prometheus base URL |
PROM_*_QUERY |
Custom PromQL templates |
Service name = Kubernetes Deployment name. Workers use the deployment's label selector for pods and replica sets.
What each worker does (live)
| Worker | Data source |
|---|---|
| Deploy correlator | kubectl get rs for deployment selector |
| Logs | Loki query_range or kubectl logs error lines |
| K8s | Pod status, restarts, warning events |
| Metrics | Prometheus instant queries |
Groq setup
- Get API key from console.groq.com
- Set in
.env:
GROQ_API_KEY=gsk_...
GROQ_API_KEY_FALLBACK=gsk_... # optional second key
GROQ_MODEL=llama-3.3-70b-versatile
When the primary key hits rate limits or quota, requests automatically retry with GROQ_API_KEY_FALLBACK.
Without a key, the heuristic synthesizer still ranks hypotheses from timeline evidence.
Contributing
See CONTRIBUTING.md. Security reports: SECURITY.md.
Next steps
- Post-incident postmortem export
- Scale action implementation
License
Apache License 2.0 — see LICENSE.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file incident_commander-0.1.0.tar.gz.
File metadata
- Download URL: incident_commander-0.1.0.tar.gz
- Upload date:
- Size: 43.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e03878c592485842c402cb117c25fb98554fbb84ef60b8aa1255012ca023095d
|
|
| MD5 |
09a9d8f65b5b43afca1f11bf52d83736
|
|
| BLAKE2b-256 |
2db2e99b01849f084e8cbef8fd58863919de748a3910b9f185fa8068597a1ed8
|
Provenance
The following attestation bundles were made for incident_commander-0.1.0.tar.gz:
Publisher:
publish.yml on sahilleth/incident-commander
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
incident_commander-0.1.0.tar.gz -
Subject digest:
e03878c592485842c402cb117c25fb98554fbb84ef60b8aa1255012ca023095d - Sigstore transparency entry: 2310552654
- Sigstore integration time:
-
Permalink:
sahilleth/incident-commander@be2e53aa938c77b6797ae01a466019a2e05ea24b -
Branch / Tag:
refs/heads/main - Owner: https://github.com/sahilleth
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@be2e53aa938c77b6797ae01a466019a2e05ea24b -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file incident_commander-0.1.0-py3-none-any.whl.
File metadata
- Download URL: incident_commander-0.1.0-py3-none-any.whl
- Upload date:
- Size: 47.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
602ee5a0152e9ad235a1176fb922e03b32554ac35d7fabd634d6270801d99dd4
|
|
| MD5 |
684db4b87eb2bb605dafc936b9b427d2
|
|
| BLAKE2b-256 |
9ff0c53e19eb8288a50b6badd28a8796d2207e2611472de205603166594f3373
|
Provenance
The following attestation bundles were made for incident_commander-0.1.0-py3-none-any.whl:
Publisher:
publish.yml on sahilleth/incident-commander
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
incident_commander-0.1.0-py3-none-any.whl -
Subject digest:
602ee5a0152e9ad235a1176fb922e03b32554ac35d7fabd634d6270801d99dd4 - Sigstore transparency entry: 2310552676
- Sigstore integration time:
-
Permalink:
sahilleth/incident-commander@be2e53aa938c77b6797ae01a466019a2e05ea24b -
Branch / Tag:
refs/heads/main - Owner: https://github.com/sahilleth
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@be2e53aa938c77b6797ae01a466019a2e05ea24b -
Trigger Event:
workflow_dispatch
-
Statement type: