Skip to main content

🐦‍⬛ qaas

A multi-agent QA system that finds real defects — and proves it

Reads your application → finds defects → reproduces each with a failing test → files the ticket → fixes it → reviews the fix → verifies it.

PyPI Python License CI Tests Built on

Quickstart · Your repo · Jira · Architecture


Most "AI QA" tools generate tests. This one behaves like a QA team.

Fifteen agents, each with its own context, tool allowlist and budget, coordinated by a state machine that is ordinary Python — because a model cannot enforce a budget it is itself spending.

      DISCOVERY LOOP                                   REMEDIATION LOOP
 ┌──────────────────────────────────────┐        ┌───────────────────────────┐
 │  MAPPER ─▶ API ─┐          │        │   FIXER ─▶ REVIEWER       │
 │   (system map)   BROWSER ─┴─▶ REPRODUCER ─┼─▶ TRIAGE│    (fix)     (review)     │
 │                 (discover)   (repro) │  (file)│                           │
 └──────────────────────────────┬───────┘        └──────────┬────────────────┘
                                │                           │
                                ▼                           ▼
                           [ TICKET ] ◀──────────────  VERIFIER (verify)

Nothing crosses between the loops except a ticket — which is also the audit trail.


✨ Why this one is different

🧠 The orchestrator is code, not a prompt A model cannot enforce a budget it is spending. Phase ordering, concurrency, retries and the loop breakers live in router.py. That is also why 686 tests run offline, free, with no API key.
🧱 Every agent is its own query() Not subagents of a shared parent. Each gets a real context boundary, an enforceable tool allowlist, and its own cost number.
🔬 Evidence or it did not happen has_evidence() and is_fileable() are methods on the envelope model, not requests in a prompt. An agent cannot talk its way past them.
📊 Measured, not asserted A deliberately buggy demo app ships with a golden ledger of 16 seeded defects + 4 planted non-defects. qaas score reports recall and precision, so a prompt change has a number attached.
🔒 Merge is impossible by construction No merge method exists anywhere. gh pr merge is refused. Pull requests open as drafts. Shipping stays a human decision.
🔍 Every action is on the record 28 kinds of ledger event — every tool call, denial, verdict and escalation. qaas trace reads it back as a timeline.

🤖 The roster

agent layer what it does
🗺️ MAPPER map services, routes, schema, ownership → the system map everything reads
🏛️ ARCHITECT discovery circular deps, layering violations, god modules, dead code
🔌 API discovery API contract drift, authz gaps, error-shape inconsistency
🖱️ BROWSER discovery drives the UI through real journeys
🗄️ DBA discovery schema constraints the code assumes and the database does not enforce
🔒 AUDITOR discovery missing authorization, secrets, vulnerable dependencies, leaks
📡 SOCKET discovery WebSocket auth, reconnect, ordering, backpressure
🧭 GUIDE discovery whether a person can find a feature, not just whether it works
⏱️ LOAD discovery N+1 queries, unindexed hot paths, unbounded results, bundle outliers
🔨 REPRODUCER triage reproduces, minimises, measures flake, commits a failing test
📝 TRIAGE triage dedupes, scores severity, routes, files — the only tracker writer
🔧 FIXER remediation the minimal fix, on a branch
⚖️ REVIEWER remediation adversarial review: APPROVE / REQUEST_CHANGES / ESCALATE
✅ VERIFIER verify re-runs the original test → VERIFIED / NOT_FIXED / REGRESSED
📊 REPORTER reporting what the run found, what recurred, and what it could not reach

ROUTER is the sixteenth. It is the Python state machine rather than an agent — a model cannot enforce a budget it is itself spending.


🚀 Quickstart in 60 seconds

pip install qaas-python
cd ~/code/your-app
qaas init .                          # inspects the repo, writes + activates a target profile
qaas validate                        # config, prompts, allowlists      ← no API call
qaas run --mode pr-check --dry-run   # every agent's exact options      ← no API call

Or point it straight at a URL:

qaas run --repo https://github.com/you/your-app --dry-run

[!TIP] Nothing above contacts an API. --dry-run prints exactly what each agent would receive — model, budget, turn cap, tool allowlist, prompt size.

Auth: the Claude Code CLI if you are signed in, otherwise ANTHROPIC_API_KEY.


🎯 Point it at your repository

qaas init https://github.com/you/your-app     # clones, inspects, writes a profile
qaas init ~/code/your-app --name your-app     # or a local path

init writes .qaas/config/targets/<name>.yaml and activates it. Every value in it is a guess you are expected to correct — it reports what it detected and what it could not find.

The load-bearing field is environment.mode:

mode meaning what agents may do
🔵 none no running instance read code, schema and spec only
🟡 external already running (staging, dev server) read and exercise, never reset
🟢 compose qaas owns the lifecycle seed, reset, tear down

[!NOTE] none is a perfectly good place to start, and where most first runs against a real repository begin. Findings stay honest about it: an agent that could not observe a behaviour says so and lowers its confidence.

Credentials never live in the profile. It names environment variables:

auth:
  mode: login
  login_endpoint: POST /api/v1/session
  roles:
    admin:  { username: qa-admin@example.com,  password_env: APP_ADMIN_PASSWORD }
    viewer: { username: qa-viewer@example.com, password_env: APP_VIEWER_PASSWORD }

🎫 File into Jira

export JIRA_BASE_URL=https://you.atlassian.net
export JIRA_EMAIL=you@example.com
export JIRA_API_TOKEN=...            # an API token, not a password
export JIRA_PROJECT_KEY=QA

QAAS_TRACKER=jira qaas tracker-check  # auth, project, permissions, workflow — creates nothing
QAAS_TRACKER=jira qaas run --mode nightly

tracker-check validates credentials, confirms the project and issue type exist, maps your workflow statuses, and prints the exact JSON it would POST. It creates nothing.

Tired of four exports in every new shell? Put them in a .env.qaas/.env is already gitignored — and every command reads it. Anything you export wins over the file, so a stale .env can never redirect a run.

  • 🔁 Dedupe across runs — tickets carry a qaas-fp-<fingerprint> label, so the next run recognises an already-filed defect and increments its occurrence count instead of filing again.
  • 🔐 Security findings are refused unless JIRA_SECURITY_PROJECT_KEY is set. A vulnerability in a project the whole company can read is a disclosure with no undo.

[!TIP] Keep the committed backend local — it writes tickets as JSON under .qaas/tickets/ so you can read what would be filed. Switch per shell with QAAS_TRACKER=jira.

📌 A view per repository, made for you

Point it at a new repository and it provisions that repository's own Jira view before the first agent starts — so there is something to watch during the run, not a report afterwards.

$ QAAS_TRACKER=jira qaas run --repo https://github.com/acme/checkout.git --mode nightly
target: checkout (none)
filter created — https://you.atlassian.net/issues/?filter=10001
every ticket from this run carries the label repo-checkout

Every ticket the system files carries repo-<target>, stamped in code rather than asked of an agent. A saved filter over exactly that label is the per-repository view, and on a company-managed project a board is built over the filter too.

qaas board                  # find or create this target's view
qaas board --no-create      # show the label and JQL, touch nothing

[!NOTE] Not a project per repository. Creating a Jira project needs administrator rights a bot account rarely has, and a project per repository is unmanageable by the tenth one. A filter needs no special grant.

Not always a board, either. Team-managed (next-gen) projects own their own board and cannot have a second one built over a filter — Jira's API will happily create one and give it no page in the UI. So the project's style is checked first, and on a team-managed project you get the filter alone. You are told which you got, and the link always opens. Point JIRA_PROJECT_KEY at a company-managed project and you get a real board per repository, with To Do / In Progress / Done.

Run it twice on the same repository and it reuses what is there. A run is never failed over this: a run that found nine defects and could not make a view has still done its job.


🔌 Bring your own MCP servers

Declare them in system.yaml. No Python to edit.

mcp_servers:
  house-lint:
    type: stdio
    command: ./tools/lint-mcp
    args: ["--strict"]
    env: { LINT_TOKEN: "${ACME_LINT_TOKEN}" }   # from the environment, never a literal
  remote-docs:
    type: http
    url: https://mcp.example/v1

Then add the name to any agent's mcp_servers: list. Declaring a server grants nothing — an agent receives it only by naming it.

$ qaas validate
                       Declared MCP servers
┃ name        ┃ kind  ┃ what it runs              ┃ used by      ┃
│ house-lint  │ stdio │ ./tools/lint-mcp --strict │ MAPPER │
│ remote-docs │ http  │ https://mcp.example/v1    │ nobody       │

[!WARNING] A server's tools are allowed wholesale once an agent names it. qaas checks that the agent declared the server; it cannot inspect what a third-party server's tools actually do. A server you declare is a server you trust.

There is deliberately no in-process Python server type — a module path from a config file would mean importing arbitrary code into the process holding your Anthropic, Jira and GitHub credentials. Wrap it in a stdio entry point instead.

Two limits worth knowing: 6 MCP servers per agent (tool-selection accuracy falls off past ~7), and FIXER is already at the cap — the agent people most want to extend.


✏️ Make the prompts yours

Agents are a prompt plus a YAML file. Both are yours to change.

qaas prompts list              # which prompt is in force, and where it came from
qaas prompts eject API     # copy it to .qaas/prompts/ and edit freely
qaas prompts diff              # what you changed vs. what shipped

Prefer adding to replacing — drop a API.append.md beside it:

## Our conventions
- Never file a finding without a curl reproduction.
- Treat any 500 on a write path as blocker severity.

That block is inserted between the agent's prompt and the shared house rules, so you keep receiving improvements to the base prompt instead of forking it forever.


🔍 Full traceability

Every tool call, denial, verdict and escalation is on the record.

$ qaas trace run-20260908T182034-c6ed26
    t+  agent         kind            detail
    0s  -             run_started     mode=nightly  agents=[7]
    0s  MAPPER  agent_started   model=claude-sonnet-5
    4s  MAPPER  tool_call ×34   Read×25, Glob×6, ToolSearch×2
    6s  MAPPER  denial          tool=Bash  reason=Bash is not in MAPPER's
                                      tool allowlist (Read, Grep, Glob).
  146s  MAPPER  system_map      version=20260907T233530  sections=[12]
  156s  MAPPER  agent_finished  subtype=success  num_turns=45
qaas trace <run-id> --follow                       # watch a run as it happens
qaas trace <run-id> --quiet                        # decisions only, no file reads
qaas trace <run-id> --agent verifier --kind verdict   # filter
qaas trace <run-id> --json                         # export
qaas show <run-id>                                 # mode, commit, tickets, escalations
qaas runs                                          # everything that ever ran

Watch a run live. --follow tails the ledger of a run in progress — start it in a second terminal the moment a run begins, or even before, and it waits.

$ qaas trace run-20260909T163240-83b11c --follow --quiet
following run-20260909T163240-83b11c — ctrl-c to stop
16:32:40 -            run_started      mode=pr-check  agents=[8]  target_sha=da19f406
16:32:40 MAPPER agent_started    model=claude-sonnet-5
16:32:46 MAPPER denial           tool=Bash  reason=Bash is not in MAPPER's
                                       tool allowlist (Read, Grep, Glob).
16:36:11 ARCHITECT     envelope         severity=major  domain=architecture
16:41:03 TRIAGE        ticket           action=created  key=QA-118  severity=major

Drop --quiet to see every file the agents read, one line each. Combine with --agent and --kind to watch one agent, or only the verdicts.

Runs are pinned to the commit of the target they examined, so a finding can be replayed against the tree that produced it.


🛡️ Safety rails

Enforced in code, not requested in a prompt:

rail what it does
📁 Path scoping Writes checked against that agent's write_paths. Most agents cannot write at all.
🌿 Branch scoping Git writes must match the agent's patterns (qa/repro/*, fix/*). main and force-push refused outright.
Forbidden classes Migrations, auth, payment, secrets, infrastructure, CI — stop at a human however small the change looks.
🎟️ Ticket rate limit Over the per-run cap the call is denied and the router escalates rather than filing.
🧪 Immutable test The agent fixing a defect may not edit the test that defines it.
🚫 No filesystem settings setting_sources=[] — a repository qaas is inspecting cannot inject settings, hooks or MCP servers into the process running it.

Denials return a reason and are logged; they never kill the turn. The agent reads the refusal and adapts.


📈 How you know it works

qaas ships with a deliberately buggy demo app and a golden ledger recording every defect seeded into it — plus planted non-defects: correct-but-suspicious code that a careless reader would report.

cd target-app && docker compose up -d
qaas run --mode nightly && qaas score

qaas score reports recall, precision and severity agreement against that ledger. That matters more than any number this page could print: it means a change to a prompt, a threshold or a model has a measurement attached rather than an opinion.

[!CAUTION] Seeded defects are easier than real ones, and this system was calibrated against them. A benchmark shows the loop works end to end and does not spray false positives. It does not show that it will find the hard bug in your codebase. Run it on something you know well and judge it on what it finds there.


🤝 Contributing

git clone https://github.com/allaabdella2-us/qa-multi-agent-system
cd qa-multi-agent-system
uv venv && uv pip install -e ".[dev]"

pytest                 # 686 tests, offline, free — keep it that way
pytest -m docker       # needs: cd target-app && docker compose up -d
qaas validate

The llm, docker, github and jira markers are deselected by default. A CI run that costs money is a CI run people switch off.

[!IMPORTANT] If you change a seeded defect in target-app/, retire its ledger entry with fixed_in: in the same commit — a stale ledger silently corrupts every score.

New to the codebase? ARCHITECTURE.md explains the entry point, the five phases, what moves between agents, and what the guardrails stop.


MIT licensed · LICENSE

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

qaas_python-0.0.1.tar.gz (412.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

qaas_python-0.0.1-py3-none-any.whl (301.1 kB view details)

Uploaded Python 3

File details

Details for the file qaas_python-0.0.1.tar.gz.

File metadata

  • Download URL: qaas_python-0.0.1.tar.gz
  • Upload date:
  • Size: 412.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.2

File hashes

Hashes for qaas_python-0.0.1.tar.gz
Algorithm Hash digest
SHA256 8ca1968705174a2280d4c3beb2379f5a5057eaf631313644607a7218ea46f1c1
MD5 9e1cb6356509a7cb1ffc24b249c1c58d
BLAKE2b-256 1f6e6c53af31bc7910ed91e1370ec4021b6b9d57201127d34b2732230d78231d

See more details on using hashes here.

File details

Details for the file qaas_python-0.0.1-py3-none-any.whl.

File metadata

  • Download URL: qaas_python-0.0.1-py3-none-any.whl
  • Upload date:
  • Size: 301.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.2

File hashes

Hashes for qaas_python-0.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 baca8c522937d59df2f4b970faaff999183871df4e29223b4ce9ffd6d4d174c8
MD5 299038cb12870b1ee04d1d2d09aed7f2
BLAKE2b-256 9ccc983baafb84e309e47b75fa885287b677b6c4be1ec88154a65ce5aa1bc74d

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.0.1 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page