🐦⬛ qaas
A multi-agent QA system that finds real defects — and proves it
Reads your application → finds defects → reproduces each with a failing test → files the ticket → fixes it → reviews the fix → verifies it.
Quickstart · Your repo · Jira · Architecture
Most "AI QA" tools generate tests. This one behaves like a QA team.
Fifteen agents, each with its own context, tool allowlist and budget, coordinated by a state machine that is ordinary Python — because a model cannot enforce a budget it is itself spending.
DISCOVERY LOOP REMEDIATION LOOP
┌──────────────────────────────────────┐ ┌───────────────────────────┐
│ MAPPER ─▶ API ─┐ │ │ FIXER ─▶ REVIEWER │
│ (system map) BROWSER ─┴─▶ REPRODUCER ─┼─▶ TRIAGE│ (fix) (review) │
│ (discover) (repro) │ (file)│ │
└──────────────────────────────┬───────┘ └──────────┬────────────────┘
│ │
▼ ▼
[ TICKET ] ◀────────────── VERIFIER (verify)
Nothing crosses between the loops except a ticket — which is also the audit trail.
✨ Why this one is different
| 🧠 The orchestrator is code, not a prompt | A model cannot enforce a budget it is spending. Phase ordering, concurrency, retries and the loop breakers live in router.py. That is also why 686 tests run offline, free, with no API key. |
🧱 Every agent is its own query() |
Not subagents of a shared parent. Each gets a real context boundary, an enforceable tool allowlist, and its own cost number. |
| 🔬 Evidence or it did not happen | has_evidence() and is_fileable() are methods on the envelope model, not requests in a prompt. An agent cannot talk its way past them. |
| 📊 Measured, not asserted | A deliberately buggy demo app ships with a golden ledger of 16 seeded defects + 4 planted non-defects. qaas score reports recall and precision, so a prompt change has a number attached. |
| 🔒 Merge is impossible by construction | No merge method exists anywhere. gh pr merge is refused. Pull requests open as drafts. Shipping stays a human decision. |
| 🔍 Every action is on the record | 28 kinds of ledger event — every tool call, denial, verdict and escalation. qaas trace reads it back as a timeline. |
🤖 The roster
| agent | layer | what it does |
|---|---|---|
| 🗺️ MAPPER | map | services, routes, schema, ownership → the system map everything reads |
| 🏛️ ARCHITECT | discovery | circular deps, layering violations, god modules, dead code |
| 🔌 API | discovery | API contract drift, authz gaps, error-shape inconsistency |
| 🖱️ BROWSER | discovery | drives the UI through real journeys |
| 🗄️ DBA | discovery | schema constraints the code assumes and the database does not enforce |
| 🔒 AUDITOR | discovery | missing authorization, secrets, vulnerable dependencies, leaks |
| 📡 SOCKET | discovery | WebSocket auth, reconnect, ordering, backpressure |
| 🧭 GUIDE | discovery | whether a person can find a feature, not just whether it works |
| ⏱️ LOAD | discovery | N+1 queries, unindexed hot paths, unbounded results, bundle outliers |
| 🔨 REPRODUCER | triage | reproduces, minimises, measures flake, commits a failing test |
| 📝 TRIAGE | triage | dedupes, scores severity, routes, files — the only tracker writer |
| 🔧 FIXER | remediation | the minimal fix, on a branch |
| ⚖️ REVIEWER | remediation | adversarial review: APPROVE / REQUEST_CHANGES / ESCALATE |
| ✅ VERIFIER | verify | re-runs the original test → VERIFIED / NOT_FIXED / REGRESSED |
| 📊 REPORTER | reporting | what the run found, what recurred, and what it could not reach |
ROUTER is the sixteenth. It is the Python state machine rather than an agent — a model cannot enforce a budget it is itself spending.
🚀 Quickstart in 60 seconds
pip install qaas-python
cd ~/code/your-app
qaas init . # inspects the repo, writes + activates a target profile
qaas validate # config, prompts, allowlists ← no API call
qaas run --mode pr-check --dry-run # every agent's exact options ← no API call
Or point it straight at a URL:
qaas run --repo https://github.com/you/your-app --dry-run
[!TIP] Nothing above contacts an API.
--dry-runprints exactly what each agent would receive — model, budget, turn cap, tool allowlist, prompt size.
Auth: the Claude Code CLI if you are signed in, otherwise ANTHROPIC_API_KEY.
🎯 Point it at your repository
qaas init https://github.com/you/your-app # clones, inspects, writes a profile
qaas init ~/code/your-app --name your-app # or a local path
init writes .qaas/config/targets/<name>.yaml and activates it. Every value
in it is a guess you are expected to correct — it reports what it detected and
what it could not find.
The load-bearing field is environment.mode:
| mode | meaning | what agents may do |
|---|---|---|
🔵 none |
no running instance | read code, schema and spec only |
🟡 external |
already running (staging, dev server) | read and exercise, never reset |
🟢 compose |
qaas owns the lifecycle | seed, reset, tear down |
[!NOTE]
noneis a perfectly good place to start, and where most first runs against a real repository begin. Findings stay honest about it: an agent that could not observe a behaviour says so and lowers its confidence.
Credentials never live in the profile. It names environment variables:
auth:
mode: login
login_endpoint: POST /api/v1/session
roles:
admin: { username: qa-admin@example.com, password_env: APP_ADMIN_PASSWORD }
viewer: { username: qa-viewer@example.com, password_env: APP_VIEWER_PASSWORD }
🎫 File into Jira
export JIRA_BASE_URL=https://you.atlassian.net
export JIRA_EMAIL=you@example.com
export JIRA_API_TOKEN=... # an API token, not a password
export JIRA_PROJECT_KEY=QA
QAAS_TRACKER=jira qaas tracker-check # auth, project, permissions, workflow — creates nothing
QAAS_TRACKER=jira qaas run --mode nightly
tracker-check validates credentials, confirms the project and issue type exist,
maps your workflow statuses, and prints the exact JSON it would POST. It
creates nothing.
Tired of four exports in every new shell? Put them in a .env — .qaas/.env is
already gitignored — and every command reads it. Anything you export wins over
the file, so a stale .env can never redirect a run.
- 🔁 Dedupe across runs — tickets carry a
qaas-fp-<fingerprint>label, so the next run recognises an already-filed defect and increments its occurrence count instead of filing again. - 🔐 Security findings are refused unless
JIRA_SECURITY_PROJECT_KEYis set. A vulnerability in a project the whole company can read is a disclosure with no undo.
[!TIP] Keep the committed backend
local— it writes tickets as JSON under.qaas/tickets/so you can read what would be filed. Switch per shell withQAAS_TRACKER=jira.
📌 A view per repository, made for you
Point it at a new repository and it provisions that repository's own Jira view before the first agent starts — so there is something to watch during the run, not a report afterwards.
$ QAAS_TRACKER=jira qaas run --repo https://github.com/acme/checkout.git --mode nightly
target: checkout (none)
filter created — https://you.atlassian.net/issues/?filter=10001
every ticket from this run carries the label repo-checkout
Every ticket the system files carries repo-<target>, stamped in code rather
than asked of an agent. A saved filter over exactly that label is the
per-repository view, and on a company-managed project a board is built over
the filter too.
qaas board # find or create this target's view
qaas board --no-create # show the label and JQL, touch nothing
[!NOTE] Not a project per repository. Creating a Jira project needs administrator rights a bot account rarely has, and a project per repository is unmanageable by the tenth one. A filter needs no special grant.
Not always a board, either. Team-managed (next-gen) projects own their own board and cannot have a second one built over a filter — Jira's API will happily create one and give it no page in the UI. So the project's style is checked first, and on a team-managed project you get the filter alone. You are told which you got, and the link always opens. Point
JIRA_PROJECT_KEYat a company-managed project and you get a real board per repository, with To Do / In Progress / Done.
Run it twice on the same repository and it reuses what is there. A run is never failed over this: a run that found nine defects and could not make a view has still done its job.
🔌 Bring your own MCP servers
Declare them in system.yaml. No Python to edit.
mcp_servers:
house-lint:
type: stdio
command: ./tools/lint-mcp
args: ["--strict"]
env: { LINT_TOKEN: "${ACME_LINT_TOKEN}" } # from the environment, never a literal
remote-docs:
type: http
url: https://mcp.example/v1
Then add the name to any agent's mcp_servers: list. Declaring a server grants
nothing — an agent receives it only by naming it.
$ qaas validate
Declared MCP servers
┃ name ┃ kind ┃ what it runs ┃ used by ┃
│ house-lint │ stdio │ ./tools/lint-mcp --strict │ MAPPER │
│ remote-docs │ http │ https://mcp.example/v1 │ nobody │
[!WARNING] A server's tools are allowed wholesale once an agent names it. qaas checks that the agent declared the server; it cannot inspect what a third-party server's tools actually do. A server you declare is a server you trust.
There is deliberately no in-process Python server type — a module path from a config file would mean importing arbitrary code into the process holding your Anthropic, Jira and GitHub credentials. Wrap it in a stdio entry point instead.
Two limits worth knowing: 6 MCP servers per agent (tool-selection accuracy falls off past ~7), and FIXER is already at the cap — the agent people most want to extend.
✏️ Make the prompts yours
Agents are a prompt plus a YAML file. Both are yours to change.
qaas prompts list # which prompt is in force, and where it came from
qaas prompts eject API # copy it to .qaas/prompts/ and edit freely
qaas prompts diff # what you changed vs. what shipped
Prefer adding to replacing — drop a API.append.md beside it:
## Our conventions
- Never file a finding without a curl reproduction.
- Treat any 500 on a write path as blocker severity.
That block is inserted between the agent's prompt and the shared house rules, so you keep receiving improvements to the base prompt instead of forking it forever.
🔍 Full traceability
Every tool call, denial, verdict and escalation is on the record.
$ qaas trace run-20260908T182034-c6ed26
t+ agent kind detail
0s - run_started mode=nightly agents=[7]
0s MAPPER agent_started model=claude-sonnet-5
4s MAPPER tool_call ×34 Read×25, Glob×6, ToolSearch×2
6s MAPPER denial tool=Bash reason=Bash is not in MAPPER's
tool allowlist (Read, Grep, Glob).
146s MAPPER system_map version=20260907T233530 sections=[12]
156s MAPPER agent_finished subtype=success num_turns=45
qaas trace <run-id> --follow # watch a run as it happens
qaas trace <run-id> --quiet # decisions only, no file reads
qaas trace <run-id> --agent verifier --kind verdict # filter
qaas trace <run-id> --json # export
qaas show <run-id> # mode, commit, tickets, escalations
qaas runs # everything that ever ran
Watch a run live. --follow tails the ledger of a run in progress — start it
in a second terminal the moment a run begins, or even before, and it waits.
$ qaas trace run-20260909T163240-83b11c --follow --quiet
following run-20260909T163240-83b11c — ctrl-c to stop
16:32:40 - run_started mode=pr-check agents=[8] target_sha=da19f406
16:32:40 MAPPER agent_started model=claude-sonnet-5
16:32:46 MAPPER denial tool=Bash reason=Bash is not in MAPPER's
tool allowlist (Read, Grep, Glob).
16:36:11 ARCHITECT envelope severity=major domain=architecture
16:41:03 TRIAGE ticket action=created key=QA-118 severity=major
Drop --quiet to see every file the agents read, one line each. Combine with
--agent and --kind to watch one agent, or only the verdicts.
Runs are pinned to the commit of the target they examined, so a finding can be replayed against the tree that produced it.
🛡️ Safety rails
Enforced in code, not requested in a prompt:
| rail | what it does |
|---|---|
| 📁 Path scoping | Writes checked against that agent's write_paths. Most agents cannot write at all. |
| 🌿 Branch scoping | Git writes must match the agent's patterns (qa/repro/*, fix/*). main and force-push refused outright. |
| ⛔ Forbidden classes | Migrations, auth, payment, secrets, infrastructure, CI — stop at a human however small the change looks. |
| 🎟️ Ticket rate limit | Over the per-run cap the call is denied and the router escalates rather than filing. |
| 🧪 Immutable test | The agent fixing a defect may not edit the test that defines it. |
| 🚫 No filesystem settings | setting_sources=[] — a repository qaas is inspecting cannot inject settings, hooks or MCP servers into the process running it. |
Denials return a reason and are logged; they never kill the turn. The agent reads the refusal and adapts.
📈 How you know it works
qaas ships with a deliberately buggy demo app and a golden ledger recording
every defect seeded into it — plus planted non-defects: correct-but-suspicious
code that a careless reader would report.
cd target-app && docker compose up -d
qaas run --mode nightly && qaas score
qaas score reports recall, precision and severity agreement against that
ledger. That matters more than any number this page could print: it means a
change to a prompt, a threshold or a model has a measurement attached rather
than an opinion.
[!CAUTION] Seeded defects are easier than real ones, and this system was calibrated against them. A benchmark shows the loop works end to end and does not spray false positives. It does not show that it will find the hard bug in your codebase. Run it on something you know well and judge it on what it finds there.
🤝 Contributing
git clone https://github.com/allaabdella2-us/qa-multi-agent-system
cd qa-multi-agent-system
uv venv && uv pip install -e ".[dev]"
pytest # 686 tests, offline, free — keep it that way
pytest -m docker # needs: cd target-app && docker compose up -d
qaas validate
The llm, docker, github and jira markers are deselected by default.
A CI run that costs money is a CI run people switch off.
[!IMPORTANT] If you change a seeded defect in
target-app/, retire its ledger entry withfixed_in:in the same commit — a stale ledger silently corrupts every score.
New to the codebase? ARCHITECTURE.md explains the entry
point, the five phases, what moves between agents, and what the guardrails stop.
MIT licensed · LICENSE
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file qaas_python-0.0.1.tar.gz.
File metadata
- Download URL: qaas_python-0.0.1.tar.gz
- Upload date:
- Size: 412.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8ca1968705174a2280d4c3beb2379f5a5057eaf631313644607a7218ea46f1c1
|
|
| MD5 |
9e1cb6356509a7cb1ffc24b249c1c58d
|
|
| BLAKE2b-256 |
1f6e6c53af31bc7910ed91e1370ec4021b6b9d57201127d34b2732230d78231d
|
File details
Details for the file qaas_python-0.0.1-py3-none-any.whl.
File metadata
- Download URL: qaas_python-0.0.1-py3-none-any.whl
- Upload date:
- Size: 301.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
baca8c522937d59df2f4b970faaff999183871df4e29223b4ce9ffd6d4d174c8
|
|
| MD5 |
299038cb12870b1ee04d1d2d09aed7f2
|
|
| BLAKE2b-256 |
9ccc983baafb84e309e47b75fa885287b677b6c4be1ec88154a65ce5aa1bc74d
|