qa-orchestrator
Inject faults, drive walks, and check the audit tool reached the verdict it should have.
bmc-sensor-audit judges firmware:
it diffs what an OpenBMC declaration promises against what the machine reports. This
drives both sides. It perturbs a machine in ways whose correct verdict is known in
advance — a sensor removed, one switched off, a reading frozen — and checks the tool
reached it.
Why this is a separate program
The audit tool is the judge of the outcome, so it cannot also be the injector. A referee that shipped its own fault injection would be certifying itself: every scenario would be written against the behaviour the tool happens to have, and a tool that stopped detecting something would quietly stop being asked about it.
So the injector lives outside and consumes only the referee's published surfaces —
exit codes, the JSON report, the attestation. referee.py does not import
bmc_sensor_audit, and a test asserts that by reading the file. If a scenario needs
something the published surface does not carry, that is a feature request against
the tool, not a reason to reach past it.
A scenario
format: qa-scenario/1
backend: mock
mode: coverage
config: fixtures/board.json # resolved beside this file, not from the cwd
machine:
sensors:
- {name: Inlet, reading: 21.0, upper_critical: 80, upper_warning: 70}
- {name: Outlet, reading: 27.5, upper_critical: 80, upper_warning: 70}
phases:
- walks: 1
expect:
audit: {exit: 0}
- action: {remove: Outlet}
walks: 1
expect:
audit:
exit: 1
finding: "not reported by the machine at all"
names: [Outlet]
not_names: [Inlet]
firmware: {Outlet: absent, Inlet: reading}
Released — 0.2.0, tagged v0.2.0, Apache-2.0, on PyPI as qa-orchestrator.
0.2.0 raises the referee floor to 0.2.0 and adds --version. The floor is
the point: from bmc-sensor-audit 0.2.0 a command that asks to verify and not
to verify at once is refused rather than run unverified, and a harness that
drives a referee should not be what pins it below its own security fix. This
package's own behaviour is unchanged. The suite now runs on every Python it
claims.
0.1.2 changes nothing this package does. It carries the repository's publication-hygiene tooling: the rules now run over commit messages as well as files, and a pre-commit hook refuses a commit whose staged content it has not read. The only differences a reader will find in the installed distribution are this paragraph and the version number. Nothing here obliges an upgrade.
0.1.1 makes the scenario schema's fail action work. It was documented from
the start and had never run: the referee's capture exits 2 both when it
cannot reach the machine and when it reached the machine and a subtree
answered with an error, and this harness raised on either — so a scenario that
induced a partial walk aborted before the referee could be asked anything.
Every walk also carries a content handle now. Needs bmc-sensor-audit 0.1.1.
pip install qa-orchestrator
qa-orchestrator check <scenario.yaml> # parse only; needs no machine
qa-orchestrator run <scenario.yaml>
The two worked scenarios are named with a placeholder above rather than by
filename, because a wheel carries only what lives under the package directory: a
pip install gives you the command and not the examples. They ship in this
repository and in the sdist, under scenarios/. From a clone,
pip install -e . puts the same script on the path and the scenarios beside it.
audit and firmware are different claims and both are worth making. The
sensor is gone and the tool noticed it is gone are separate facts. A scenario
that could only express the second could never tell a broken injector from a blind
referee — the run would be green either way.
not_names is half of every detection claim. Naming the frozen sensor shows
the tool found something; showing the sensor beside it was not named is what
makes that evidence of detection rather than of a check that flags everything.
The format refuses what it cannot run
A phase with no walks, an expect block that sets nothing, an action verb this
build does not implement, a drive with fewer values than walks, exit_code where
exit was meant — all rejected at parse time with the phase named. Each of them
would otherwise produce a run that executed, reported clean, and tested nothing,
which is indistinguishable from a real pass at the exit code.
Three tiers, and one of them says so
| Backend | What it is | State |
|---|---|---|
mock |
the audit tool's own MockBMC, in process |
working; needs no hardware |
qemu |
attaches to a running instance, injects over QMP | wire format tested, integration not |
testbed |
relay boards and real fans | not implemented, and refuses |
The tiers exist so one scenario runs at increasing cost and realism without being
rewritten. That only holds if a tier which cannot do something refuses instead of
approximating it. testbed raises at construction and its refusal lists what a
real implementation needs — a specification, not an apology. A stub that accepted
injections and did nothing would report that real fans were pulled having touched
nothing, and every phase after it would judge an unperturbed machine and pass.
qemu attaches to an already-running instance whose QMP socket the scenario names.
It does not boot one. Owning the boot recipe — image build id, machine type, FRU
provisioning — is real work that is not done here, and claiming it would be worse
than not having it. The QMP conversation is tested against a fake socket, so the
greeting, handshake, framing and error path are covered; no real QEMU has run it.
Exit codes
The same three-valued contract the audit tool uses, because this sits in the same pipeline and a fourth vocabulary at this layer is one more thing for a gate to get wrong:
0 |
every expectation held |
1 |
a verdict disagreed with the scenario |
2 |
the run could not be completed |
2 never reads as clean. A scenario that could not reach its backend has judged
nothing. Could-not-complete outranks disagreement, because a run that stopped early
has not evaluated the phases it never reached.
A run also reports how many phases asserted anything. A scenario of phases that assert nothing is green and worthless, and the exit code cannot tell that from a real pass — so it is said in words instead.
The acceptance scenario
scenarios/stuck-at.yaml reproduces the experiment the audit tool already ships: a
sensor driven to a new value before each of twelve walks, then left alone for
sixteen. The engine is silent while everything moves and names exactly the frozen
sensor once one freezes, with a control sensor still moving beside it.
That experiment is the acceptance test for this harness, because its outcome is
already known. If the orchestrator cannot express it, the DSL is wrong — and at
first it could not: drive took one sensor and a phase may carry only one action,
so there was no way to keep a control moving beside the frozen one. The general
drive: {sensors: {...}} form exists because this file needed it.
Two things it does not show. Freezing a reading through a mock is an experiment, not a sensor failing on its own. And the declaration it runs against carries thresholds deliberately: a sensor declaring none is excluded from the liveness model entirely, so a thresholdless board produces an empty model and a clean run that checked nothing.
A partial walk is evidence, not a failed run
The audit tool's capture exits 2 for two different facts: it could not reach
the machine, and it reached the machine while one subtree answered with an error.
The second is a walk the tool writes and keeps on purpose, because knowing
which subtree failed is the point.
This harness used to raise on any non-zero, so the scenario schema's fail
action — make a subtree answer with an HTTP status — aborted the run before the
referee could be asked anything. No shipped scenario used it and no test
exercised it, so it had never once worked. scenarios/partial-walk.yaml is that
scenario, and it now runs.
The fix is to judge the file rather than the exit code. validate-walk says
whether what was written is a well-formed walk/1 — a question about the artifact
rather than about the run — and the walk's own error list says whether the machine
answered for all of it. A capture that produced no readable walk is still a
failure and still stops the run.
What it is really testing is the referee's honesty about not knowing. A subtree it
could not read is indistinguishable from a subtree with nothing in it, so the tool
must not report absence on a partial walk. Exit 2 is it saying so, and a
tool that rendered a network error as two sensors missing would be worse than
useless on a line.
Every walk gets a content handle
evidence:
walk 001 complete sha256:81422480ba090695df9f8dabb6aba4cd849048f137ca5707558fbcd03e43950e
walk 002 PARTIAL sha256:19cb099c107403ad68a8ce0c29cfc6d0a6724dde688ad815347a8a2192e91835
Printed by every run, including one that stopped early. A clean run deletes its workdir, so the run that needs no further explanation is exactly the one whose walks are gone — the handles outlive them.
They are the tool's own, from capture --print-digest: a SHA-256 over the file's
bytes, which sha256sum reproduces in any language. Anyone who kept a walk can
match it with no tooling and nothing to trust, and a walk that does not match is
not the walk that was judged. This program reads the handle rather than computing
one, because two definitions of one number is how the two come to disagree.
Not built yet
- The redundancy scenario against the audit tool's supplemental template.
- The
qemuboot recipe; this build attaches to a running instance. - The
testbedtier, which needs a lab.
Apache-2.0.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file qa_orchestrator-0.2.0.tar.gz.
File metadata
- Download URL: qa_orchestrator-0.2.0.tar.gz
- Upload date:
- Size: 47.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
14ccb8532f76ba8c1bcecbf609188955bcb45ee5a859af0e9dc14444eff87fc2
|
|
| MD5 |
37a5b4e3b5261e602bed7ef748a9a44f
|
|
| BLAKE2b-256 |
66c12c8f5b2f7a34415311f1adf9b4ff5598886c3211e0907438f3deeebab15f
|
File details
Details for the file qa_orchestrator-0.2.0-py3-none-any.whl.
File metadata
- Download URL: qa_orchestrator-0.2.0-py3-none-any.whl
- Upload date:
- Size: 34.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bc4c12d7e75032eb8883575284799fc60d0ec8c2d701554508c8dbcf8aeca4b9
|
|
| MD5 |
d0d4bb7fbe049c74b84aa8fc17ddbd9d
|
|
| BLAKE2b-256 |
9c41ed401696d1658e712f8e771c5f23bf52897086a4ac06ef5ff011cc803a9d
|