Reproof
Reproof is a self-hosted mobile QA platform with no STF dependency. It records QA issues as video, user actions, and initial conditions; replays them deterministically on Android and iOS; applies AI-generated repairs; and re-verifies a repaired candidate against the same approved original recording.
For wheel installation and running outside a development checkout, see the installation guide. Progress on the general-app product path is tracked in the current execution plan (the plan document is at r51; see HANDOFF.md for physical-device progress since then).
Status
What is proven and what is still open:
- The sample path (record, reproduce, repair, re-verify) runs end to end on
Android devices, iOS Simulator, and a physical iPhone. In the counter defect
(one tap on
Addraises the count by 2), the original reproduced 3/3 and the repaired candidate passed 3/3; these results include real Claude-generated patches. - The shared QA path (recording, video, issue packages, approved reproduction specifications) is verified with two real worker processes driving a synthetic app on one Mac, plus real MP4 and browser capture (guide).
- Protected repair of declared product files is implemented (G9 path); the 92 G9 software checks pass.
- On a physical iPhone the protected lifecycle has completed end to end:
device qualification, service composition, issue recording, specification
approval, 3/3 replay reproduction, and a
verifiedAI repair — checked by an independent observer process against fixture-prepared device state. - The same lifecycle completed against a real product app (not the bundled
sample): the fix candidate verified, and a candidate that also performed
network egress was rejected
egress_violation— fail-closed on device byte counters with packet-level capture evidence, under both Wi-Fi and cellular-only connectivity. - Still open: the isolated VM build lane (the verified lane is host-build — builds run directly on the host and are not reported as isolated), and two-Mac acceptance.
Product direction
The goal is a self-hosted mobile test platform: remote manual control, automation, and farm operation on one common session layer, connected to recording and AI reproduce/repair. See the platform direction and design, the comparison with official documentation, and record/replay features.
Automatic observation of general UIKit and Android Views apps is prepared by declaring public build inputs and UI IDs (UIKit, Android Views). The observation profile preserves the original, stays separate from fixtures and repair policy, and is attached to general-app sessions on the shared service.
Scope of the sample commands
The commands below target only the bundled Android sample app and the iOS
Simulator sample app. They record a defect where Add increments the count by
2, then re-verify a candidate built from the same buggy variant with only the
business logic patched. The Python host runs with no external packages.
iOS Simulator
The Swift recording SDK, a UIKit sample, and an XCUITest batch runner form one pipeline. Across three cases — counter, duplicate submit, initialization failure — real Claude patches passed the existing regression tests, with the original reproducing 3/3 and the repair passing 3/3. The same three cases were verified on a physical iPhone with real Claude repairs, original 3/3 reproduction, and candidate 3/3 verification. The counter case is also connected end to end from live recording to AI repair (live repair).
bash scripts/ios-demo.sh <BOOTED_SIMULATOR_UUID> artifacts/my-ios-demo
See the iOS runbook and results, the three real-AI bug cases, and the QA report for failure paths. The older Android bundle format (v1) and iOS bundle format (v2) are verified separately.
Android quick start
Requirements: Python 3.11+, JDK 17, Android SDK 35 with build-tools and platform-tools, Gradle 8.14.5 (wrapper included), and an Android device running API 26+ with USB debugging enabled. The build below assumes a prepared offline cache; a fresh environment needs the Android plugin, Kotlin, and JUnit dependencies fetched once.
cd <reproof clone path>
python3 -m reproof doctor
python3 -m reproof build --receipt artifacts/build.json
build builds the sample app and the ID-based input driver together and stores
source/APK hashes in the receipt. Default tool paths are discovered
automatically in standard macOS install locations; elsewhere pass --gradle,
--java-home, --sdk-home, or set JAVA_HOME and ANDROID_HOME.
Running record with --scripted performs synthetic QA actions and produces
the first bundle automatically:
python3 -m reproof record \
--apk android/sample/build/outputs/apk/buggy/debug/sample-buggy-debug.apk \
--driver-apk android/driver/build/outputs/apk/debug/driver-debug.apk \
--receipt artifacts/build.json \
--scripted --output artifacts/qa-bundle
python3 -m reproof validate artifacts/qa-bundle
python3 -m reproof replay artifacts/qa-bundle --output artifacts/original-runs
For manual QA, drop --scripted. In the sample app, type QA, tap Add once,
then press Enter in the terminal — the session freezes on the Report action.
Record-only buttons never enter replay events. With several devices attached,
pick one with --serial.
The dedicated sample package io.reproof.sample is reset before each run.
Run with the device screen on and unlocked.
Repair loop
Run the full pipeline offline with the prepared reference patch; this run is not recorded as an AI execution in the results:
python3 -m reproof repair artifacts/qa-bundle \
--patch-file scripts/sample-fix.json \
--output artifacts/offline-repair
With a signed-in Claude CLI, send the sample source and synthetic QA recording to generate a real patch:
python3 -m reproof repair artifacts/qa-bundle \
--agent claude --output artifacts/claude-repair
The pipeline confirms the same defect across 3 original runs, then applies the
patch in a separate work directory. A run is verified only when the installed
repaired APK's hash matches the build result, the protected JUnit regression
tests actually pass, and the candidate passes 3/3 runs. Mixed results are
inconclusive; successful runs are never cherry-picked.
Repair scope is limited to the integer increment expression in
CounterLogic.kt. If the agent touches other code, instrumentation, fixtures,
tests, or build settings, the build is blocked beforehand. Repair of declared
product files in general projects uses the separate
G9 path.
To run the offline example in one shot: bash scripts/demo.sh artifacts/my-demo. The output directory must be a new path; existing evidence
is never overwritten.
Reading results
Check report.html and job.json/result.json in the output directory.
Every repair attempt leaves a source copy, edits.json, patch.diff, a build
receipt, and repeated-run evidence. The original sample source is preserved as
the baseline.
Exit codes: 0 for successful reproduce/repair, 2 for unmet conditions or
environment blocks, 130 for CLI cancel.
Tests
python3 -m unittest discover -s tests -t . -v
# Regression tests on the fixed build
cd android
./gradlew --offline :sample:testFixedDebugUnitTest :sample:assembleFixedDebug :driver:assembleDebug
The original's :sample:testBuggyDebugUnitTest intentionally fails because of
the defect. In a copy where the repair loop patched CounterLogic, the same test
must pass. NO-SOURCE or skipped tests never count as verification success.
For extra scripts that check input, scroll, navigation, back, and
sensitive-input refusal on a physical device, see scripts/device_smoke.py --help.
Layout and scope
android/sdk: opt-in recording SDK — ordered JSONL, session freeze, and incomplete-record markersandroid/sample: buggy and fixed builds plus protected regression testsandroid/driver: observe, tap, input, scroll, and back through UiAutomation resource IDsreproof: bundle validation, compilation, ADB execution, repeat verdicts, repair orchestrator, HTML reportsios: Swift Recorder, UIKit sample, protected XCUITest and logic regressionreproof/ios_*: Simulator builds, v2 bundles, batch runs, Swift repair verificationschemas: public formats for recording and verdict datatests: failure, tampering, repeat-result, and patch-boundary checks
See the iOS support design, execution contracts and current limits, implementation and verification status, and the full plan. Per-app observation, fixtures, and independent verification have been exercised on a physical iPhone against a real product app; two-Mac acceptance and the isolated VM build lane are still open. Check support scope separately for the sample path and the shared QA path.
Live console
python3 -m reproof live-serve --demo starts the local console. Real iOS
Simulator connection, record/replay, and Python-script export are covered in
the live runbook; results are in
Live QA.
To inspect the work queue and recording library without a device, run
python3 -m reproof live-serve --demo --demo-count 2. See the
operations, CLI, and agent tools documentation.
For continuous touch, streaming, and worker execution on real Android — and for physical-iPhone signing, install, live record/replay results, and their limits — see the device live documentation.
The central coordinator for multiple users and registered hosts is a separate mode from the local console. Roles, project permissions, one-time host enrollment, TLS/CSRF, configuration files that contain no secrets, and the admin CLI are covered in the shared coordinator operations documentation.
Metadata
Release files for reproof-qa 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| reproof_qa-0.1.0.tar.gz | 1.7 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| reproof_qa-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.0 MB
Release files / reproof_qa-0.1.0.tar.gz
| Download URL | reproof_qa-0.1.0.tar.gz |
|---|---|
| Size | 1.7 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
68643715233d57cc0c6aa8c4d21923111e98578b2367b1ed3682358d507c1ea9
|
|
BLAKE2b-256 checksum How to use checksums |
0cdd76de6956659b99f4aa9aeb80a11e9d56dab74eb98fdd136d62d9cdbcca7c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency logRelease files / reproof_qa-0.1.0-py3-none-any.whl
| Download URL | reproof_qa-0.1.0-py3-none-any.whl |
|---|---|
| Size | 1.3 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
73a20369a4ce1ed1208e0208cc9696c99b9ffce47be0e1364e9e7711c40947f8
|
|
BLAKE2b-256 checksum How to use checksums |
bc2fe09b937b4148f018a6a181782e12ae47fef5f7e2949826cc6301e857e94e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.
Transparency log