Robot Data Audit (RDA)
Independent quality assessment for robot data. Runs locally — your data never leaves your machine.
RDA is a diagnostic tool. It does not guarantee training success-rate improvements.
RDA audits robot manipulation datasets (LeRobot format) for integrity, temporal consistency, motion quality and distribution coverage, then generates optimization recommendations calibrated to your target model architecture. Use it as an independent check before you accept a vendor dataset, train a policy, or publish a benchmark.
Install
pip install robot-data-audit
Optional dependency tiers:
| Tier | Extra | Unlocks |
|---|---|---|
| Core (default) | — | parquet audits: integrity + temporal/motion + dataset-utility metrics |
| Visual | pip install robot-data-audit[video] or pip install av |
the 4 visual metrics (freeze / timestamp-alignment / stream-sync / quality) |
| Lerobot | [lerobot] |
.parquet dataset loading via the lerobot package |
| UI | [ui] |
the web dashboard |
Visual metrics without PyAV are reported as "not audited", never as
"pass": the JSON report carries a top-level skipped_by_missing_dep
field and the CLI prints a warning listing the skipped checks. The web
dashboard shows the same guarantee — a dep-missing dataset renders a
"not audited ≠ pass" banner, and the Health Overview adds a Visual
Integrity section (checked / flagged / not-checked episode counts) while
Episode Explorer gains a per-episode Visual Audit panel (freeze-region
table, per-camera quality penalties, per-sample quality curves).
Acceptance summary (v0.8.0): rda audit --format json emits an
acceptance_summary block — dataset-level P10/P50/P90 baselines,
Tukey-IQR runtime outliers, the Tier-1 calibration layer and the
not-checked inventory — and the dashboard adds an Acceptance page that
presents it as one deliverable for the data-acceptance party. RDA measures
and presents; the accept/reject decision stays with the human reviewer.
Quick start
# 1. Audit a dataset — 18 metrics incl. visual-stream integrity, three-tier verdicts (PASS / REVIEW / EXCLUDE)
rda audit /path/to/lerobot/dataset
# 2. Recommendations calibrated to your model type
rda recommend /path/to/dataset --policy temporal # or frame-wise
# 3. Optional web dashboard
rda ui
rda audit is fully offline. rda recommend computes all metrics locally and sends only aggregated statistics (<1KB) to the rules API — cached for offline reuse, and RDA_API_URL can point to your own server for private deployments. The usable_retention statistic is fps-aware (v0.7.3): a "usable" run must span ≥16 frames and ≥1 second (DROID's min_non_idle_len=16 encodes 1 s at 15-30 Hz; on 50 Hz data the floor rises to 50 frames so sub-second twitches no longer count).
Why RDA
12 local datasets, 4,959 episodes, one set of default thresholds, zero per-dataset tuning — full table in docs/benchmark.md.
Full audit of lerobot/libero_10 (v3.0) — 379 episodes, 101,469 frames: all 12 applicable integrity checks clean, 0 hard defects; the one REVIEW signal (low-motion heuristic) is discussed honestly. Read the report →
Blind test — we injected 50 defective episodes (5 defect classes, seed=42) into lerobot/pusht and kept 156 as controls. RDA caught all 50 under the broad criterion, precision 1.000 (zero false alarms on controls) under the strict one. Read the blind-test report →
Validated on AgiBotWorld2026 — third-party audit of AgiBot's Phase 3 dataset: all 5 simulation tasks + a real-robot RL package, 1,112 episodes, 4,448 integrity checks with 0 failures, and a 3.1× enrichment of RDA's discontinuity spikes at official human-takeover boundaries. Zero adaptation needed. Read the case study →
Every audit can also render into a shareable single-file HTML report and a README badge:
python tools/rda_render.py rda_report.json --html report.html --badge badge.svg
More: CLI reference & metrics table · experiments · real-world feedback form
Metric provenance
Every metric ships with a four-file provenance record (docs/provenance/<metric>/): algorithm.md (how it works), source.md (public precedents consulted — ideas only), implementation_origin.md (original implementation, zero third-party code), license.md (compliance notes). Covered: all 18 metrics (missing_dropout, invalid_values, schema_consistency, timestamp_validity, joint_limit, video_frame_integrity, video_freeze, video_timestamp_alignment, video_stream_sync, visual_quality, sensor_synchronization, sampling_jitter, velocity_acceleration, action_discontinuity, temporal_sufficiency, idle_ratio, distribution, coverage). Index: docs/provenance/.
Citation
@software{robot_data_audit,
title = {Robot Data Audit: Quality Auditing for Robot Manipulation Datasets},
author = {Niu Su Tech},
year = {2026},
url = {https://github.com/liesliy/rda}
}
MIT License.
Release files for robot-data-audit 0.9.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| robot_data_audit-0.9.2.tar.gz | 235.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| robot_data_audit-0.9.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 464.3 kB
Release files / robot_data_audit-0.9.2.tar.gz
| Download URL | robot_data_audit-0.9.2.tar.gz |
|---|---|
| Size | 235.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
22841c074d37bf9c3497891554f0134f13ef892f4f326b9753523e1f39786674
|
|
BLAKE2b-256 checksum How to use checksums |
981f9eaf7a7472dce51d6698c3a770c233bebb955c357dd2896b5c6a88b3e65d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.12
|
Release files / robot_data_audit-0.9.2-py3-none-any.whl
| Download URL | robot_data_audit-0.9.2-py3-none-any.whl |
|---|---|
| Size | 228.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
44fd6cc50cb490eaab1e18b4ca54d05145e7523dc1bf87305ef8e79f770c4910
|
|
BLAKE2b-256 checksum How to use checksums |
11f86affa7ebed0445333e81875ee4464ac6119af768c5db146ac7f82f7fd835
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.12
|