Robot Data Audit (RDA)
Independent quality assessment for robot data. Runs locally — your data never leaves your machine.
RDA is a diagnostic tool. It does not guarantee training success-rate improvements.
RDA audits robot manipulation datasets (LeRobot format) and reports a three-tier verdict per episode — PASS / REVIEW / EXCLUDE — together with measured diagnostics. Use it as an independent check before you accept a vendor dataset, train a policy, or publish a benchmark.
Current release: v0.9.7 — pip install robot-data-audit.
⭐ Support RDA
If RDA helps you ship better robot data, star the repo — it takes one click and is the single best way to help us grow.
Every star tells us we're on the right track. Bug reports, feature requests and
real-world feedback are even better — see the
feedback form
or run rda feedback to submit one.
The four-layer audit
RDA runs every episode through four sequential layers. The key design rule: only hard integrity checks can flip an episode to EXCLUDE; diagnostic measurements never do. This separates "this data is broken" from "this data looks unusual", so observational signals never masquerade as fatal defects.
| Layer | Role | # Metrics | Can set verdict? |
|---|---|---|---|
| L1 — Integrity Gate | Deterministic hard checks (missing / NaN / limit / video-stream) | 9 | ✅ PASS → REVIEW / EXCLUDE |
| L2 — Trajectory Diagnostics | Observational motion & video anomalies | 8 | ❌ findings only |
| L3 — Dataset Profile | Training-data efficiency & coverage | 4 | ❌ findings only |
| L4 — Dataset Summary | Dataset-level P10/P50/P90 aggregation | — | 📊 report only |
An episode is EXCLUDE only if a critical L1 check fails (e.g. missing data, NaN, timestamp errors). Some L1 checks like video_freeze may also flag REVIEW when the issue is advisory rather than fatal. L2/L3 surface observational measurements and findings for the human reviewer. RDA measures and presents — the accept/reject decision stays with you.
Install
pip install robot-data-audit
Optional dependency tiers:
| Tier | Extra | Unlocks |
|---|---|---|
| Core (default) | — | parquet audits: integrity + temporal/motion + dataset-utility metrics |
| Visual | pip install robot-data-audit[video] or pip install av |
the video visual metrics (freeze / timestamp-alignment / stream-span/offset/drift / quality) |
| Lerobot | [lerobot] |
.parquet dataset loading via the lerobot package |
| UI | [ui] |
the web dashboard |
| Everything | [all] |
all of the above |
Visual metrics without PyAV are reported as "not audited", never as
"pass": the JSON report carries a top-level skipped_by_missing_dep field
and the CLI prints a warning listing the skipped checks. The web dashboard
shows the same guarantee — a dep-missing dataset renders a "not audited ≠
pass" banner.
Quick start
# 1. Audit a dataset — 21 metrics across four layers, three-tier verdicts
# Default: Fast Audit (all metrics except visual_quality)
rda audit /path/to/lerobot/dataset
# 2. Include visual quality analysis
rda audit /path/to/lerobot/dataset --video-quality
# 3. Recommendations calibrated to your model type
rda recommend /path/to/dataset --policy temporal # or frame-wise
# 4. Optional web dashboard
rda ui
rda audit is fully offline and emits a structured JSON report: per-episode
verdicts, every metric's measurement/findings, plus a dataset-level
acceptance_summary (P10/P50/P90 baselines, runtime outliers, and the
not-checked inventory). rda recommend computes all metrics locally and
sends only aggregated statistics (<1 KB) to the rules API — cached for
offline reuse, and RDA_API_URL can point to your own server for private
deployments.
Execution tiers
Visual quality analysis (visual_quality) involves frame decoding and is
significantly slower than other metrics. RDA v0.9.4 introduces tiered
execution so you can control which metrics run:
| Mode | Flag | What runs |
|---|---|---|
| Fast Audit | (default) | All 21 metrics except visual_quality |
| Video Quality | --video-quality |
All 21 metrics, including visual_quality |
| No Video | --no-video |
All metrics except video-related (9 video metrics skipped) |
| Video Only | --video-only |
Only the 9 video-related metrics |
| Full Audit | --full |
All 21 metrics, including visual_quality |
Flags are mutually exclusive. The JSON report (schema v1.2) includes an
execution_tier field and a video_quality block indicating whether
visual quality was executed and why. The text report header shows the
active tier and whether visual quality was skipped.
Programmatic use
import numpy as np
from rda.io.schema import EpisodeData
from rda.audit.episode_audit import EpisodeAuditor
episode = EpisodeData(
episode_index=0,
num_frames=n_frames,
timestamps=np.array(timestamps), # seconds
observation={"state": state_array}, # shape (T, DoF)
action={"joint_pos": action_array}, # shape (T, DoF)
meta={"fps": 10, "source": "my/dataset"},
)
result = EpisodeAuditor().audit(episode)
print(result.verdict) # PASS / REVIEW / EXCLUDE
for name, metric in result.metrics.items():
if metric.has_finding:
print(name, metric.measurement)
The 21 metrics
L1 — Integrity Gate (hard checks, can EXCLUDE)
missing_dropout · invalid_values (NaN/Inf) · schema_consistency ·
temporal_validity · joint_limit (three-level PASS/REVIEW/EXCLUDE with
configurable approach_threshold / consecutive_frames / jump_multiplier)
· video_frame_integrity · video_freeze · video_timestamp_alignment ·
video_stream_presence
L2 — Trajectory Diagnostics (observational, never EXCLUDE)
sensor_sync · sampling_jitter · velocity_acceleration ·
action_discontinuity (MAD-based spike detection) · visual_quality ⚡ ·
video_stream_span_consistency · video_stream_temporal_offset ·
video_stream_temporal_drift
⚡
visual_qualityrequires frame decoding and is skipped by default (Fast Audit). Enable with--video-qualityor--full.
L3 — Dataset Profile (efficiency & coverage)
idle_ratio (three-tier fallback: 30-bin valley → 3×MAD → 1e-6 floor) ·
distribution · coverage · temporal_structure
Metrics are also classified by cross-platform portability (MVP Spec
v0.2.0 §1.5): Tier-1 universal (duration_sec, spike_count,
effective_motion_ratio — comparable across any robot), Tier-2 normalizable
(velocity / acceleration / jerk / path-length — need platform scaling),
Tier-3 platform-specific (joint limits, workspace, torque/force/tactile).
Validated on real datasets
13 local datasets, 4,940 episodes, one set of default thresholds, zero per-dataset tuning — full table in docs/benchmark.md.
Four popular LeRobot datasets audited (v0.9.7) — we ran RDA against
lerobot/pusht, aloha_sim_transfer_cube_human, xarm_lift_medium and
droid_100 (1,156 episodes across a 2-DOF sim, a 14-DOF bimanual sim, a
4-DOF arm and a 7-DOF Franka). All episodes pass L1 integrity under v0.9.7;
the dataset profiles differ sharply — median idle frames range from 20.8%
(xArm) to 81.7% (PushT), and state space occupancy ranges from 2.4% to
39%. Under the new verdict pipeline, all four datasets now achieve 100% PASS
rate. Every figure is reproducible straight from the PyPI package
(pip install robot-data-audit==0.9.7).
Full audit of lerobot/libero_10 (v3.0) — 379 episodes, 101,469 frames: all applicable integrity checks clean, 373 PASS / 6 REVIEW / 0 EXCLUDE (6 REVIEW from video_freeze detection only). Read the report →
Blind test (v0.9.7) — we injected 50 defective episodes (5 defect classes,
seed=42) into lerobot/pusht and kept 156 as controls. Precision 1.000
(zero false alarms on controls), recall 0.800 strict and broad (40/50
caught; frozen episodes are a known regression — idle_ratio findings no
longer auto-escalate to REVIEW verdicts).
Read the blind-test report →
Validated on AgiBotWorld2026 — third-party audit of AgiBot's Phase 3 dataset: all 5 simulation tasks + a real-robot RL package, 1,112 episodes, 4,448 integrity checks with 0 failures, and a 3.1× enrichment of RDA's discontinuity spikes at official human-takeover boundaries. Zero adaptation needed. Read the case study →
Every audit can also render into a shareable single-file HTML report and a README badge:
python tools/rda_render.py rda_report.json --html report.html --badge badge.svg
More: CLI reference & metrics table · experiments · real-world feedback form
Governance
RDA's metric I/O is pinned by a core-parameter spec, and every change is guarded by semantic-invariant tests run in CI. Seven invariant guard tests (INV-003 … INV-009) protect the core architecture — e.g. "L2 diagnostics never set an EXCLUDE verdict", "no inference is reported as a measurement", "the report always carries a tool version". The full suite runs on every push across Python 3.10–3.12.
Metric provenance
Every metric ships with a four-file provenance record
(docs/provenance/<metric>/): algorithm.md (how it works),
source.md (public precedents consulted — ideas only),
implementation_origin.md (original implementation, zero third-party
code), license.md (compliance notes). Index:
docs/provenance/.
Citation
@software{robot_data_audit,
title = {Robot Data Audit: Quality Auditing for Robot Manipulation Datasets},
author = {Niu Su Tech},
year = {2026},
url = {https://github.com/liesliy/rda}
}
MIT License.
Release files for robot-data-audit 0.9.9
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| robot_data_audit-0.9.9.tar.gz | 254.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| robot_data_audit-0.9.9-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 502.2 kB
Release files / robot_data_audit-0.9.9.tar.gz
| Download URL | robot_data_audit-0.9.9.tar.gz |
|---|---|
| Size | 254.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2c21f3acca91e4a06d7865744c1e9760d6a7b2f3470c81ab7d574b61a902f30c
|
|
BLAKE2b-256 checksum How to use checksums |
5c0cd1388b33fb41b06a82cd057cbd7315b106b79978f8700850b8dea9af3bab
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.12
|
Release files / robot_data_audit-0.9.9-py3-none-any.whl
| Download URL | robot_data_audit-0.9.9-py3-none-any.whl |
|---|---|
| Size | 248.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
050e14c6205dbf265469482399bb2cbaa64ae1020860f99af52c48d061b48082
|
|
BLAKE2b-256 checksum How to use checksums |
79b7173240b6b520248f361b25f5719ea22e721bd95ff4b88fb8e537728d012a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.12
|