Skip to main content

Robot Data Audit (RDA)

PyPI Python License: MIT Downloads Downloads/month Tests

audit: lerobot/pusht audit: AgiBotWorld2026 RL

Independent quality assessment for robot data. Runs locally — your data never leaves your machine.

RDA is a diagnostic tool. It does not guarantee training success-rate improvements.

RDA audits robot manipulation datasets (LeRobot format) and reports a three-tier verdict per episode — PASS / REVIEW / EXCLUDE — together with measured diagnostics. Use it as an independent check before you accept a vendor dataset, train a policy, or publish a benchmark.

Current release: v0.9.7pip install robot-data-audit.

⭐ Support RDA

If RDA helps you ship better robot data, star the repo — it takes one click and is the single best way to help us grow.

Star RDA on GitHub

Every star tells us we're on the right track. Bug reports, feature requests and real-world feedback are even better — see the feedback form or run rda feedback to submit one.

The four-layer audit

RDA runs every episode through four sequential layers. The key design rule: only hard integrity checks can flip an episode to EXCLUDE; diagnostic measurements never do. This separates "this data is broken" from "this data looks unusual", so observational signals never masquerade as fatal defects.

Layer Role # Metrics Can set verdict?
L1 — Integrity Gate Deterministic hard checks (missing / NaN / limit / video-stream) 9 ✅ PASS → REVIEW / EXCLUDE
L2 — Trajectory Diagnostics Observational motion & video anomalies 8 ❌ findings only
L3 — Dataset Profile Training-data efficiency & coverage 4 ❌ findings only
L4 — Dataset Summary Dataset-level P10/P50/P90 aggregation 📊 report only

An episode is EXCLUDE only if a critical L1 check fails (e.g. missing data, NaN, timestamp errors). Some L1 checks like video_freeze may also flag REVIEW when the issue is advisory rather than fatal. L2/L3 surface observational measurements and findings for the human reviewer. RDA measures and presents — the accept/reject decision stays with you.

Install

pip install robot-data-audit

Optional dependency tiers:

Tier Extra Unlocks
Core (default) parquet audits: integrity + temporal/motion + dataset-utility metrics
Visual pip install robot-data-audit[video] or pip install av the video visual metrics (freeze / timestamp-alignment / stream-span/offset/drift / quality)
Lerobot [lerobot] .parquet dataset loading via the lerobot package
UI [ui] the web dashboard
Everything [all] all of the above

Visual metrics without PyAV are reported as "not audited", never as "pass": the JSON report carries a top-level skipped_by_missing_dep field and the CLI prints a warning listing the skipped checks. The web dashboard shows the same guarantee — a dep-missing dataset renders a "not audited ≠ pass" banner.

Quick start

# 1. Audit a dataset — 21 metrics across four layers, three-tier verdicts
#    Default: Fast Audit (all metrics except visual_quality)
rda audit /path/to/lerobot/dataset

# 2. Include visual quality analysis
rda audit /path/to/lerobot/dataset --video-quality

# 3. Recommendations calibrated to your model type
rda recommend /path/to/dataset --policy temporal   # or frame-wise

# 4. Optional web dashboard
rda ui

rda audit is fully offline and emits a structured JSON report: per-episode verdicts, every metric's measurement/findings, plus a dataset-level acceptance_summary (P10/P50/P90 baselines, runtime outliers, and the not-checked inventory). rda recommend computes all metrics locally and sends only aggregated statistics (<1 KB) to the rules API — cached for offline reuse, and RDA_API_URL can point to your own server for private deployments.

Execution tiers

Visual quality analysis (visual_quality) involves frame decoding and is significantly slower than other metrics. RDA v0.9.4 introduces tiered execution so you can control which metrics run:

Mode Flag What runs
Fast Audit (default) All 21 metrics except visual_quality
Video Quality --video-quality All 21 metrics, including visual_quality
No Video --no-video All metrics except video-related (9 video metrics skipped)
Video Only --video-only Only the 9 video-related metrics
Full Audit --full All 21 metrics, including visual_quality

Flags are mutually exclusive. The JSON report (schema v1.2) includes an execution_tier field and a video_quality block indicating whether visual quality was executed and why. The text report header shows the active tier and whether visual quality was skipped.

Programmatic use

import numpy as np
from rda.io.schema import EpisodeData
from rda.audit.episode_audit import EpisodeAuditor

episode = EpisodeData(
    episode_index=0,
    num_frames=n_frames,
    timestamps=np.array(timestamps),          # seconds
    observation={"state": state_array},      # shape (T, DoF)
    action={"joint_pos": action_array},      # shape (T, DoF)
    meta={"fps": 10, "source": "my/dataset"},
)
result = EpisodeAuditor().audit(episode)
print(result.verdict)                        # PASS / REVIEW / EXCLUDE
for name, metric in result.metrics.items():
    if metric.has_finding:
        print(name, metric.measurement)

The 21 metrics

L1 — Integrity Gate (hard checks, can EXCLUDE) missing_dropout · invalid_values (NaN/Inf) · schema_consistency · temporal_validity · joint_limit (three-level PASS/REVIEW/EXCLUDE with configurable approach_threshold / consecutive_frames / jump_multiplier) · video_frame_integrity · video_freeze · video_timestamp_alignment · video_stream_presence

L2 — Trajectory Diagnostics (observational, never EXCLUDE) sensor_sync · sampling_jitter · velocity_acceleration · action_discontinuity (MAD-based spike detection) · visual_quality ⚡ · video_stream_span_consistency · video_stream_temporal_offset · video_stream_temporal_drift

visual_quality requires frame decoding and is skipped by default (Fast Audit). Enable with --video-quality or --full.

L3 — Dataset Profile (efficiency & coverage) idle_ratio (three-tier fallback: 30-bin valley → 3×MAD → 1e-6 floor) · distribution · coverage · temporal_structure

Metrics are also classified by cross-platform portability (MVP Spec v0.2.0 §1.5): Tier-1 universal (duration_sec, spike_count, effective_motion_ratio — comparable across any robot), Tier-2 normalizable (velocity / acceleration / jerk / path-length — need platform scaling), Tier-3 platform-specific (joint limits, workspace, torque/force/tactile).

Validated on real datasets

13 local datasets, 4,940 episodes, one set of default thresholds, zero per-dataset tuning — full table in docs/benchmark.md.

Four popular LeRobot datasets audited (v0.9.7) — we ran RDA against lerobot/pusht, aloha_sim_transfer_cube_human, xarm_lift_medium and droid_100 (1,156 episodes across a 2-DOF sim, a 14-DOF bimanual sim, a 4-DOF arm and a 7-DOF Franka). All episodes pass L1 integrity under v0.9.7; the dataset profiles differ sharply — median idle frames range from 20.8% (xArm) to 81.7% (PushT), and state space occupancy ranges from 2.4% to 39%. Under the new verdict pipeline, all four datasets now achieve 100% PASS rate. Every figure is reproducible straight from the PyPI package (pip install robot-data-audit==0.9.7).

Full audit of lerobot/libero_10 (v3.0) — 379 episodes, 101,469 frames: all applicable integrity checks clean, 373 PASS / 6 REVIEW / 0 EXCLUDE (6 REVIEW from video_freeze detection only). Read the report →

Blind test (v0.9.7) — we injected 50 defective episodes (5 defect classes, seed=42) into lerobot/pusht and kept 156 as controls. Precision 1.000 (zero false alarms on controls), recall 0.800 strict and broad (40/50 caught; frozen episodes are a known regression — idle_ratio findings no longer auto-escalate to REVIEW verdicts). Read the blind-test report →

Validated on AgiBotWorld2026 — third-party audit of AgiBot's Phase 3 dataset: all 5 simulation tasks + a real-robot RL package, 1,112 episodes, 4,448 integrity checks with 0 failures, and a 3.1× enrichment of RDA's discontinuity spikes at official human-takeover boundaries. Zero adaptation needed. Read the case study →

Every audit can also render into a shareable single-file HTML report and a README badge:

python tools/rda_render.py rda_report.json --html report.html --badge badge.svg

More: CLI reference & metrics table · experiments · real-world feedback form

Governance

RDA's metric I/O is pinned by a core-parameter spec, and every change is guarded by semantic-invariant tests run in CI. Seven invariant guard tests (INV-003 … INV-009) protect the core architecture — e.g. "L2 diagnostics never set an EXCLUDE verdict", "no inference is reported as a measurement", "the report always carries a tool version". The full suite runs on every push across Python 3.10–3.12.

Metric provenance

Every metric ships with a four-file provenance record (docs/provenance/<metric>/): algorithm.md (how it works), source.md (public precedents consulted — ideas only), implementation_origin.md (original implementation, zero third-party code), license.md (compliance notes). Index: docs/provenance/.

Citation

@software{robot_data_audit,
  title = {Robot Data Audit: Quality Auditing for Robot Manipulation Datasets},
  author = {Niu Su Tech},
  year = {2026},
  url = {https://github.com/liesliy/rda}
}

MIT License.

Release files for robot-data-audit 0.9.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for robot-data-audit 0.9.9
File Size Uploaded
robot_data_audit-0.9.9.tar.gz 254.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for robot-data-audit 0.9.9
File Interpreter ABI Platform
robot_data_audit-0.9.9-py3-none-any.whl Python 3 none any Details

Total release size: 502.2 kB

Release files / robot_data_audit-0.9.9.tar.gz

Download URL robot_data_audit-0.9.9.tar.gz
Size 254.2 kB
Tags Source
SHA-256 checksum
How to use checksums
2c21f3acca91e4a06d7865744c1e9760d6a7b2f3470c81ab7d574b61a902f30c
BLAKE2b-256 checksum
How to use checksums
5c0cd1388b33fb41b06a82cd057cbd7315b106b79978f8700850b8dea9af3bab
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.12

Release files / robot_data_audit-0.9.9-py3-none-any.whl

Download URL robot_data_audit-0.9.9-py3-none-any.whl
Size 248.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
050e14c6205dbf265469482399bb2cbaa64ae1020860f99af52c48d061b48082
BLAKE2b-256 checksum
How to use checksums
79b7173240b6b520248f361b25f5719ea22e721bd95ff4b88fb8e537728d012a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.12

Release history Release notifications | RSS feed

0.9.14

2 release files

0.9.13

2 release files

0.9.12

2 release files

0.9.11

2 release files

This release

0.9.9 This release

2 release files

0.9.8

1 release file

0.9.7

2 release files

0.9.6

2 release files

0.9.5

2 release files

0.9.4

2 release files

0.9.3

1 release file

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.0

1 release file

0.5.9

1 release file

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.15

2 release files

0.4.14

2 release files

0.4.13

1 release file

0.4.12

1 release file

0.4.11

1 release file

0.4.10

1 release file

0.4.9

1 release file

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.2

2 release files

0.2.1

1 release file

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page