Streaming dataloader for robotics trajectory datasets
Project description
Traceplane
Python SDK for the Traceplane trajectory data platform.
Installation
pip install traceplane
With framework extras:
pip install traceplane[torch] # PyTorch DataLoader
pip install traceplane[jax] # JAX support
pip install traceplane[training] # Diffusion policy training
pip install traceplane[sim] # Isaac Sim backend (NVIDIA GPU required)
pip install traceplane[mujoco] # MuJoCo backend (CPU-friendly, works on Blackwell)
pip install traceplane[all] # Everything
Picking a sim backend
| Backend | Install | Use when |
|---|---|---|
traceplane[sim] (Isaac Sim) |
pip install isaacsim --extra-index-url https://pypi.nvidia.com then pip install traceplane[sim] |
Photorealistic rendering, USD assets, NVIDIA GPU available, host is pre-Blackwell |
traceplane[mujoco] (MuJoCo) |
pip install traceplane[mujoco] |
CPU eval, CI runs, Blackwell hosts (RTX 5080+) where Isaac Sim 5.x currently segfaults, no NVIDIA dependency |
Both backends share the same config + metrics layer; pick at runtime:
traceplane-sim-eval --sim-backend=mujoco --task=pick_place_cube --num-episodes=10
Quick Start
from traceplane import TraceplaneClient
client = TraceplaneClient("https://api.traceplane.ai", api_key="tp_live_...")
# Register a dataset
client.register("my_data", "/path/to/dataset", include_data=True)
# Query with SQL
rows = client.sql_rows("SELECT * FROM my_data WHERE frame_count > 100")
# Upload data
client.upload_dataset("my_data", "/path/to/parquet/files/")
# Vector search
results = client.search_similar("my_data", episode_index=0, k=5)
Features
- SQL query engine -- register datasets and query with full SQL, including vector UDFs (
vec_mean,vec_norm,vec_cosine_sim, etc.) - Streaming dataloaders -- PyTorch, JAX, and TensorFlow adapters with windowed sampling
- LeRobot format -- native reader and writer for LeRobot v2/v3 datasets (Parquet + MP4)
- Embodiment registry -- 8 built-in robot profiles (Franka, ALOHA, Unitree G1/H1, UR5e, KUKA iiwa, xArm 7, Stretch 3) with joint-limit + action-space validation, shared verbatim with the Rust backend
- Similarity search -- find related episodes via embedding-based vector search; DINOv2 keyframe embeddings for visual similarity (
traceplane[vision]) - Dataset upload -- push local Parquet files to the platform
- Retargeting -- XR hand poses to robot action space via calibration bridge
- Training -- built-in diffusion policy training with
traceplane-trainCLI - Isaac Sim round-trip -- capture headless rollouts as canonical LeRobot v2 episodes via
SimEpisodeWriter
Training Integration
from traceplane import LeRobotReader
from traceplane.torch import TorchEpisodeLoader
reader = LeRobotReader("/path/to/lerobot/dataset")
loader = TorchEpisodeLoader(reader, batch_size=32, window_size=16)
for batch in loader:
observations = batch["observation"]
actions = batch["action"]
# ... your training loop
Embodiment Profiles
Built-in kinematic profiles for 8 robots, each with joint-limit + action-space validation. The same YAML files drive both this SDK and the Rust backend, so cross-embodiment queries and retargeting validation stay consistent end-to-end.
from traceplane.embodiments import get_profile, list_profiles
print(list_profiles())
# ['franka_panda', 'aloha_v2', 'unitree_g1', 'unitree_h1',
# 'ur5e', 'kuka_iiwa', 'xarm7', 'stretch3']
g1 = get_profile("unitree_g1")
print(g1.total_dof()) # 23
print(g1.group_names()) # ['left_arm', 'right_arm', 'waist', ...]
arm_action = g1.project_action(full_action_23d, "left_arm") # 7-dim subvector
Sim Rollout Round-trip
SimEpisodeWriter turns Isaac Sim rollouts (or any simulator's per-step
actions/states) into canonical LeRobot v2 datasets, ready to be re-ingested
by Traceplane or consumed by any LeRobot-compatible tool.
from traceplane.embodiments import get_profile
from traceplane.sim.writer import SimEpisodeWriter
profile = get_profile("franka_panda")
with SimEpisodeWriter("./rollouts", profile, fps=30.0, task="pick-cube") as w:
for ep in rollouts:
w.add_episode(
actions=ep.actions, # (T, profile.action_dim)
states=ep.joint_positions, # (T, profile.state_dim)
timestamps=ep.timestamps, # optional; defaults to 1/fps spacing
metadata={"success": ep.success, "seed": ep.seed},
)
Vision Frame Embeddings
Encode keyframes from a dataset's videos with a pretrained DINOv2 model for
visual similarity search and near-duplicate detection. This is the vision
counterpart to traceplane.embeddings (trajectory statistics + text labels).
Requires the optional vision extra:
pip install traceplane[vision]
# Embed every 5th keyframe of each episode's video to a Parquet shard
python -m traceplane.frame_embeddings embed ./my-dataset --stride 5
# Find the 10 nearest frames to a query frame
python -m traceplane.frame_embeddings search \
./my-dataset/embeddings/frame_embeddings.parquet \
--query-episode 0 --query-frame 50 --k 10
from traceplane.frame_embeddings import compute_frame_embeddings, search_frames
path = compute_frame_embeddings("./my-dataset", stride=5) # -> Parquet
hits = search_frames(path, query_episode=0, query_frame=50, k=10)
Embeddings are the L2-normalised DINOv2 CLS token (768-dim), stored as Parquet
and searched with brute-force cosine — consistent with the existing
vec_cosine_sim path. (ANN/Lance is deferred until frame counts require it.)
Once embeddings exist, traceplane check --ml adds ML-powered QA to the
standard dataset CI report: outlier flagging (k-NN cosine distance, z-scored)
and near-duplicate clustering, layered on top of the existing 50+ rule-based
checks.
# After running `python -m traceplane.frame_embeddings embed ./my-dataset`:
traceplane check ./my-dataset --ml
# Optional knobs:
traceplane check ./my-dataset --ml --outlier-z 3.0 --dedupe-threshold 0.97
traceplane check ./my-dataset --ml --embeddings /elsewhere/fe.parquet
API Reference
Full documentation: docs.traceplane.ai
License
Apache-2.0
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file traceplane-0.4.0.tar.gz.
File metadata
- Download URL: traceplane-0.4.0.tar.gz
- Upload date:
- Size: 108.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0906b26b1a5b0b1885e2586ba01d46e641662817dcda8ea65514323b97b7e3a6
|
|
| MD5 |
d55dc7be8e5ad3601c90b66de1dde8f3
|
|
| BLAKE2b-256 |
254a89bfcec74fa80f86ba7f93622470ead81ee82b6f36ebcd55b749421fb7dc
|
File details
Details for the file traceplane-0.4.0-py3-none-any.whl.
File metadata
- Download URL: traceplane-0.4.0-py3-none-any.whl
- Upload date:
- Size: 136.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1e962bce009c94db81308731421a41d76f4b75bb5e0e19ed9129e53b626b2354
|
|
| MD5 |
2be60476b51fa20108672b149aa6aa89
|
|
| BLAKE2b-256 |
f39203336f7dffb5d049b3cbcbf6c337bdff3a652bf74f44a4258328d6e1d045
|