Physical-AI data processing on Daft, starting with hand tracking.
Project description
daft-physical-ai
Physical-AI data processing on Daft: hand tracking and reward scoring. The methods run as Daft UDFs, so they slot into any Daft pipeline and execute lazily, batched, and distributed.
Available on PyPI:
pip install "daft-physical-ai[mediapipe]"
API
The package operates on a Daft image column and returns a hand-pose column. A
LeRobot dataset is a natural source: Daft's native reader
daft.datasets.lerobot (added in Daft #7090)
decodes each camera into an image column with load_video_frames.
import daft
from daft.datasets import lerobot
from daft_physical_ai.hands import track_hands
# one row per frame; the camera key is decoded into an image column.
# egodex-test is a tiny EgoDex sample (3 episodes / 632 frames) in LeRobot v3 format.
df = lerobot.read("pepijn223/egodex-test", load_video_frames="observation.image")
# pick a method (each returns the same schema):
# mediapipe -> CPU, 2D only, permissive license, no weights to supply
# wilor -> GPU, 3D MANO keypoints (MANO weights user-supplied)
df = df.with_column("hands", track_hands(df["observation.image"], method="mediapipe"))
df.write_parquet("annotated/")
Raw EgoDex releases
For Apple's original EgoDex HDF5+MP4 release, use the extension's lazy reader:
from daft_physical_ai.datasets import egodex
episodes = egodex.raw("/data/egodex", tasks="fold_towel").limit(2)
poses = egodex.trajectory(episodes, fields=["transforms/leftHand", "transforms/rightHand"])
frames = egodex.camera_frames(poses, width=224, height=224, sample_interval_seconds=1.0)
EgoDex is CC-BY-NC-ND, so the package does not download, extract, or redistribute it. Download and extract the archives from the official EgoDex repository, then point raw() at your copy. See the runnable example.
Install the method you need as an extra: pip install "daft-physical-ai[mediapipe]"
(CPU, 2D), pip install "daft-physical-ai[wilor]" (GPU, 3D), or
pip install "daft-physical-ai[all]" for both. WiLoR additionally needs a CUDA
torch build and chumpy from git
(pip install 'chumpy @ git+https://github.com/mattloper/chumpy', omitted from the
extra because PyPI metadata can't carry direct references), plus a user-supplied
MANO_RIGHT.pkl (research-gated).
Output schema
One unified output schema regardless of method: each frame yields a list of 0-2 detected hands. A single hand value (MediaPipe):
{
"handedness": "right", # "left", "right", or "unknown"
"confidence": 0.979,
"kp2d": [[1412.1, 1111.1], # 21 image-space [x, y] keypoints
[1357.9, 1075.9],
...],
"kp3d": None, # 21 [x, y, z] keypoints, or null for 2D-only methods
}
The Daft type is list[struct{ handedness: string, confidence: float32, kp2d: list[list[float32]], kp3d: list[list[float32]] }], defined as HANDS_DTYPE in
daft_physical_ai/hands/schema.py.
Reward scoring
Score episodes with a reward model (Robometer-4B) - per-frame task progress (0-1) plus success probability, written back as a dataset column. Use it to filter failed or stalled episodes before BC training, as dense reward for RL post-training, or to catch mislabeled tasks.
from daft_physical_ai.rewards import score_rewards
# one row per episode: task text, length, and where its frames live in the video
# (e.g. from daft.datasets.lerobot.read_episodes - the video column can be a
# Daft file handle or a local path string)
df = df.with_column(
"rewards",
score_rewards(
df["task"], df["length"], df["from_ts"], df["to_ts"], df["video"],
url="http://localhost:8001", # any running Robometer eval server
max_frames=8, # frames sampled per episode
),
)
Scoring is a pure HTTP call: the package never imports the model - you bring a
running Robometer eval server and
pass its URL. daft-physical-ai rewards scaffolds a complete demo plus the two
server scripts to run one yourself (run_robometer_server.py for any NVIDIA
GPU, modal_eval_server.py for Modal). The output type:
struct {
reward_score: list[float64] # per-frame task progress, 0-1
robometer_success: list[float64] # per-frame success probability
reward_frames: list[struct{index, timestamp_s}] # which frames were scored
}
Example
A complete walkthrough - read a dataset, run track_hands (MediaPipe), draw the
keypoints, and score against EgoDex ground truth:
Available in three equivalent forms:
- examples/hands/demo.md - read it start to finish; code and outputs inline.
- examples/hands/demo.ipynb - runnable notebook (outputs included).
- examples/hands/demo.py - plain script.
Generate your own (other methods, a Modal GPU runtime, with/without eval) with the
daft-physical-ai hands command - run it with no flags for an interactive
walkthrough, or pass flags:
# No flags - interactive walkthrough that asks a few questions
uvx daft-physical-ai hands
# --no-input skips all prompts; flags supply the answers, the rest use defaults
uvx daft-physical-ai hands --method mediapipe --output-dir my-demo --no-input
uvx daft-physical-ai hands --method wilor --runtime modal --mano-path ./MANO_RIGHT.pkl --no-input
uvx runs the CLI without installing anything (scaffolding needs no inference
deps). If the PyPI package is
already installed (pip install daft-physical-ai), plain daft-physical-ai hands
works too; from a clone of this repo, uv sync installs it (uv run daft-physical-ai).
Each capability is its own subcommand - daft-physical-ai hands and
daft-physical-ai rewards so far (daft-physical-ai with no arguments lists
what's available). The rewards scaffold also writes the Robometer server
scripts next to the demo, so one directory holds everything: score the
episodes, and serve the model locally or on Modal.
To run a generated demo you also need its inference stack. uvx covers that
too - one line, nothing installed:
uvx --from jupyterlab --with "daft-physical-ai[mediapipe]" --with matplotlib --with scipy \
jupyter-lab hand-tracking-demo/demo.ipynb
(scipy is only needed if the demo includes the ground-truth eval.)
In a clone, uv sync already brings a Daft with the LeRobot reader; install the
extras into the venv, then run from the activated venv - not uv run, which
re-syncs the env and would drop them:
source .venv/bin/activate
uv pip install -U av mediapipe scipy opencv-python matplotlib jupyterlab
jupyter lab hand-tracking-demo/demo.ipynb
Development
uv sync # set up env + install deps
uv run pre-commit install # install lint/format hooks
uv run pytest tests/ -v # run the test suite
Versioning
Versions are derived from git tags via hatch-vcs. Tag releases as v0.1.0,
v0.2.0, etc.
Publishing
Publishing a GitHub release triggers .github/workflows/publish-package.yml,
which builds a wheel and sdist with uv build and uploads both to PyPI via
trusted publishing. Configure the
trusted publisher on PyPI for this repository before the first release.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file daft_physical_ai-0.2.0.tar.gz.
File metadata
- Download URL: daft_physical_ai-0.2.0.tar.gz
- Upload date:
- Size: 1.1 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
27e52a8c4f85086fc5ecdac612c198fc2eaacfd980bee711bb5bc866cb1ba03e
|
|
| MD5 |
31e822768ed7fab5b10074e441449dfb
|
|
| BLAKE2b-256 |
9f71d0933dc1837ef0e6faa62d946606df1bc476459515167ec7a998d8e3ab9e
|
Provenance
The following attestation bundles were made for daft_physical_ai-0.2.0.tar.gz:
Publisher:
publish-package.yml on Eventual-Inc/daft-physical-ai
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
daft_physical_ai-0.2.0.tar.gz -
Subject digest:
27e52a8c4f85086fc5ecdac612c198fc2eaacfd980bee711bb5bc866cb1ba03e - Sigstore transparency entry: 2192170731
- Sigstore integration time:
-
Permalink:
Eventual-Inc/daft-physical-ai@1409829627b2fc59be950a6d7b26a80b6386bf0d -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/Eventual-Inc
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-package.yml@1409829627b2fc59be950a6d7b26a80b6386bf0d -
Trigger Event:
push
-
Statement type:
File details
Details for the file daft_physical_ai-0.2.0-py3-none-any.whl.
File metadata
- Download URL: daft_physical_ai-0.2.0-py3-none-any.whl
- Upload date:
- Size: 50.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
27dd92ac7a2ee9e02b4a9f10a2a33eb9e69b6f30d256ab8f57a79b56259e6f21
|
|
| MD5 |
8e27a25e55725123787191f70a2fe5de
|
|
| BLAKE2b-256 |
082f284a9639d3b2fd2ff36a305d378f9e6d4e7fbabac565f0125e54360aa590
|
Provenance
The following attestation bundles were made for daft_physical_ai-0.2.0-py3-none-any.whl:
Publisher:
publish-package.yml on Eventual-Inc/daft-physical-ai
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
daft_physical_ai-0.2.0-py3-none-any.whl -
Subject digest:
27dd92ac7a2ee9e02b4a9f10a2a33eb9e69b6f30d256ab8f57a79b56259e6f21 - Sigstore transparency entry: 2192170743
- Sigstore integration time:
-
Permalink:
Eventual-Inc/daft-physical-ai@1409829627b2fc59be950a6d7b26a80b6386bf0d -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/Eventual-Inc
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-package.yml@1409829627b2fc59be950a6d7b26a80b6386bf0d -
Trigger Event:
push
-
Statement type: