Skip to main content

Tiny Python SDK for the Drift dataset API

Project description

Drift

Drift is an open-source Python SDK for working with real-world driving data and training end-to-end driving models.

The goal is to make embodied driving research easier to start. Instead of spending days wiring up dataset parsers, route metadata, and segment downloads, you can install one package, load a dataset, and begin building a training pipeline from real driving data.

Today, Drift includes a complete real-world driving dataset with:

  • 15.08 hours of human driving
  • 933 synchronized driving segments
  • 68 routes
  • 74.9 GiB of synchronized camera, vehicle telemetry, and navigation data

Drift exposes datasets, routes, and segments as simple Python objects while hiding the storage details underneath. It is best understood as a clean access layer over driving data and artifacts, rather than a fully prepared ML dataset.

If you are new to the project, start with the included example notebook:

  • examples/drift_train_driving.ipynb

That notebook walks through the end-to-end flow: loading the dataset, inspecting routes and segments, downloading data, training a simple driving policy, and exporting a model.

Drift aims to do for end-to-end driving data what Hugging Face Datasets did for NLP datasets: provide a simple, reproducible interface that makes experimentation more accessible.

The intended SDK flow is:

  1. Load dataset metadata.
  2. List routes.
  3. Inspect route segment references.
  4. Resolve each segment.
  5. Download the specific files your pipeline needs.

Installation

pip install drift

API endpoint

The hosted API used by the examples below is:

https://drift-api-production-1d47.up.railway.app

You do not need to run the backend locally to try the examples.

Core objects

The SDK exposes three data objects plus one client:

  • drift.client.DriftClient
  • drift.dataset.Dataset
  • drift.route.Route
  • drift.segment.Segment

The top-level convenience helpers are:

  • drift.load(dataset_id, base_url=None) -> Dataset
  • drift.route(route_id, base_url=None) -> Route
  • drift.segment(segment_id, base_url=None) -> Segment

Client methods

DriftClient is the full method-level API:

from drift.client import DriftClient

client = DriftClient(base_url="https://drift-api-production-1d47.up.railway.app")

dataset = client.load_dataset("jordy-v0")
routes = client.list_routes("jordy-v0")
route = client.get_route(routes[0].route_id)
segment = client.get_segment(route.segments[0]["segment_id"])

client.close()

Available methods:

  • load_dataset(dataset_id) returns a Dataset
  • list_routes(dataset_id) returns list[Route]
  • get_route(route_id) returns a Route
  • get_segment(segment_id) returns a Segment
  • close() closes the underlying HTTP client

Dataset methods

Use drift.load(...) or DriftClient.load_dataset(...) when you need dataset summary metadata.

import drift

dataset = drift.load("jordy-v0")
print(dataset.id)
print(dataset.routes_count)
print(dataset.segments_count)

Dataset fields:

  • id
  • routes_count
  • segments_count

Route methods

Use drift.route(...) or DriftClient.get_route(...) when you need the segment references for a route.

import drift

client = drift.client.DriftClient()
routes = client.list_routes("jordy-v0")

route = client.get_route(routes[0].route_id)
print(route.route_id)
print(route.num_segments)
print(route.dataset)
print(route.segments[:3])

client.close()

Route fields:

  • route_id
  • num_segments
  • created_at
  • dataset
  • segments

Each route.segments item is currently a small reference dictionary such as:

{"segment_id": "...", "segment_index": 0}

Segment methods

Use drift.segment(...) or DriftClient.get_segment(...) when you need the downloadable files for one segment.

import drift

segment = drift.segment("<segment-id>")
print(segment.segment_id)
print(segment.route_id)
print(segment.segment_index)
print(segment.files)
print(segment.download_url)

Segment fields:

  • segment_id
  • route_id
  • segment_index
  • files
  • download_url

segment.files is a dictionary of alias to presigned URL. Current aliases may include:

  • qcamera
  • qlog
  • rlog
  • ecamera
  • fcamera
  • navigation

Downloading files

Segment.download(destination=None, chunk_size=8192) downloads the segment's primary file. Today that is download_url if present, otherwise files["fcamera"].

import os
import drift

os.makedirs("downloads", exist_ok=True)

segment = drift.segment("<segment-id>")
path = segment.download("downloads/segment.bin")
print(path)

Important limitation:

  • Segment.download(...) downloads one file per call.
  • It does not download every file listed in segment.files.
  • If your pipeline needs navigation, qlog, or another alias specifically, fetch the URL from segment.files and download it yourself.

End-to-end metadata workflow

This is the real SDK usage pattern that a training pipeline should start from.

from pathlib import Path
import drift

BASE_URL = "https://drift-api-production-1d47.up.railway.app"
DATASET_ID = "jordy-v0"
CACHE_DIR = Path("training-cache")
CACHE_DIR.mkdir(parents=True, exist_ok=True)

client = drift.client.DriftClient(base_url=BASE_URL)

dataset = client.load_dataset(DATASET_ID)
print(f"dataset={dataset.id} routes={dataset.routes_count} segments={dataset.segments_count}")

routes = client.list_routes(DATASET_ID)
segment_ids = [
    segment_ref["segment_id"]
    for route in routes
    for segment_ref in route.segments
]

print(f"resolved {len(segment_ids)} segment references")

for segment_id in segment_ids[:8]:
    segment = client.get_segment(segment_id)
    target = CACHE_DIR / f"{segment.segment_id}.bin"
    segment.download(str(target))
    print(f"downloaded {target}")

client.close()

Training-pipeline pattern

The SDK does not expose download_dataset(...), and it does not materialize a ready-made parquet or CSV training table. A trainer should explicitly:

  1. Query dataset and route metadata.
  2. Expand route segment references into segment IDs.
  3. Download the needed segment files.
  4. Parse those raw artifacts into model-ready tensors or tables.

This is a simplified version of the same control flow used by train_driving.py:

from pathlib import Path
import drift

BASE_URL = "https://drift-api-production-1d47.up.railway.app"
DATASET_ID = "jordy-v0"
CACHE_DIR = Path(".cache/drift-data")
CACHE_DIR.mkdir(parents=True, exist_ok=True)

client = drift.client.DriftClient(base_url=BASE_URL)
routes = client.list_routes(DATASET_ID)

segment_ids = []
for route in routes:
    for segment_ref in route.segments:
        segment_ids.append(segment_ref["segment_id"])

for segment_id in segment_ids[:32]:
    segment = client.get_segment(segment_id)

    if "navigation" in segment.files:
        print("navigation url:", segment.files["navigation"])

    target = CACHE_DIR / f"{segment.segment_id}.bin"
    segment.download(str(target))

client.close()

If your trainer expects .parquet, .csv, .jsonl, or image tensors, you still need a conversion step after download.

PyTorch example

drift.torch.DriftDataset is a metadata dataset. It does not decode frames or controls for you. It resolves segment metadata lazily and can optionally download one primary file per sample.

from torch.utils.data import DataLoader
from drift.torch import DriftDataset

dataset = DriftDataset(
    dataset_id="jordy-v0",
    base_url="https://drift-api-production-1d47.up.railway.app",
    download_dir=".cache/drift-torch",
)

print("segments:", len(dataset))

sample = dataset[0]
print(sample["segment_id"])
print(sample["files"])
print(sample.get("download_path"))

loader = DataLoader(dataset, batch_size=4, shuffle=False)
batch = next(iter(loader))
print(batch["segment_id"])

The payload returned by each item is shaped for a downstream parser:

  • segment_id
  • route_id
  • segment_index
  • files
  • download_url
  • download_path when download_dir is set

TensorFlow / Keras example

drift.tensorflow.DriftDataset is also metadata-first. Iterate over it directly or convert it into a tf.data.Dataset if you want a Keras-compatible input stream.

import tensorflow as tf
from drift.tensorflow import DriftDataset

segments = DriftDataset(
    dataset_id="jordy-v0",
    base_url="https://drift-api-production-1d47.up.railway.app",
)

for item in segments:
    print(item["segment_id"], item["files"])
    break

keras_dataset = segments.to_keras_dataset()
keras_dataset = keras_dataset.take(4)

for item in keras_dataset:
    tf.print(item["segment_id"], item["download_url"])

This is useful when your actual model input pipeline lives in a later mapping stage that downloads or decodes the referenced artifacts.

Common mistake

This is wrong:

client.load_dataset("jordy-v0")

That only returns summary metadata. It does not download data.

This is the correct conceptual flow:

dataset = client.load_dataset("jordy-v0")
routes = client.list_routes(dataset.id)
route = client.get_route(routes[0].route_id)
segment = client.get_segment(route.segments[0]["segment_id"])
segment.download("downloads/example.bin")

Release notes

0.1.1

  • polished package metadata and documentation
  • added release workflow and repository scaffolding
  • improved SDK examples for downloads and framework integrations

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

drift_sdk-0.1.3.tar.gz (10.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

drift_sdk-0.1.3-py3-none-any.whl (8.7 kB view details)

Uploaded Python 3

File details

Details for the file drift_sdk-0.1.3.tar.gz.

File metadata

  • Download URL: drift_sdk-0.1.3.tar.gz
  • Upload date:
  • Size: 10.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for drift_sdk-0.1.3.tar.gz
Algorithm Hash digest
SHA256 acd26e9ad68e4f4ab13c8a6fd8cd44b86ce70f923043b8671ddc9d7da33d36bc
MD5 6b0e9b10397aab704614025028841648
BLAKE2b-256 dae292a9e11107f47fdd48a226fe6163d31844c15d5427b98751a67d471d54a7

See more details on using hashes here.

File details

Details for the file drift_sdk-0.1.3-py3-none-any.whl.

File metadata

  • Download URL: drift_sdk-0.1.3-py3-none-any.whl
  • Upload date:
  • Size: 8.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for drift_sdk-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 13f1edf047df7692dc655ad2be4e97db3e76445f4d86401a43155e856cd5d2fe
MD5 ded8dac7e1ccc32bc65cac363c998351
BLAKE2b-256 ac793f0e7c38b36ac4a85815f2d6d239bb39f901c30e6d92766c7c908870634b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page