Skip to main content

A preprocessing library for foundation time series models

Project description

Time Series Preprocessing Library

A comprehensive library for building streaming data pipelines for time series datasets, providing tools for downloading, transforming, and combining time series data in real-time.

Features

  • Streaming Dataset Transforms: Build composable data pipelines that process time series data as it streams
  • Dataset Downloaders: Download datasets from Hugging Face Hub with caching support
  • Synthetic Data Generation: Generate synthetic time series data for testing and development
  • Flexible Data Pipeline Builder: Chain multiple transforms together using a fluent builder pattern
  • PyTorch Integration: Full compatibility with PyTorch's IterableDataset interface
  • Type Safety: Comprehensive type annotations for better development experience

Installation

Using poetry:

poetry install

Using pip:

pip install .

Core Components

Dataset Transforms

The library provides a rich set of streaming dataset transforms that can be chained together:

Basic Transforms

  • TransformingDataset: Apply arbitrary functions to each item in a dataset
  • BatchingIterableDataset: Group items into batches with configurable batch sizes
  • UnbatchingIterableDataset: Flatten batched datasets back to individual items
  • SlidingWindowIterableDataset: Create sliding windows over time series data

Combining Transforms

  • ConcatDataset: Concatenate multiple datasets sequentially
  • CombiningDataset: Combine multiple datasets using custom operations (e.g., element-wise addition)
  • ProbabilisticMixingDataset: Mix datasets with configurable probabilities and seeding

Pipeline Builder

  • Builder: Fluent interface for building complex data pipelines

Dataset Downloaders

  • HuggingFaceDownloader: Download datasets from Hugging Face Hub with automatic caching
  • GiftEvalWrapperDataset: Wrapper for Salesforce's GiftEval pretraining datasets

Synthetic Data

  • LinearTrendDataset: Generate synthetic time series with linear trends and configurable noise

Serialization

  • serialize_tensor_stream: Save tensor streams to disk in sharded format
  • SerializedTensorDataset: Load serialized tensor streams with lazy loading support

Usage Examples

Basic Transform Pipeline

from preprocessing.transform.dataset_builder import Builder
from preprocessing.transform.batching_dataset import BatchingIterableDataset
from preprocessing.transform.transforming_dataset import TransformingDataset

# Create a simple pipeline
pipeline = (
    Builder(your_dataset)
    .batch(batch_size=32)
    .map(lambda x: x * 2)  # Double all values
    .build()
)

# Iterate over the pipeline
for batch in pipeline:
    print(batch.shape)  # (32, sequence_length, features)

Combining Multiple Datasets

from preprocessing.transform.combining_dataset import CombiningDataset

# Combine two datasets element-wise
def add_operation(x, y):
    return x + y

combined = CombiningDataset([dataset1, dataset2], op=add_operation)

for result in combined:
    print(result)  # x + y for each pair

Concatenating Datasets

from preprocessing.transform.concat_dataset import ConcatDataset

# Concatenate datasets sequentially
concatenated = ConcatDataset([dataset1, dataset2, dataset3])

for item in concatenated:
    # Items from dataset1, then dataset2, then dataset3
    print(item)

Probabilistic Mixing

from preprocessing.transform.probabilistic_mixing_dataset import ProbabilisticMixingDataset

# Mix datasets with custom probabilities
mixed = ProbabilisticMixingDataset(
    datasets={"train": train_data, "val": val_data},
    probabilities={"train": 0.8, "val": 0.2},
    seed=42  # For reproducibility
)

for item in mixed:
    # 80% chance from train, 20% chance from val
    print(item)

Sliding Windows

from preprocessing.transform.sliding_window_dataset import SlidingWindowIterableDataset

# Create sliding windows over time series
windowed = SlidingWindowIterableDataset(
    dataset=your_dataset,
    window_size=100,
    step=50
)

for window in windowed:
    print(window.shape)  # (100, features)

Downloading Datasets

from preprocessing.downloader.huggingface import HuggingFaceDownloader
from preprocessing.config import DatasetConfig

# Configure dataset download
config = DatasetConfig(
    name="air-passengers",
    repo_id="duol/airpassengers",
    files=["AP.csv"],
    cache_dir="data/cache"
)

# Download dataset
downloader = HuggingFaceDownloader(config)
data = downloader.download()

Synthetic Data Generation

from preprocessing.synthetic.linear_trend import LinearTrendDataset

# Generate synthetic time series
synthetic_data = LinearTrendDataset(
    sequence_length=1000,
    num_sequences=100,
    trend_slope=0.01,
    noise_std=0.1,
    seed=42
)

for sequence in synthetic_data:
    print(sequence.shape)  # (1000, 1)

Serialization

from preprocessing.serialization.serialize import serialize_tensor_stream
from preprocessing.serialization.deserialize import SerializedTensorDataset

# Save dataset to disk
serialize_tensor_stream(
    dataset=your_dataset,
    output_dir="data/serialized",
    max_tensors_per_file=1000
)

# Load dataset from disk
loaded_dataset = SerializedTensorDataset(
    filepaths=["data/serialized/shard_00000.pt"],
    lazy=True  # Load on-demand
)

Project Structure

preprocessing/
├── common/              # Common types and utilities
│   └── tensor_dataset.py
├── config/              # Configuration management
│   ├── __init__.py
│   └── examples/
├── downloader/          # Dataset downloaders
│   ├── huggingface.py
│   └── gift_eval.py
├── serialization/       # Data serialization
│   ├── serialize.py
│   └── deserialize.py
├── synthetic/           # Synthetic data generation
│   └── linear_trend.py
└── transform/           # Dataset transforms
    ├── batching_dataset.py
    ├── combining_dataset.py
    ├── concat_dataset.py
    ├── dataset_builder.py
    ├── probabilistic_mixing_dataset.py
    ├── sliding_window_dataset.py
    ├── transforming_dataset.py
    └── unbatching_dataset.py

Development

Setup Development Environment

poetry install --with dev

Run Tests

poetry run pytest

Run Tests with Coverage

poetry run pytest --cov=preprocessing

Key Design Principles

  1. Streaming First: All transforms work with streaming data, enabling processing of large datasets that don't fit in memory
  2. Composability: Transforms can be easily chained together using the Builder pattern
  3. Type Safety: Comprehensive type annotations for better development experience
  4. PyTorch Integration: Full compatibility with PyTorch's data loading ecosystem
  5. Reproducibility: Built-in support for random seeds and deterministic operations
  6. Flexibility: Support for custom operations and transformations

License

MIT License - see LICENSE file for details

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ts_preprocessing-0.1.1.tar.gz (10.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ts_preprocessing-0.1.1-py3-none-any.whl (15.5 kB view details)

Uploaded Python 3

File details

Details for the file ts_preprocessing-0.1.1.tar.gz.

File metadata

  • Download URL: ts_preprocessing-0.1.1.tar.gz
  • Upload date:
  • Size: 10.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.13

File hashes

Hashes for ts_preprocessing-0.1.1.tar.gz
Algorithm Hash digest
SHA256 2feae7af43098b062d50cd06c95a899dcdcf6b1e93a93f437ce0ea89a66d128e
MD5 bd9a095993f60a7274cbbfc1c7896bff
BLAKE2b-256 8b648436d81fef527d8d63732e47d16db5b8b498bf17a95b5b3286bcff4dd870

See more details on using hashes here.

File details

Details for the file ts_preprocessing-0.1.1-py3-none-any.whl.

File metadata

File hashes

Hashes for ts_preprocessing-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 e923d976218d5a4b3dc0c7e9c4edde229ed813d0512438e622df030eb3b77e33
MD5 a5309e717a26db70dff9f790ea961e07
BLAKE2b-256 e7c78749149d78c8763bf362b4e30a1b2f77fd9a2f0f4af7088533a6d7852484

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page