Tsxtract
High-Performance Time-Series Feature Extraction. Rust Core. Python Ease.
What if time-series feature engineering was 800× faster and used zero defensive memory copies?
Tsxtract is a minimalistic, dependency-light time-series feature extraction library designed to make extracting statistical, temporal, and spectral features across large datasets blazingly fast, memory-efficient, and effortless. It combines a zero-copy Rust engine with a clean, Scikit-Learn-compatible Python interface—ideal for machine learning pipelines, quantitative finance, real-time sensor telemetry, and high-throughput research.
Features • Installation • Quickstart • Benchmarks • Streaming & Sliding Windows • Scikit-Learn Integration • Documentation
Why Tsxtract?
Traditional Python time-series feature libraries (tsfresh, TSFEL, catch22) force a painful trade-off: wait minutes to hours for feature extraction, or risk Out-Of-Memory (OOM) crashes from defensive copies. Tsxtract eliminates that trade-off.
- ⚡ Blazing Fast: Computes up to 800,256 series/second on standard hardware—outperforming
catch22by 820× andtsfreshby 14,000×. - 🧠 Zero-Copy Ingestion: Directly borrows contiguous NumPy buffer pointers via PyO3. No data duplication, no DataFrame melting, and zero intermediate memory ballooning.
- 🎯 33 Curated, High-Signal Features: Avoids the curse of dimensionality. Features are mathematically non-redundant ($|r| < 0.70$ for 83.3% of pairs), spanning distribution moments, quantiles, crossings, spectral power, and permutation entropy.
- 🔓 Full Multi-Core Scaling (GIL-Free): Releases Python’s Global Interpreter Lock (GIL) across the entire computation region, saturating all CPU cores with Rayon's work-stealing scheduler.
- 🔄 Real-Time Streaming Ready: Compute streaming features with constant memory $O(1)$ state updates using the built-in
StreamingExtractor.
Key Features
- Zero-Copy Hybrid Architecture: PyO3 bindings pass 2D NumPy pointer references directly into native Rust SIMD and multi-core loops without copying a single byte.
- Batch-First Parallelism: Processes $N$ series in parallel across hardware threads instead of running serial Python loops.
- Dual API Support: Extract raw 2D NumPy matrices for maximum speed, or labeled Pandas/Polars DataFrames for immediate exploratory analysis.
- Scikit-Learn Compatible: Seamlessly drop
TsxtractTransformerinto anysklearn.pipeline.Pipelineor cross-validation grid search. - Realfft & Branchless Primitives: Preallocated thread-local FFT workspaces and branchless quantile quickselects ensure predictable sub-millisecond execution.
- Streaming & Sliding Windows: Extract rolling features over continuous data streams without reallocating buffers.
Installation
Prebuilt Wheels (Recommended)
Precompiled binary wheels are available on PyPI for Linux (x86_64, aarch64), macOS (Apple Silicon arm64, Intel x86_64), and Windows (x64). No Rust compiler required!
# Core install (NumPy only)
pip install tsxtract-rs
# With optional Pandas DataFrame support
pip install "tsxtract-rs[pandas]"
Using uv or conda:
uv add tsxtract-rs
From Source (Development)
git clone https://github.com/Aamod007/Tsxtract.git
cd Tsxtract
pip install maturin
maturin develop --release
Quickstart
1. Batch Feature Extraction (2D NumPy)
Extract 33 features from 100,000 series in under a second:
import numpy as np
import tsxtractor as tsx
# 1,000 series of 500 time-steps (float64)
X = np.random.randn(1000, 500)
# Extract 33 features (zero-copy, multi-threaded)
features = tsx.extract_features(X)
print("Output shape:", features.shape) # (1000, 33)
print("Feature names:", tsx.feature_names()[:5])
# ['mean', 'std', 'var', 'min', 'max', ...]
2. Labeled Pandas DataFrame
# Returns a labeled pandas DataFrame with clean column headers
df = tsx.extract_features_df(X)
print(df.head())
3. Ragged Series of Different Lengths
# Sequences of varying lengths are supported natively
arr1 = np.random.randn(300)
arr2 = np.random.randn(500)
arr3 = np.random.randn(120)
features = tsx.extract_features([arr1, arr2, arr3])
print(features.shape) # (3, 33)
4. Sliding Windows over a Long Signal
# Extract rolling window features from a 1D continuous sensor stream
signal = np.random.randn(100_000)
windowed_features = tsx.sliding_features(signal, window=256, stride=64)
Scikit-Learn Pipeline
Integrate directly into standard classification, regression, or clustering pipelines:
import numpy as np
import tsxtractor as tsx
from sklearn.base import BaseEstimator, TransformerMixin
from sklearn.pipeline import Pipeline
from sklearn.ensemble import RandomForestClassifier
from sklearn.preprocessing import StandardScaler
class TsxtractTransformer(BaseEstimator, TransformerMixin):
"""Extract 33 Tsxtract features per input row (one series per row)."""
def fit(self, X, y=None):
return self
def transform(self, X):
X_contig = np.ascontiguousarray(X, dtype=np.float64)
return tsx.extract_features(X_contig)
# Assemble end-to-end reproducible pipeline
pipeline = Pipeline([
("features", TsxtractTransformer()),
("scaler", StandardScaler()),
("classifier", RandomForestClassifier(n_estimators=100))
])
# Fit on raw time-series training data (n_samples, time_steps)
pipeline.fit(X_train, y_train)
y_pred = pipeline.predict(X_test)
Streaming & Real-Time Telemetry
Maintain running statistical features in real-time embedded systems or trading loops without recomputing from scratch:
from tsxtractor import StreamingExtractor
# Initialize streaming extractor with buffer capacity
stream = StreamingExtractor(capacity=500)
# Ingest points one by one with O(1) state updates
for tick in incoming_data_feed:
stream.push(tick)
current_features = stream.compute()
Benchmarks
Tested on a 16-core system across 1,000 series of 500 steps (500,000 data points total):
| Library | Features | Runtime | Series / sec | Per-Feature Cost | Speedup vs Tsxtract |
|---|---|---|---|---|---|
Tsxtract (tsxtract-rs) |
33 | 1.25 ms | 800,256 | 0.038 µs | Baseline (1.0×) |
catch22 |
22 | 1,024.8 ms | 976 | 46.58 µs | 820× slower |
TSFEL |
156 | 7,154.0 ms | 140 | 45.86 µs | 5,725× slower |
tsfresh |
777 | 17,683.3 ms | 57 | 22.76 µs | 14,151× slower |
Memory Footprint (100,000 series × 500 steps):
- Tsxtract: +25.2 MiB allocated memory (strictly the output matrix, zero input duplication).
- tsfresh / Pandas: +1,250 MiB memory ballooning due to melted DataFrame indices.
The 33 Curated Features
Tsxtract deliberately computes 33 high-signal, non-redundant features spanning all temporal domains:
- Distribution Moments: Mean, Standard Deviation, Variance, Skewness, Kurtosis.
- Extrema & Spans: Min, Max, Peak-to-Peak Range, Quantiles (q05, q25, median, q75, q95), Interquartile Range (IQR).
- Dynamics & Crossing: Zero Crossing Rate, Mean Crossing Rate, Root Mean Square (RMS), Crest Factor, Median Absolute Deviation (MAD).
- Temporal Differences: Mean Absolute Change, Mean Consecutive Change, Number of Local Peaks.
- Autocorrelation Structure: Lag-1, Lag-2, Lag-3, Lag-5, Lag-10 Autocorrelation.
- Spectral Domain: Energy, Spectral Energy, Dominant Frequency, Spectral Centroid, Spectral Spread, Spectral Roll-off.
- Complexity: Permutation Entropy (order 3, delay 1).
Contributing
Contributions, bug reports, and PRs are welcome!
Please check CONTRIBUTING.md for details on setting up the local Rust/Python development environment and running the benchmark suites.
License
Distributed under the MIT License. See LICENSE for details.
Built with 🦀 Rust & 🐍 Python by Aamod.
Metadata
Release files for tsxtract-rs 0.3.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tsxtract_rs-0.3.1.tar.gz | 4.7 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tsxtract_rs-0.3.1-cp310-abi3-win_amd64.whl | CPython 3.10 | abi3 | Windows x86-64 | Details |
Total release size: 5.3 MB
Release files / tsxtract_rs-0.3.1.tar.gz
| Download URL | tsxtract_rs-0.3.1.tar.gz |
|---|---|
| Size | 4.7 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6132679e6072cca191a9b1e68adf5fd73862eb49bbcb77333f131de625e31f0e
|
|
BLAKE2b-256 checksum How to use checksums |
43602852f7ec272be401c01713d101b454f8a3540fd6171456bf8da56cb45a6d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.14.1
|
Release files / tsxtract_rs-0.3.1-cp310-abi3-win_amd64.whl
| Download URL | tsxtract_rs-0.3.1-cp310-abi3-win_amd64.whl |
|---|---|
| Size | 535.0 kB |
| Tags | CPython 3.10 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
7ac5584d11e715d1cdf50b864e0b0656c97f0eef4b282bd199cac0482a5ca0ed
|
|
BLAKE2b-256 checksum How to use checksums |
45ab8a590c72dfb0278181ca74bafa7a5b65efe71c2fea080814a89ecac140fe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.14.1
|