QABBA
QABBA is a production-oriented Python package for Quantized ABBA, a symbolic time-series representation that combines adaptive piecewise-linear approximation, symbolic digitization, low-bitwidth center quantization, and optional lossless symbol-layer coding.
The package turns long numerical time series into compact symbolic sequences with quantized reconstruction parameters. This is useful when storage, transmission, and downstream symbolic analysis matter: edge sensing, long time-series archives, approximate reconstruction, motif-oriented mining, and LLM-facing time-series tokenization.
Features
- Compact symbolic representations: convert dense float time series into adaptive symbol sequences and a small table of quantized centers.
- Storage-aware by design: inspect total bits, storage ratio, compression factor, and symbol-layer storage with a unified API.
- Optional lossless symbol codecs: choose
none,fixed,huffman,lzw,zlib,gzip,bz2,lzma, or optionalzstdfor second-stage symbol coding. - Scikit-learn style API: use familiar
fit,transform,fit_transform,inverse_transform,get_params, andset_params. - Cython acceleration with fallback: build fast kernels when Cython and a compiler are available; otherwise use pure Python implementations.
- Research-to-software bridge: package APIs follow the paper methodology while keeping experimental scripts separate under
exps.
Installation
Install the released package:
pip install qabba
Install optional extras:
pip install "qabba[plot,zstd]"
Install from source for development:
git clone https://github.com/chenxinye/qabba_code.git
cd qabba_code
pip install -e ".[dev]"
To force the pure Python kernels:
QABBA_DISABLE_CYTHON=1 pip install -e .
Quick Start
import numpy as np
from qabba import QABBA, mean_squared_error
x = np.sin(np.linspace(0, 8 * np.pi, 512))
model = QABBA(
tol=0.05,
alpha=0.4,
bits_for_len=8,
bits_for_inc=12,
symbol_codec="fixed",
)
symbols = model.fit_transform(x)
x_hat = model.inverse_transform(symbols)
print("first symbols:", "".join(symbols[:20]))
print("storage ratio:", model.storage_.storage_ratio)
print("compression factor:", model.storage_.compression_factor)
print("MSE:", mean_squared_error(x, x_hat))
Common Workflows
1. Fit, Encode, and Reconstruct
model = QABBA(tol=0.05, alpha=0.5)
symbols = model.fit_transform(x)
x_hat = model.inverse_transform(symbols)
2. Compare Symbol Codecs
model = QABBA(tol=0.05, alpha=0.5).fit(x)
for codec in ["none", "fixed", "lzw", "huffman"]:
stats = model.encode_symbols(codec=codec)
print(codec, stats.total_bits, stats.bits_per_symbol)
The symbol codec is lossless with respect to the symbolic sequence; it changes storage accounting, not reconstruction error.
3. Batch Multiple Time Series
X = np.vstack([
np.sin(np.linspace(0, 6 * np.pi, 256)),
np.cos(np.linspace(0, 6 * np.pi, 256)),
])
model = QABBA(tol=0.05, alpha=0.5, n_jobs=2)
symbols = model.fit_transform(X)
reconstructions = model.inverse_transform(symbols)
Rows of a two-dimensional array are treated as independent univariate time series that share one learned symbolic codebook.
4. Inspect Storage
storage = model.storage_
print(storage.total_bits)
print(storage.symbol_bits)
print(storage.center_bits)
print(storage.storage_ratio)
print(storage.compression_factor)
Applications
QABBA is designed for workflows where a time series should be both compact and semantically usable after compression.
- Edge and IoT sensing: transmit symbolic sequences and quantized centers instead of dense float streams.
- Scientific monitoring: archive long signals while preserving trend and shape information for approximate reconstruction.
- Symbolic time-series mining: feed downstream algorithms for motifs, anomaly screening, clustering, and pattern matching.
- LLM-oriented time-series processing: expose adaptive symbolic sequences as text-like inputs for language-model pipelines.
- Compression research: compare rate-distortion behavior against classical dimensionality reduction and lossy scientific compressors.
QABBA is not intended to replace strict error-bounded floating-point compressors when every value must be recovered under a prescribed absolute or relative error bound. It is strongest when very compact symbolic structure is valuable.
Documentation
The ReadTheDocs source lives in docs/source. Build it locally with:
cd docs
make html
The documentation includes installation, quick start, examples, applications, mathematical formulation, API reference, and license pages.
Development
Run the test suite:
pytest
Build a wheel without dependencies:
python -m pip wheel . -w dist --no-deps
Build the documentation:
cd docs
make clean
make html
Continuous Integration
The GitHub Actions workflow in .github/workflows/ci.yml runs on pull requests and pushes to master. It checks:
- package import and API tests on Python 3.9, 3.10, 3.11, and 3.12;
- source distribution and wheel builds;
- Sphinx documentation builds with
make html.
Cite
@misc{carson2025quantizedsymbolictimeseries,
title={Quantized symbolic time series approximation},
author={Erin Carson and Xinye Chen and Fei He and Cheng Kang},
year={2026},
eprint={2411.15209},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2411.15209},
}
License
QABBA is released under the MIT License. See LICENSE.
Release files for qabba 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| qabba-0.2.0.tar.gz | 561.7 kB | Details |
Release files / qabba-0.2.0.tar.gz
| Download URL | qabba-0.2.0.tar.gz |
|---|---|
| Size | 561.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5b76eea02f81aae211fce97b4cf939b4dacdc1ced27088b0a740332ea6559723
|
|
BLAKE2b-256 checksum How to use checksums |
9509199b1fdf164ce7ece36f328f73852e39ef56791938a4797f2e48adc25fdb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.5
|