Skip to main content

Based on the important perspective that time series are external manifestations of complex dynamical systems, we propose a bimodal generative mechanism for time series data that integrates both symbolic and series modalities. This mechanism enables the unrestricted generation of a vast number of complex systems represented as symbolic expressions $f(\cdot)$ and excitation time series $X$. By inputting the excitation into these complex systems, we obtain the corresponding response time series $Y=f(X)$. This method allows for the unrestricted creation of high-quality time series data for pre-training the time series foundation models.

🔥 News

[Aug. 2026] Recently, multivariate time series generation algorithms based on structural causal models (SCM) have been widely used: TabPFNv3, CauKer, and TiRex-2. Therefore, we specifically reproduced the corresponding generation algorithms in the scm module. See the notebooks.

[Jun. 2026] We extend the learnable white-noise-to-signal simulator family with KalmanFilterSimulator (state-space AR + Kalman filtering) and MarkovSwitchingSimulator (Markov-switching autoregression for regime-dependent dynamics).

[Feb. 2026] Since all stationary time series can be obtained by exciting a linear time-invariant system with white noise, we propose a learnable series generation method based on the ARIMA model. This method ensures the generated series is highly similar to the inputs in autocorrelation and power spectrum density.

[Sep. 2025] Our paper "Synthetic Series-Symbol Data Generation for Time Series Foundation Models" has been accepted by NeurIPS 2025, where SymTime pre-trained on the $S^2$ synthetic dataset achieved SOTA results in fine-tuning of forecasting, classification, imputation and anomaly detection tasks.

🚀 Installation

We have highly encapsulated the algorithm and uploaded the code to PyPI:

pip install s2generator

We used NumPy, Pandas, and Scipy to build the data science environment, Matplotlib for data visualization, and Statsmodels for time series analysis and statistical processing.

✨ Usage

We provide a unified data generation interface SeriesSymbolGenerator, two parameter modules SeriesParams and SymbolParams, as well as auxiliary modules for the generation of excitation time series and complex system. We first specify the parameters or use the default parameters to create parameter objects, and then pass them into our SeriesSymbolGenerator respectively. finally, we can start data generation through the run method after instantiation.

import numpy as np

# Importing data generators object
from s2generator.symbol import SeriesSymbolGenerator, SeriesParams, SymbolParams
from s2generator.utils import plot_symbol_series

# Creating a random number object
rng = np.random.RandomState(0)

# Create the parameter control modules
series_params = SeriesParams()
symbol_params = SymbolParams()  # specify specific parameters here or use the default parameters

# Create an instance
generator = SeriesSymbolGenerator(series_params=series_params, symbol_params=symbol_params)

# Start generating symbolic expressions, sampling and generating series
symbols, inputs, outputs = generator.run(
    rng, input_dimension=1, output_dimension=1, n_inputs_points=256
)

# Print the expressions
print(symbols)
# Visualize the time series
fig = plot_symbol_series(inputs, outputs)

For a user-specified symbolic expression (complex system), use CustomSymbolGenerator: bind f(.) once, then sample excitation X and obtain Y=f(X).

from s2generator.symbol import CustomSymbolGenerator

# Bind a user-specified complex system, then generate via the excitation module
custom = CustomSymbolGenerator("(x_0 add sin(x_0))")
symbols, inputs, outputs = custom.run(rng, n_inputs_points=256)

(73.5 add (x_0 mul (((9.38 mul cos((-0.092 add (-6.12 mul x_0)))) add (87.1 mul arctan((-0.965 add (0.973 mul rand))))) sub (8.89 mul exp(((4.49 mul log((-29.3 add (-86.2 mul x_0)))) add (-2.57 mul ((51.3 add (-55.6 mul x_0)))**2)))))))

The input and output dimensions of the multivariate time series and the length of the sampling sequence can be adjusted in the run method.

rng = np.random.RandomState(512)  # Change the random seed

# Try to generate the multi-channels time series
symbols, inputs, outputs = generator.run(rng, input_dimension=2, output_dimension=2, n_inputs_points=336)

print(symbols)
fig = plot_symbol_series(inputs, outputs)

(-9.45 add ((((0.026 mul rand) sub (-62.7 mul cos((4.79 add (-6.69 mul x_1))))) add (-0.982 mul sqrt((4.2 add (-0.14 mul x_0))))) sub (0.683 mul x_1))) | (67.6 add ((-9.0 mul x_1) add (2.15 mul sqrt((0.867 add (-92.1 mul x_1))))))

Two symbolic expressions are connected by " | ".

🧮 Algorithm

The advantage of $S^2$ data lies in its diversity and unrestricted generation capacity. On the one hand, we can build a complex system with diversity based on binary trees (right); on the other hand, we combine 5 different methods to generate excitation series, as follows:

  • MixedDistribution: Sampling from a mixture of distributions can show the random of time series;
  • ARMA: The sliding average and autoregressive processes can show obvious temporal dependencies;
  • ForecastPFN and KernelSynth: The decomposition and combination methods can reflect the dynamics of time series;
  • IntrinsicModeFunction: The excitation generated by the modal combination method has obvious periodicity.

By generating diverse complex systems and combining multiple excitation generation methods, we can obtain high-quality, diverse time series data without any constraints. For detailed on the data generation process, please refer to our paper or documentation.

🎖️ Citation

If you find this $S^2$ data generation method helpful, please cite the following paper:

@inproceedings{
    wang2026synthetic,
    title={Synthetic Series-Symbol Data Generation for Time Series Foundation Models},
    author={Wenxuan Wang and Kai Wu and Yujian Betterest Li and Dan Wang and Xiaoyu Zhang},
    booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
    year={2026},
    url={https://openreview.net/forum?id=xB1ZNgq0Xp}
}

Release files for S2Generator 0.0.19

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for S2Generator 0.0.19
File Size Uploaded
s2generator-0.0.19.tar.gz 896.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for S2Generator 0.0.19
File Interpreter ABI Platform
s2generator-0.0.19-py3-none-any.whl Python 3 none any Details

Total release size: 1.8 MB

Release files / s2generator-0.0.19.tar.gz

Download URL s2generator-0.0.19.tar.gz
Size 896.7 kB
Tags Source
SHA-256 checksum
How to use checksums
2b13cff199b0b17f7dffb402da2adddbbef20e79403ef6d9f420a12970c38a0c
BLAKE2b-256 checksum
How to use checksums
fe44a80698ed4cf9a5be7af51b3342c714afb324872d27fa9b4a080433a26aea
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release files / s2generator-0.0.19-py3-none-any.whl

Download URL s2generator-0.0.19-py3-none-any.whl
Size 864.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
dfe358f6dc91e5f3f0ab36911bc6870ae71faab7f321d143c840676734945ba0
BLAKE2b-256 checksum
How to use checksums
5b2961161b8d06f53771acde183b4ed494d2210f5485a3b6bc60acadad1f2ed7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release history Release notifications | RSS feed

0.0.23

2 release files

0.0.21

2 release files

0.0.20

2 release files

This release

0.0.19 This release

2 release files

0.0.18

2 release files

0.0.14

2 release files

0.0.13

2 release files

0.0.11

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page