slearn: learning symbolic sequences
slearn is a research package for symbolic sequence generation, symbolic time-series representation, string-distance evaluation, and controlled sequence-learning experiments. It was originally developed around LZW-controlled symbolic strings and LSTM/GRU forecasting; the current experiment suite also compares minimal recurrent, Transformer, efficient-attention, and RWKV-style models under one finite-context prediction protocol.
Install
Core package:
pip install slearn
or:
conda install -c conda-forge slearn
Manuscript experiment environment:
git clone https://github.com/chenxinye/slearn.git
cd slearn
bash exps/scripts/install_experiment_deps.sh
source exps/.venv/bin/activate
Core Features
LZW-controlled symbolic string generation with lzw_string_generator and lzw_string_seeds.
Symbolic time-series transforms including SAX, SAX-TD, eSAX, mSAX, aSAX, ABBA, and fABBA-style representations.
String distances and similarities including Damerau-Levenshtein, Jaro-Winkler, Hamming, cosine, LCS, Dice, and Smith-Waterman variants.
A reproducible neural benchmark for finite-context symbolic prediction and recursive rollout.
Quick Example
from slearn import lzw_string_generator, symbolicML
from slearn.dmetric import normalized_damerau_levenshtein_distance
seed, complexity = lzw_string_generator(
nr_symbols=4,
target_complexity=30,
random_state=7,
)
model = symbolicML(classifier_name="MLPClassifier", ws=4, random_seed=0)
X, y = model.encode(seed * 4)
pred = model.forecast(X, y, step=10, hidden_layer_sizes=(32,), max_iter=500)
target = (seed * 5)[len(seed * 4):len(seed * 4) + 10]
print(complexity)
print(normalized_damerau_levenshtein_distance(target, "".join(pred)))
Benchmark Smoke Test
python exps/symbolic_sequence_benchmark.py --smoke --device cpu
For Slurm runs, submit from exps/:
cd exps
sbatch scripts/run_symbolic_benchmark_slurm.sh
After completion, merge shards and generate figures:
bash scripts/merge_symbolic_results.sh results_symbolic/slurm_<array_job_id>
bash scripts/run_symbolic_visualizations.sh results_symbolic/slurm_<array_job_id>/results_merged.csv
Documentation
The full documentation covers installation, quick start examples, application workflows, experiment reproduction, API references, license, and citations. Build it locally with:
python -m pip install -r docs/requirements.txt
sphinx-build -b html docs/source docs/build/html
Citation
If you use slearn or the LZW symbolic string library, please cite:
@inproceedings{cahuantzi2023comparison,
title = {A Comparison of LSTM and GRU Networks for Learning Symbolic Sequences},
author = {Cahuantzi, Roberto and Chen, Xinye and Guettel, Stefan},
booktitle = {Intelligent Computing},
pages = {771--785},
year = {2023},
publisher = {Springer Nature Switzerland}
}
License
This project is licensed under the MIT License.
Metadata
Release files for slearn 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| slearn-0.3.0.tar.gz | 35.1 kB | Details |
Release files / slearn-0.3.0.tar.gz
| Download URL | slearn-0.3.0.tar.gz |
|---|---|
| Size | 35.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6cd20d8a4a96ac868929f9a2ea92ae1ad3f57b1f3dc11af2c423996bfe43ea46
|
|
BLAKE2b-256 checksum How to use checksums |
bf9e4c9085d3ed1b2a142f9f4c975ba08c57facfc690d6062a889dd05243a2cc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.5
|