Skip to main content

Learn rule lists from data for classification, regression or subgroup discovery

Project description

MDL Rule Lists for prediction and subgroup discovery.

PyPI version PyPI - Python Version License: MIT

This repository contains the code for using rule lists for univariate or multivariate classification or regression and its equivalents in Data Mining and Subgroup Discovery. These models use the Minimum Description Length (MDL) principle as optimality criteria.


This project was written for Python 3.7. All required packages from PyPI are specified in the requirements.txt.

NOTE: This list of packages includes the gmpy2 package.


The latest release can be installed using pip:

pip install rulelist

If you run into issues regarding the gmpy2 package mentioned above, please refer to their documentation for help.

For the current version, you can clone the repository and install the dependencies locally:

git clone
cd RuleList
pip install -r requirements.txt

Example of usage for prediction:

import pandas as pd
from rulelist import RuleList
from sklearn import datasets
from sklearn.model_selection import train_test_split

task = 'prediction'
target_model = 'categorical'

data = datasets.load_breast_cancer()
Y = pd.Series(
X = pd.DataFrame(

X_train, X_test, y_train, y_test = train_test_split(X, Y, test_size = 0.3)

model = RuleList(task = task, target_model = target_model), y_train)

y_pred = model.predict(X_test)
from sklearn.metrics import accuracy_score


Example of usage for subgroup discovery:

import pandas as pd
from rulelist.rulelist import RuleList
from sklearn import datasets

task = 'discovery'
target_model = 'gaussian'

data = datasets.load_boston()
y = pd.Series(
X = pd.DataFrame(

model = RuleList(task = task, target_model = target_model), y)



If there are any questions or issues, please contact me by mail at or open an issue here on Github.


In a machine learning (prediction) context for problems of classification, regression, multi-label classification, multi-category classification, or multivariate regression cite the corresponding bibtex of the first classification application of MDL rule lists:

  title={Interpretable multiclass classification by MDL-based rule lists},
  author={Proen{\c{c}}a, Hugo M and van Leeuwen, Matthijs},
  journal={Information Sciences},

in the context of data mining and subgroup discovery please refer to subgroup lists:

  title={Discovering outstanding subgroup lists for numeric targets using MDL},
  author={Proen{\c{c}}a, Hugo M and Gr{\"u}nwald, Peter and B{\"a}ck, Thomas and van Leeuwen, Matthijs},
  journal={arXiv preprint arXiv:2006.09186},


  title={Robust subgroup discovery},
  author={Proen{\c{c}}a, Hugo Manuel and B{\"a}ck, Thomas and van Leeuwen, Matthijs},
  journal={arXiv preprint arXiv:2103.13686},


Project details

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

rulelist-0.2.0.tar.gz (38.3 kB view hashes)

Uploaded source

Built Distribution

rulelist-0.2.0-py3-none-any.whl (55.3 kB view hashes)

Uploaded py3

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page