An efficient Python implementation of the Apriori algorithm.

These details have not been verified by PyPI

Project links

Project description

Efficient-Apriori

An efficient pure Python implementation of the Apriori algorithm.

The apriori algorithm uncovers hidden structures in categorical data. The classical example is a database containing purchases from a supermarket. Every purchase has a number of items associated with it. We would like to uncover association rules such as {bread, eggs} -> {bacon} from the data. This is the goal of association rule learning, and the Apriori algorithm is arguably the most famous algorithm for this problem. This repository contains an efficient, well-tested implementation of the apriori algorithm as described in the original paper by Agrawal et al, published in 1994.

The code is stable and in widespread use. It's cited in the book "Mastering Machine Learning Algorithms" by Bonaccorso.

The code is fast. See timings in this PR.

Example

Here's a minimal working example. Notice that in every transaction with eggs present, bacon is present too. Therefore, the rule {eggs} -> {bacon} is returned with 100 % confidence.

from efficient_apriori import apriori
transactions = [('eggs', 'bacon', 'soup'),
                ('eggs', 'bacon', 'apple'),
                ('soup', 'bacon', 'banana')]
itemsets, rules = apriori(transactions, min_support=0.5, min_confidence=1)
print(rules)  # [{eggs} -> {bacon}, {soup} -> {bacon}]

If your data is in a pandas DataFrame, you must convert it to a list of tuples. Do you have missing values, or does the algorithm run for a long time? See this comment. More examples are included below.

Installation

The software is available through GitHub, and through PyPI. You may install the software using pip.

pip install efficient-apriori

Contributing

You are very welcome to scrutinize the code and make pull requests if you have suggestions and improvements. Your submitted code must be PEP8 compliant, and all tests must pass. See list of contributors here.

More examples

Filtering and sorting association rules

It's possible to filter and sort the returned list of association rules.

from efficient_apriori import apriori
transactions = [('eggs', 'bacon', 'soup'),
                ('eggs', 'bacon', 'apple'),
                ('soup', 'bacon', 'banana')]
itemsets, rules = apriori(transactions, min_support=0.2, min_confidence=1)

# Print out every rule with 2 items on the left hand side,
# 1 item on the right hand side, sorted by lift
rules_rhs = filter(lambda rule: len(rule.lhs) == 2 and len(rule.rhs) == 1, rules)
for rule in sorted(rules_rhs, key=lambda rule: rule.lift):
  print(rule)  # Prints the rule and its confidence, support, lift, ...

Transactions with IDs

If you need to know which transactions occurred in the frequent itemsets, set the output_transaction_ids parameter to True. This changes the output to contain ItemsetCount objects for each itemset. The objects have a members property containing is the set of ids of frequent transactions as well as a count property. The ids are the enumeration of the transactions in the order they appear.

from efficient_apriori import apriori
transactions = [('eggs', 'bacon', 'soup'),
                ('eggs', 'bacon', 'apple'),
                ('soup', 'bacon', 'banana')]
itemsets, rules = apriori(transactions, output_transaction_ids=True)
print(itemsets)
# {1: {('bacon',): ItemsetCount(itemset_count=3, members={0, 1, 2}), ...

Project details

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

This version

2.0.6

May 25, 2025

2.0.5

Sep 2, 2024

2.0.4 yanked

Sep 1, 2024

Reason this release was yanked:

Error - .py files were not included

2.0.3

Feb 2, 2023

2.0.2

Jan 16, 2023

2.0.1

Oct 19, 2021

2.0.0

Oct 18, 2021

1.1.1

Mar 19, 2020

1.1.0

Dec 28, 2019

1.0.0

May 20, 2019

0.4.5

Nov 4, 2018

0.4.4

Jun 20, 2018

0.4.3

Jun 20, 2018

0.4.2

Jun 18, 2018

0.4

Jun 18, 2018

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

efficient_apriori-2.0.6.tar.gz (14.7 kB view details)

Uploaded May 25, 2025 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

efficient_apriori-2.0.6-py3-none-any.whl (14.8 kB view details)

Uploaded May 25, 2025 Python 3

File details

Details for the file efficient_apriori-2.0.6.tar.gz.

File metadata

Download URL: efficient_apriori-2.0.6.tar.gz
Upload date: May 25, 2025
Size: 14.7 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for efficient_apriori-2.0.6.tar.gz
Algorithm	Hash digest
SHA256	`5ad8ce4b4f8aee53cc0ef99a804145071bd8bdd652fdfb78cdf2e43831baad6e`
MD5	`13038925d6d93ced6038971d676fb4b1`
BLAKE2b-256	`eceb3a3049d64fae1b51f116a31d37bf81d40cfc5a397f492bc71ee0cc999d39`

See more details on using hashes here.

File details

Details for the file efficient_apriori-2.0.6-py3-none-any.whl.

File metadata

Download URL: efficient_apriori-2.0.6-py3-none-any.whl
Upload date: May 25, 2025
Size: 14.8 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for efficient_apriori-2.0.6-py3-none-any.whl
Algorithm	Hash digest
SHA256	`dad484bb1fb55966a4890f2dc8c7f51515ad3638b64b01580744435c1d46181b`
MD5	`ad49500502dbd4a440950882524043e9`
BLAKE2b-256	`8af2fe62e78214643e2b187fcbf7462dc17eb8071da2f04e1a3349a3297df799`

See more details on using hashes here.

efficient-apriori 2.0.6

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Efficient-Apriori

Example

Installation

Contributing

More examples

Filtering and sorting association rules

Transactions with IDs

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes