Skip to main content

pyHordcoin - Python interface for Hordcoin.jl

This is a python interface for the Julia package Hordcoin.jl which provides methods for finding probability distributions with maximal Shannon entropy given a fixed marginal distribution or entropy up to a chosen order, and to compute the Connected Information.

Installing

Installation is easy, just:

  1. (optional but very advised) create and activate a new virtual environment:
    python -m venv hordcoin
    source hordcoin/bin/activate
    
  2. change directory to the unpacked repository:
    cd /path/to/pyHordcoin
    
  3. install the package:
    1. if you only want to use its components for your measures:
      pip install pyHordcoin
      
    2. if you want to contribute to the development download the sources and then, from the unpacked folder, run:
      pip install -e .[dev]
      

Running

The first time you import the package it will take some time to install the Julia dependencies (and possibly Julia itself), it might take a few minutes. From the second import on it's much faster, although the import of Julia still takes some time. This will be repaid during the maximisation of the entropies.

In case you face a segmentation fault during the first import see the instructions below.

Usage

This package offers an interface to Hordcoin.jl which implements methods that maximize the Shannon entropy of a probability distribution with marginal distribution or entropic constraints and compute the Connected Information.

The input data must satisfy the following requirements:

  • The probability distributions are stored as multidimensional Numpy arrays;
  • Probabilities are non-negative and sum up to 1;
  • OR are provided as (non-normalised) counts;
  • The maximal order of the fixed marginal distributions has to be in [2, n-1], where n is the number of dimensions of the probability distribution.

Connected Information

The main function of the package is connected_information that uses the the maximum entropy with constraints at different orders to compute the Connected Information. It takes as input the probability distribution or the (non-normalised) counts, along with the desired orders of Connected Information and the optimisation method.

When computing multiple Connected Information values for the same probability distribution, it is possible to pass the sizes (desired orders) as an array. This will speed up the process by chaining the computations, thereby reducing the number of maximizations.

If no method is passed, the kind of optimisation is decided by the data type of the input:

  • int input triggers constraints on the marginal entropy and the RawPolymatroid method;
  • float input triggers constraints on the marginal distributions and the Ipfp method.

It is possible to have complete control on the kind of constraints by passing a method explicitly:

  • Gradient, Cone and Ipfp trigger constraints on the marginal distributions with both int and float inputs;
  • RawPolymatroid and GPolymatroid require constraints on the marginal entropy. However, GPolymatroid will raise an error if used with float inputs as it needs the counts to compute the correction.

It's possible to pass a precomputed dictionary of entropies to speed up the computation when using entropic constraints. See note below on the structure of this dictionary.

The basic usage of connected_information is the following:

import pyHordcoin as hc
import numpy as np

counts=np.array([[[1, 2],[3, 4]], [[4, 2], [1, 3]]], dtype=int)
hc.connected_information(counts, 2)

Which will optimise (maximize entropy) constraining the marginal entropies (up to order 2) and should give a result similar to ({2: 0.09310598013744764}, None). The second element of the tuple is None by default. When calling connected_information with the argument full_output=True the second element contains a dictionary with information about the maximally entropic distribution at all orders computed. In case of fixed marginal constraints, each element is the optimized distribution at that order. In case of fixed marginal entropies, each element is the entropy vector that maximises the entropy at that order. This information can be passed on to subsequent calls of functions from this package or used for further analysis.

Warning: for high dimensional probability distribution using full_output=True can consume a lot of memory ($\sim k^n$ for fixed marginal constraints, $k$ states and $n$ dimensions, $\sim2^n$ for fixed marginal entropies).

Notably, the following operations all give the same results:

hc.connected_information(counts, [2])
hc.connected_information(counts, 2, hc.RawPolymatroid())

frequencies = counts.astype(float) / np.sum(counts)
hc.connected_information(frequencies, 2, hc.RawPolymatroid())

Alternatively, it's possible to trigger the marginal distribution constraints with these equivalent lines:

hc.connected_information(frequencies, 2)
hc.connected_information(counts, 2, hc.Ipfp())
hc.connected_information(frequencies, [2], hc.Ipfp())

Or similar results with:

hc.connected_information(frequencies, 2, hc.Gradient())
hc.connected_information(frequencies, 2, hc.Cone())
hc.connected_information(frequencies, 2, hc.Cone(hc.SCS()))
hc.connected_information(frequencies, 2, hc.Cone(hc.Mosek()))

Where the last one requires a Mosek license. (Academic licence easy to obtain at https://www.mosek.com/products/academic-licenses/).

Other useful parameters for the Polymatroid methods are:

  • zhang_yeung: to enable the Zhang-Yeung inequalities complementing the Shannon inequalities and improving the approximation at higher orders (see paper),
  • optimizer: to chose between the SCS and the Mosek optimiser
  • mle_correction: (only RawPolymatroid) enables a rough correction for the finite sample
  • tolerance: (only GPolymatroid) enables a relaxation of the constraints to help convergence (sometimes if fails with the corrected entropies). Note: CI estimate can become negative due to the relaxed constraints.

Other functions

This interface currently implements a single function to access the entropy maximisation. The function maximise_entropy works for marginal constraints entropic constraints selecting the appropriate set of constraints using the same rules as connected_information (with the exception that fixed entropy maximisation from a normalised distribution is not allowed). maximise_entropy takes as an input a probability distribution and the order of marginal distributions to constrain (or the order up to which the marginal entropies must be fixed). The optimiser is an optional parameter that can have further specified parameters (such as the number of iterations, etc.). The function returns the maximum entropy and, in case of fixed marginals, the probability distribution with maximal entropy as an np.ndarray. It's possible to pass a precomputed dictionary of entropies to speed up the computation when using entropic constraints.

The basic usage is the following:

import pyHordcoin as hc

probability_distribution = np.array([[[1/16, 3/16], [3/16, 1/16]], [[1/16, 3/16], [3/16, 1/16]]])
marginal_size = 2
hc.maximise_entropy(probability_distribution, marginal_size)

Running the code with the optional parameter method:

hc.maximise_entropy(probability_distribution, marginal_size, method = hc.Gradient(10, hc.SCS()))

The package also contains one utility function: distribution_entropy computes the information entropy of a probability distribution.

Usage of the functions:

hc.distribution_entropy(probability_distribution)

NOTES

  • All the entropies are measured in bits. This applies also to the precalculated entropies provided by the user.

  • The precalculated entropies must be stored in a dictionary whose keys are tuples of dimensions indexed from 1. These are the dimensions kept in the marginal from which the entropy is calculated. In a 4-D distribution, suppose I sum away dimensions 2 and 4 (marginal=distribution.sum((2,4,))), then the dictionary will contain {(1,3,):-np.sum(marginal*np.log2(marginal))}

A full example with precalculated entropies may look like this:

from itertools import combinations

A = np.random.randint(1000, size=[2, 2, 2]).astype(np.float64)
A /= A.sum()

marginal_entropies = {}
for i in range(4):
    for a in combinations(range(3), i):
        m = tuple(set(range(3)) - set(a))
        tmp = A.sum(m)
        k = tuple(b + 1 for b in a)
        marginal_entropies[k] = hc.distribution_entropy(tmp)

hc.connected_information(
    A, 2, hc.RawPolymatroid(), precalculated_entropies=marginal_entropies,
    full_output=True,
)

Recommendations

When computing with fixed entropies and a small number of samples, the recommended method is the GPolymatroid with Mosek. When the distribution is sampled enough, you can use RawPolymatroid to estimate the entropy with the plug-in estimator. More information can be found in the paper.

Without a MOSEK license, use the Ipfp method (default). It is accurate and fast. It can also be parametrized with the maximum number of iterations, but it is not necessary. The default value is 1000, but it will typically converge in just a few.

When computing with fixed marginal distributions you might want to consider the Cone method with Mosek optimiser. This requires a license to use the MOSEK solver. Without the license, it is possible to use SCS instead, but it is less accurate and slower.

The Gradient method is the slowest and may fail during execution due to limitations of Second Order Cone constraints in solvers.

Compilation issues

If during the first import you experience segmentation fault, this is most likely due to issues during the precompilation of Julia packages (and the disgraceful way Juliacall handles them).

Try the following steps:

  1. Locate the Julia executable that Juliacall is using and the path to the Julia project of Hordcoin
    • If you Python crashes before any output, try ipython (pip install ipython), you should see a lot of output including lines like:
      [juliapkg] Using Julia 1.13.0 at /home/path/to/your/.venv/julia_env/pyjuliapkg/install/bin/julia
      Activating project at `~/path/to/your/.venv/lib/python3.XX/site-packages/pyHordcoin/julia`
      
  2. Assuming the Julia executable is at some/path/julia and the project at other/path/site-packages/pyHordcoin/julia, run in a terminal:
    $ some/path/julia
    julia> using Pkg
    julia> Pkg.activate("other/path/site-packages/pyHordcoin/julia")
    julia> Pkg.instantiate()
    
  3. This time you should see the errors without segfault. Follow the instructions there, often the fix for failing package ProblematicPackage is:
    julia> Pkg.build("ProblematicPackage")
    
  4. Now all the packages should be precompiled, when importing pyHordcoin they'll be already available and no issues should follow.

Reading the documentation

If I didn't upload it somewhere else, from the pyHordcoin folder in the cloned repository run:

pip install -r docs/requirements.txt

to get the right packages.

Now you can read the documentation running:

mkdocs serve

And opening your browser at: http://127.0.0.1:8000/. Notice however that there isn't anything you can't find already in the sources.

How to cite

If you use this code for a scientific publication, please cite:

Tani Raffaelli G., Kislinger J., Kroupa T., and Hlinka J., "HORDCOIN: A Software Library for Higher Order Connected Information and Entropic Constraints Approximation"

Metadata

Release files for pyHordcoin 1.2.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyHordcoin 1.2.2
File Size Uploaded
pyhordcoin-1.2.2.tar.gz 38.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyHordcoin 1.2.2
File Interpreter ABI Platform
pyhordcoin-1.2.2-py3-none-any.whl Python 3 none any Details

Total release size: 71.8 kB

Release files / pyhordcoin-1.2.2.tar.gz

Download URL pyhordcoin-1.2.2.tar.gz
Size 38.9 kB
Tags Source
SHA-256 checksum
How to use checksums
6d54caec5b3c1b437db8fd97b863be6421f6e75981bf6dfc3cf7a827080cfda6
BLAKE2b-256 checksum
How to use checksums
a7a82fdf362433360e0396befe09b2821536e1d8d7b227d2b3fbb42f6c4e0990
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.11

Release files / pyhordcoin-1.2.2-py3-none-any.whl

Download URL pyhordcoin-1.2.2-py3-none-any.whl
Size 32.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a665c8dcf2f6e1c1a176a52aa646978423d1c221092423443621c6c25ae8cab2
BLAKE2b-256 checksum
How to use checksums
f4428a38da0bd18e312ae414ce398ebd9eab98cb21727900ede48fc19ee1118d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.11

Release history Release notifications | RSS feed

This release

1.2.2 This release

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.9

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page