Skip to main content

Tool to generate biclustering and triclustering datasets

Project description

Nclustgen

Nclustgen is a python tool to generate biclustering and triclustering datasets programmatically.

It wraps two java packages G-Bic, and G-Tric, that serve as backend generators. If you are interested on a GUI version of this generator or on using this generator in a java environment check out those packages.

This tool adds some functionalities to the original packages, for a more fluid interaction with python libraries, like:

  • Conversion to numpy arrays
  • Conversion to sparse tensors
  • Conversion to networkX or dgl n-partite graphs

Installation

This tool can be installed from PyPI:

pip install nclustgen

NOTICE: Nclustgen installs by default the dgl build with no cuda support, in case you want to use gpu you can override this by installing the correct dgl build, more information at: https://www.dgl.ai/pages/start.html.

Getting started

Here are the basics, the full documentation is available at: http://nclustgen.readthedocs.org.

## Generate biclustering dataset

from nclustgen import BiclusterGenerator

# Initialize generator
generator = BiclusterGenerator(
     dstype='NUMERIC',
     patterns=[['CONSTANT', 'CONSTANT'], ['CONSTANT', 'NONE']],
     bktype='UNIFORM',
     in_memory=True,
     silence=True
)

# Get parameters
generator.get_params()

# Generate dataset
x, y = generator.generate(nrows=50, ncols=100, nclusters=3)

# Build graph
graph = generator.to_graph(x, framework='dgl', device='cpu')

# Save data files
generator.save(file_name='example', single_file=True)

## Generate triclustering dataset

from nclustgen import TriclusterGenerator

# Initialize generator
generator = TriclusterGenerator(
     dstype='NUMERIC',
     patterns=[['CONSTANT', 'CONSTANT', 'CONSTANT'], ['CONSTANT', 'NONE', 'NONE']],
     bktype='UNIFORM',
     in_memory=True,
     silence=True
)

# Get parameters
generator.get_params()

# Generate dataset
x, y = generator.generate(nrows=50, ncols=100, ncontexts=10, nclusters=25)

# Build graph
graph = generator.to_graph(x, framework='dgl', device='cpu')

# Save data files
generator.save(file_name='example', single_file=True)

License

GPLv3

Documentation

The documentation is available at: https://nclustgen.readthedocs.org.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Files for nclustgen, version 1.0.1
Filename, size File type Python version Upload date Hashes
Filename, size nclustgen-1.0.1-py3-none-any.whl (25.3 MB) File type Wheel Python version py3 Upload date Hashes View
Filename, size nclustgen-1.0.1.tar.gz (37.0 kB) File type Source Python version None Upload date Hashes View

Supported by

AWS AWS Cloud computing Datadog Datadog Monitoring DigiCert DigiCert EV certificate Facebook / Instagram Facebook / Instagram PSF Sponsor Fastly Fastly CDN Google Google Object Storage and Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Salesforce Salesforce PSF Sponsor Sentry Sentry Error logging StatusPage StatusPage Status page