Skip to main content


Comparison-based Machine Learning in Python

DOI PyPI version Documentation Test status Test Coverage

Comparison-based learning methods are machine learning algorithms using similarity comparisons ("A and B are more similar than C and D") instead of featurized data.

from sklearn.datasets import load_iris
from sklearn.model_selection import cross_val_score

from cblearn.datasets import make_random_triplets
from cblearn.embedding import SOE

X = load_iris().data
triplets = make_random_triplets(X, result_format="list-order", size=2000)

estimator = SOE(n_components=2)
# Measure the fit with scikit-learn's cross-validation
scores = cross_val_score(estimator, triplets, cv=5)
print(f"The 5-fold CV triplet error is {sum(scores) / len(scores)}.")

# Estimate the scale on all triplets
embedding = estimator.fit_transform(triplets)
print(f"The embedding has shape {embedding.shape}.")

Getting Started

Contribute

We are happy about your bug reports, questions or suggestions as Github Issues and code or documentation contributions as Github Pull Requests. Please see our Contributor Guide.

Related packages

There are more Python packages for comparison-based learning:

  • metric-learn is a collection of algorithms for metric learning. The weakly supervised algorithms learn from triplets and quadruplets.
  • salmon is a package for efficiently collecting triplets in crowd-sourced experiments. The package implements ordinal embedding algorithms and sampling strategies to query the most informative comparisons actively.

Authors and Acknowledgement

cblearn was initiated by current and former members of the Theory of Machine Learning group of Prof. Dr. Ulrike von Luxburg at the University of Tübingen.

Author: David-Elias Künstle

Maintainers: Guillermo Aguilar, Vivek Anand

Contributors: Alexander Conzelmann, Michaël Perrot, Mojtaba Barzegari

We want to thank all the contributors here on GitHub. This work has been supported by the Machine Learning Cluster of Excellence, funded by EXC number 2064/1 - Project number 390727645. The authors would like to thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting David-Elias Künstle.

License

This library is free to use, share, and adapt under the MIT License conditions.

Citation

Please cite our JOSS paper if you publish work using cblearn:

Künstle et al., (2024). cblearn: Comparison-based Machine Learning in Python. Journal of Open Source Software, 9(98), 6139, https://doi.org/10.21105/joss.06139

@article{Künstle2024, 
    doi = {10.21105/joss.06139}, 
    url = {https://doi.org/10.21105/joss.06139}, 
    year = {2024}, 
    publisher = {The Open Journal}, 
    volume = {9}, number = {98}, pages = {6139}, 
    author = {David-Elias Künstle and Ulrike von Luxburg}, 
    title = {cblearn: Comparison-based Machine Learning in Python}, 
    journal = {Journal of Open Source Software} 
} 

Changelog

Upcoming

0.4

  • Feature: embedding.LORE, a low-rank ordinal embedding that estimates the intrinsic dimensionality
  • Feature: Support for Python 3.12 and 3.13
  • Improvement: Compatibility with scikit-learn 1.6 and newer, which replaced the estimator tag dictionary by the __sklearn_tags__ API. scikit-learn 1.6 is now the minimum supported version.
  • Improvement: Readable errors for invalid embedding dimensions, raised in fit instead of __init__, as required by scikit-learn
  • Improvement: Validation of test_dimensions in embedding.estimate_dimensionality_cv
  • Improvement: The food, nature and vogue datasets are downloaded from OSF mirrors, since the original hosts no longer serve the archives
  • Improvement: Externally hosted archives are pinned to fixed files or commits, so that their checksums no longer change on every upstream push
  • Fix: Gradient of the STE embedding, which weighted each triplet by P * (1 - P) instead of (1 - P)
  • Fix: estimate_dimensionality_cv returns the estimated dimension as a Python int
  • Fix: utils.torch_device returns and validates an explicit device instead of None
  • Fix: Sparse input is detected with scipy.sparse.issparse, so that the newer sparray classes are recognized
  • Fix: The numpy converter for rpy2 is registered directly, since numpy2ri.activate raises in rpy2 3.5.12 and newer
  • Others: Extended unit tests and coverage, seeded triplet sampling in the test suite

0.3

  • Feature: JOSS paper
  • Feature: Quickstart guide in documentation
  • Feature: Data point sampling from manifolds.
  • Improvement: Extended documentation
  • Improvement: cblearn logo and new style in documentation
  • Improvement: Filter invalid responses in datasets
  • Improvement: Full compatibility to sklearn estimator tests

0.2

  • Improvement: Extended documentation
  • Feature: embedding.estimate_dimensionality_cv function (Künstle et al., 2022)
  • Fix: Avoid numpy deprecation warning for scalar variables in fetch_similarity_matrix
  • Fix: Various errors in the examples
  • Fix: Minor errors in the unit tests
  • Others: Updated dependencies

0.1

0.1.2

  • support python 3.11
  • update core dependencies

0.1.1

  • Minor fixes in the documentation.
  • Adapt loading of food and imagenet dataset to solve problems caused by changes in externally hosted files

0.1.0

  • Support python 3.9 and 3.10.
  • Introduce semantic versioning
  • Publish to PyPI

MIT License

Copyright (c) 2020-2021 The cblearn developers.

Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.

Metadata

Release files for cblearn 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cblearn 0.4.0
File Size Uploaded
cblearn-0.4.0.tar.gz 422.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cblearn 0.4.0
File Interpreter ABI Platform
cblearn-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 540.0 kB

Release files / cblearn-0.4.0.tar.gz

Download URL cblearn-0.4.0.tar.gz
Size 422.1 kB
Tags Source
SHA-256 checksum
How to use checksums
a2788bb03d8da7b6c9c30c178ca7d465a14de16a6d29f527bbf5dd01207c603e
BLAKE2b-256 checksum
How to use checksums
64bb876a6c59fce0f30fb0a20840c7b5faab6b6a1a4e71dc78eb4900c082a58d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / cblearn-0.4.0-py3-none-any.whl

Download URL cblearn-0.4.0-py3-none-any.whl
Size 118.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8224e02c56013ceb5452a588d80b75a66d22ff5a351cf23f46715bc1fa7c73a5
BLAKE2b-256 checksum
How to use checksums
8f63bd05b60dc070e095b1d19da9de7ae973bd98163dd84da75fb0a4f70507a6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page