Quantitative taxonomy methods of A.A. Lyubishchev (1943) — continuous multivariate classification for biological systematics.
Project description
lyubishchev
Quantitative taxonomy methods of Alexander Alexandrovich Lyubishchev (1890–1972)
pip install lyubishchev
Open a Python terminal and type:
from scipy.spatial.distance import sokalsneath, sokalmichener
Two distance functions named after Sokal, memorialized in the standard scientific computing library. They compute dissimilarity between organisms described as binary character vectors — presence or absence of a trait.
Now type:
from lyubishchev import divergence_coefficient, scatter_ellipse, classify
Lyubishchev was doing the same thing in 1943. With continuous measurements. With full covariance structure. Twenty years before Sokal & Sneath published Principles of Numerical Taxonomy (1963).
He deserves a citation. This library is it.
What this is
Implementations of Lyubishchev's multivariate classification methods for biological taxonomy, as described in his 1943 manuscript:
Lyubishchev, A.A. (1943). Programma obshchey sistematiki [Program of General Systematics].
Manuscript, 22 November 1943. Digitized by ZIN RAS Coleoptera Laboratory.
http://www.zin.ru/animalia/coleoptera/rus/lyubis05.htm
His Western publication, one year before the canonical Sokal & Sneath text:
Lubischew, A.A. (1962). On the use of discriminant functions in taxonomy.
Biometrics, 18(4), 455–477.
Functions
divergence_coefficient(a, b)
Lyubishchev's D from the 1943 manuscript:
D = (M₁ - M₂)² / (σ₁² + σ₂²)
Squared distance between two group means, normalized by pooled variance. Works on continuous measurements. When D > 1.0, groups are cleanly separated. When D < 0.5, you have transgression — the boundary between taxa breaks down.
import numpy as np
from lyubishchev import divergence_coefficient
rng = np.random.default_rng(42)
oleracea = rng.multivariate_normal([3.2, 1.5], [[0.04, 0.01], [0.01, 0.04]], 20)
carduorum = rng.multivariate_normal([3.8, 1.9], [[0.04, 0.01], [0.01, 0.04]], 20)
D = divergence_coefficient(oleracea, carduorum)
print(f"D = {D:.3f}") # D >> 1.0 → clean separation
scatter_ellipse(X, y)
Fit covariance ellipses per class — the geometric object Lyubishchev drew by hand in Fig. 1 of his 1943 manuscript. Each class is represented by its centroid and covariance matrix. Overlap between ellipses is transgression.
from lyubishchev import scatter_ellipse
X = np.vstack([oleracea, carduorum])
y = ['oleracea'] * 20 + ['carduorum'] * 20
ellipses = scatter_ellipse(X, y)
print(ellipses['oleracea']['mean']) # centroid
print(ellipses['oleracea']['cov']) # covariance matrix
transgression(ellipses, class_a, class_b)
Compute whether two scatter ellipses overlap. Lyubishchev called this transgression — the failure of measurement space to cleanly separate two taxa.
from lyubishchev import transgression
result = transgression(ellipses, 'oleracea', 'carduorum')
print(result['transgression']) # False → clean separation
print(result['separation_ratio']) # > 1.0 → well separated
classify(specimen, ellipses)
The Edgeworth-Pearson multivariate probability function — the mathematical core of Lyubishchev's paper nomograms (Fig. 3, 1943 manuscript). Given a specimen's measurements and a set of reference groups, returns posterior probability of belonging to each group.
This is what his field nomograms computed by hand. This is a deployed model on paper.
from lyubishchev import classify
specimen = np.array([3.75, 1.85])
result = classify(specimen, ellipses)
best = max(result, key=lambda k: result[k]['posterior'])
print(best) # 'carduorum'
print(result['carduorum']['posterior']) # e.g. 0.923
print(result['carduorum']['mahalanobis_distance']) # distance to centroid
LyubishchevClassifier — scikit-learn estimator
New in 0.2.0. The 1943 method wrapped as a drop-in scikit-learn classifier, so it plugs straight into Pipeline, GridSearchCV, cross_val_score, and the rest of the ecosystem. Each class is modeled by its centroid and full (regularized) covariance; a specimen is assigned to the class under which its measurements are most probable — the multivariate posterior Lyubishchev computed with paper nomograms.
from lyubishchev import LyubishchevClassifier
clf = LyubishchevClassifier(standardize=True) # z-score features first
clf.fit(X_train, y_train)
labels = clf.predict(X_test)
proba = clf.predict_proba(X_test) # vectorized over all rows
score = clf.score(X_test, y_test)
Parameters:
standardize(defaultFalse) — z-score each feature using the training mean/std before fitting. Use it when measurements live on different scales (mm vs kg) so the covariance structure isn't dominated by one feature.reg_covar(default1e-6) — ridge added to each class covariance diagonal, preventingLinAlgErrorwhen features are collinear or a class has fewer samples than features (the same trickGaussianNBandGaussianMixtureuse).priors(default'uniform') —'uniform'(Lyubishchev's equal-prior assumption),'empirical'(observed class frequencies), or an explicit array.
Input is validated through check_X_y / check_array: NaN, infinite values, single-class targets, and feature-count mismatches at predict time all raise. The estimator passes scikit-learn's check_estimator conformance suite.
from sklearn.model_selection import cross_val_score
cross_val_score(LyubishchevClassifier(standardize=True), X, y, cv=5)
Plotting
New in 0.2.0. A digital reproduction of Lyubishchev's Fig. 1 (Рис. 1) — covariance ellipses per taxon — plus decision-region plots for a fitted classifier. Requires the [plot] extra.
from lyubishchev import scatter_ellipse, LyubishchevClassifier
from lyubishchev.plot import plot_ellipses, plot_classification
ellipses = scatter_ellipse(X, y)
plot_ellipses(ellipses, X, y) # the 1943 figure, digitally
clf = LyubishchevClassifier().fit(X, y)
plot_classification(clf, X, y) # decision regions (2-D)
Comparison with SciPy
scipy.sokalsneath / sokalmichener |
lyubishchev |
|
|---|---|---|
| Input | Binary presence/absence vectors | Continuous measurements |
| Covariance | Not modeled (independence assumed) | Full covariance matrix |
| Distance metric | Simple matching / weighted mismatch | Mahalanobis distance |
| Classification | No | Yes — posterior probability per class |
| Primary source | Sokal & Sneath (1963) | Lyubishchev (1943, 1962) |
Lyubishchev explicitly criticized the independence assumption in his 1943 manuscript. The covariance structure between measurements is not noise — it is information.
Installation
pip install lyubishchev
With plotting support:
pip install lyubishchev[plot]
Requirements
- Python ≥ 3.9
- numpy ≥ 1.24
- scipy ≥ 1.10
- scikit-learn ≥ 1.2
- matplotlib ≥ 3.6 (optional, for the
[plot]extra)
License
MIT
Citation
If you use this library, you are citing Lyubishchev. That is the point.
@misc{lyubishchev1943,
author = {Lyubishchev, Alexander Alexandrovich},
title = {Programma obshchey sistematiki
[Program of General Systematics]},
year = {1943},
note = {Manuscript, 22 November 1943.
Digitized by ZIN RAS Coleoptera Laboratory.
Available at: http://www.zin.ru/animalia/coleoptera/rus/lyubis05.htm}
}
@article{lubischew1962,
author = {Lubischew, A.A.},
title = {On the use of discriminant functions in taxonomy},
journal = {Biometrics},
year = {1962},
volume = {18},
number = {4},
pages = {455--477},
}
@software{lyubishchev_python,
author = {Berdeyev, Akzhan},
title = {lyubishchev: Quantitative taxonomy methods of A.A. Lyubishchev},
year = {2026},
url = {https://github.com/akzhanberdi/lyubishchev},
}
Bad Dog Data — baddogdata.com
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file lyubishchev-0.2.2.tar.gz.
File metadata
- Download URL: lyubishchev-0.2.2.tar.gz
- Upload date:
- Size: 19.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
982916009c0322e875bc9c42530f5a4f09ac900e3e3ef5919c81bc81fd8bd049
|
|
| MD5 |
59daef2c6ef67b77c962858f53c20049
|
|
| BLAKE2b-256 |
05fd7af80da546e85d79aea555d2b0a79a620e93f02fa3b3c0908be18866b4bd
|
File details
Details for the file lyubishchev-0.2.2-py3-none-any.whl.
File metadata
- Download URL: lyubishchev-0.2.2-py3-none-any.whl
- Upload date:
- Size: 15.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c2be4e7d7bca3026f309d3359f44cd169934b454a4e84e306f8679894584150d
|
|
| MD5 |
7bcb3e6885e9a49f904bd36162becb14
|
|
| BLAKE2b-256 |
3d67ed50ad7898b2f75d6f5560908da5b2167561ca6ea06098749f42c20df2ab
|