pysimlr
A professional, high-performance PyTorch implementation of SIMLR (Structured Identification of Multimodal Low-rank Relationships), based on the reference R implementation.
Features
- Multi-modal Low-rank Analysis: Integrate multiple data modalities into a shared latent space.
- PyTorch Backend: High-performance numerical operations with GPU support (where available).
- Professional Engineering: Clean, modular structure with comprehensive testing.
- R Parity: Functionality aligned with the reference R implementation.
- Robust SVD: Automated fallback to randomized SVD for numerical stability.
- Flexible Optimization: Multiple optimizers including Hybrid Adam with backtracking line search.
Installation
# From source
git clone https://github.com/stnava/pysimlr.git
cd pysimlr
pip install .
Quick Start
import torch
from pysimlr import simlr
# Generate some dummy multimodal data
n, p1, p2 = 100, 50, 40
x1 = torch.randn(n, p1)
x2 = torch.randn(n, p2)
# Run SIMLR
result = simlr([x1, x2], k=5, iterations=20, verbose=True)
# Access shared latent space
u = result['u'] # n x 5
# Access modality-specific basis matrices
v1, v2 = result['v'] # p1 x 5 and p2 x 5
Development and Testing
The project uses pytest for unit and parity testing.
pytest # fast suite (~2.5 min); this is the default
pytest -m slow # statistical calibration + multi-seed benchmarks only
pytest -m "" # everything
pytest.ini sets pythonpath and excludes the slow marker by default, so no
PYTHONPATH export is needed.
Parity Note
This implementation aims for functional parity with the reference R code. Any identified bugs in the original logic are noted in the source with a BUG comment and addressed in the Python version.
License
Apache-2.0
What Could Go Wrong? (Audit Findings)
Failure modes to watch for, with the diagnostics that detect them.
-
Mixing Alpha Starvation: If
mixing_alpha(scheduled projection) is not annealed to 1.0, the model may perform well while losing interpretability, because the latent scores drift off the Stiefel manifold. Verifyinvariant_orthogonality_defectin your results. -
Latent Collapse: In deep models (NED/NEDPP), insufficient VICReg regularization can leave several latents highly correlated, making the deep consensus redundant. Check
u_off_diag_covin the training diagnostics. -
Sparsity vs. Signal: Aggressive quantile sparsification in
simlr_sparsenesscan zero out predictive features before they reach the deep layers. Monitor the feature importance map. -
Shared-Private Starvation: In NEDPP models, too short a shared-latent schedule lets the private latents absorb the shared variation, giving poor consensus recovery.
-
Orthogonality defect: measured invariant defect after 40 iterations on a synthetic 3-factor problem is roughly
1e-4forconstraint="NewtonSchulz"and3e-4forconstraint="Stiefel". The softorthofamily is looser (~2e-3) by construction — it is a partial projection, not a retraction. If your defect exceeds1e-2, reduce the step size or raise the constraint iteration count. -
The orthogonality weight is a hyperparameter, not a free improvement. On a synthetic 3-factor benchmark, latent recovery was better without the soft orthogonality term (
orthox0) than with it, and degraded further as the weight rose. Tune it against your own criterion. -
The energy is only comparable within a run.
simlrreports the lowest-energy iterate (best_energy,best_iteration), because alternating minimization is not monotone:uis recomputed after each sweep over the views, so the total typically falls steeply for a few iterations and then drifts slowly upward. Convergence usually happens within 5–10 iterations; the defaultiterations=100mostly spends compute after the optimum. -
Interpretability R² is reported both in-sample and cross-validated. Use the
*_cvfields. The in-sample values come from an effectively unregularized fit and approach 1 whenever the predictor is wide relative to the sample count — on pure noise with n=40 and p=35, in-sample R² is 0.92 while cross-validated R² is −367. -
Results depend on whether the optional NSA-Flow backend is installed.
pysimlr.nsa_backend.backend_report()tells you which backend resolved; record it alongside your results. -
Check
simlragainst its own initialization on your data. On well-conditioned multi-view data the SVD initialization is already at the optimum of the objective, and 40 iterations change latent recovery by ~0.0002. The iteration earns its keep only where the shared signal is not each view's leading singular direction — in a regime where it is buried under view-private structure, 60 iterations improved recovery from 0.076 to 0.095.result["best_iteration"]tells you which iterate was actually returned; if it is 0, the optimizer found nothing better than where it started. SeeCORRECTNESS_AUDIT.mdfor the measurements.
Statistical notes
simlr_permandpaths.permutation_testreport add-one permutation p-values,(1 + #{null ≥ observed}) / (1 + n_perms). The smallest attainable value is1 / (1 + n_perms), so choosen_permsfor the resolution you need.energy_type="acc"and"logcosh"sum over the full K×K cross-covariance, including off-diagonal terms, so they reward cross-component coupling as well as per-component alignment."regression"and"nc"are diagonally aligned. On a synthetic benchmark the choice made little difference to latent recovery (0.981 vs 0.980) but the diagonal-only variant had lower cross-component leakage.
Release files for pysimlr 0.2.11
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pysimlr-0.2.11.tar.gz | 276.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pysimlr-0.2.11-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 449.4 kB
Release files / pysimlr-0.2.11.tar.gz
| Download URL | pysimlr-0.2.11.tar.gz |
|---|---|
| Size | 276.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
60db3607bad39d57b5666deae9b030831f1e51547f00a74c079174dae8b752e8
|
|
BLAKE2b-256 checksum How to use checksums |
6ddc4e2e77d0045bb05c58155973518337243790e83ad07057eb6632ce262266
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.12
|
Release files / pysimlr-0.2.11-py3-none-any.whl
| Download URL | pysimlr-0.2.11-py3-none-any.whl |
|---|---|
| Size | 173.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7f28b412a2053d9a1726e9d78fd0e83a5ba5ebed8f760dfb924d8d19c4371db8
|
|
BLAKE2b-256 checksum How to use checksums |
2cf0a68bbf2886b82a656ce1b7dd954a712ed670cd06bb51f0544fce6eab54c0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.12
|