RobustSignalMaker
NaN-aware, leakage-free stability selection for scientific signals and spectra. RSM identifies the important parts of a series (time series, spectra, diffraction patterns, mass spectra, sensor recordings) to retain and removes the rest, returning a reproducible region of the sampling axis and an honest estimate of what that region can predict.
Install
pip install robustsignalmaker
Python 3.9 or newer. Three required dependencies, all of which you almost
certainly already have: numpy>=1.24, scipy>=1.10, scikit-learn>=1.3.
There is no deep-learning dependency; the gated engine is numpy with
hand-derived gradients.
One optional extra, for writing result tables through pandas rather than the built-in csv writer:
pip install "robustsignalmaker[tables]"
Verify the install:
import robustsignalmaker as rsm
print(rsm.__version__)
The package ships a py.typed marker, so type checkers will use its
annotations without a stub package.
Why RSM
Most band-selection tools report a single selection and leave you to guess how much of it would survive a different sample. RSM's deliverable is the reproducible band set and an honest account of what discarding the rest costs:
- Leakage-safe stability selection is the product. Selection is
aggregated over resamples into stability frequencies
(Meinshausen-Buhlmann style, with Shah-Samworth complementary pairs
available), wrapped in nested cross-validation where every fitted
statistic lives inside the training fold, and compared against a matched
full-signal baseline with a preserved / significantly-better /
significantly-worse verdict. On six real datasets every selection
compresses hard and returns
preserved: tecator 7 of 20 segments, corn 4 of 35, ovarian SELDI-TOF 6 of 64, RRUFF Raman 8.5 of 64 per fold, BasicMotions 2.75 of 10, and four of eighteen bond-angle bands on a gold nanoparticle dataset.
Real instrument data is where that gets hard, because a selection frequency is not well defined when different samples observe different parts of the axis. The rest of the machinery exists to make it well defined:
- No value is ever fabricated. Every pooling, lens, and model operation is renormalised over the evidence that actually exists (validity masks travel with the data end to end). Detector dropout, saturation, instrument-range gaps and not-detected zeros are absence, not zeros.
- Three-valued verdicts. Every candidate band ends as selected, rejected, or UNASSESSABLE. A region nobody measured can never read as "stably rejected", and every selection frequency is reported beside its observation support.
- Honesty guards. Samples with too little evidence are refused, not guessed. If the missingness pattern alone predicts the target (a non-ignorable, instrument-linked confound), RSM warns and stamps the result; imputation-first pipelines destroy exactly that alarm.
- Your choice of predictor. Selection is done by the gated head, but the
final model on the chosen bands is yours:
refit_model=takeslin,rdg,las,eln,log,svm,rf,xgb,mlp, orplsfor PLS regression and PLS-DA. Swapping it never changes the selection, and the refusal rule survives: a random forest refuses the same thin samples the head refuses rather than predicting from imputed values. - Lens marginalisation. Band importance depends on how a signal is represented (raw intensity, derivatives, smoothed scales). RSM can marginalise the selection over a lens ensemble and reports how much the lenses agree.
Input contract
X : float array (n_samples, n_channels, n_points)
All samples share one sampling axis of length n_points. A single-channel
spectrum is n_channels=1, so reshape a plain (n_samples, n_points)
matrix with X[:, None, :].
Missingness is declared in one of two ways, and they are equivalent:
import numpy as np
# 1. NaN in X means "not measured"
X = np.random.default_rng(0).normal(size=(20, 1, 64))
X[3, 0, 10:14] = np.nan
# 2. or pass an explicit boolean validity mask, True = observed
V = np.isfinite(X)
V[5, 0, 20:24] = False # censored, though a value is present
Pass the mask as the third argument to fit, or leave it out and let the
NaNs speak: sel.fit(X, y) and sel.fit(X, y, V) are the same call.
Use the explicit mask when a missing value is not representable as NaN, for example a mass-spectrometry zero that means "not detected" rather than "intensity zero", or a saturated detector reading you want treated as censored.
y may be regression, binary or multiclass, and may contain NaN for
unlabelled samples: they are excluded from every split, loss, resample and
score, and counted in the result rather than silently absorbed.
Multi-channel series get one gate per band shared across channels, with per-channel missingness handled underneath the gate.
Quick start
This example is self-contained and runs as written.
import numpy as np
from robustsignalmaker import BootstrapMaskSelector, make_signal_control
# a control with a known informative band, so the answer is checkable
c = make_signal_control(n=200, n_points=256, task="binary", seed=0)
X, y = c.X, c.y # X is (200, 1, 256)
X = X.copy()
X[:80, 0, 180:210] = np.nan # 80 samples never saw this stretch
sel = BootstrapMaskSelector(task="binary", segment=16,
n_bootstrap=30, seed=0).fit(X, y)
print("kept bands :", sel.selected_segments_)
print("point indices :", sel.selected_points()[:12], "...")
print("coverage :", round(sel.coverage(), 3))
truth = sorted(set(np.flatnonzero(c.informative) // 16))
print("truth was :", truth) # the planted band, for comparison
Reading the output
This is the part worth five minutes, because a selection frequency without its support is the thing RSM exists to stop you reporting.
import numpy as np
for k in range(sel.grid.n_segments):
print(f"band {k:3d} pi={sel.pi_[k]!s:>6} "
f"support={sel.support_[k]:.2f} {sel.verdicts_[k]}")
pi_[k]is the selection frequency for bandk, computed only over the resamples in which that band was ASSESSABLE. It isnan, never0, when too few resamples could assess it.support_[k]is that denominator: the fraction of resamples that could assess the band. Readpi_andsupport_together. Api_of 0.9 at support 0.2 is a much weaker statement than the same value at support 1.0.verdicts_[k]is"selected","rejected"or"unassessable". The third value is the point of the library: absence of evidence is reported as such, not as evidence of absence.sel.stability_report()gives the run-level diagnostics in one dict:coverage,assessable_universe,unassessable_fraction,ambiguous_fraction(bands stuck between 0.2 and 0.8, so no threshold helps),n_failed_fits,strataanddegenerate.
ambiguous_fraction is the go/no-go number. If it is high, pi_ is diffuse
rather than bimodal, and no choice of tau will produce a trustworthy set;
the honest move is to report that rather than to pick a threshold.
The honest verdict: nested cross-validation
BootstrapMaskSelector tells you which bands are stable. It does not tell
you what the selection is worth, because it saw all the data. For that, wrap
it:
from robustsignalmaker import NestedCV
result = NestedCV(k_outer=5, segment=16, random_state=0).run(X, y)
s = result.summary()
print(s["score_name"], round(s["score_mean"], 4),
"vs full signal", round(s["baseline_score_mean"], 4))
print("verdict:", s["verdict"],
" bands kept per fold:", round(s["mean_n_selected"], 2))
print("missingness informative:", s["missingness_informative"])
result.save("rsm_results") # CSV tables + JSON summary + pickle
Every statistic, including the selection itself, is refitted inside each training fold. Nothing crosses a fold boundary.
Read preserved correctly. It means nothing SIGNIFICANT was lost, not
that nothing was lost. On several real datasets the full signal is ahead on
the point estimate and the paired test cannot resolve the difference at five
folds, which is why both means are printed beside the verdict. The verdict
values are preserved, sig.better, sig.worse and undecidable.
If your samples are not independent (repeated measurements of the same specimen, several instruments per physical sample), pass groups so a specimen is never split across folds:
result = NestedCV(k_outer=5, segment=16, random_state=0).run(X, y, groups=specimen_id)
Failing to do this is the most common way to get an optimistic number out of spectral data.
Recipes
Choosing the penalty. Sparsity strength is per-dataset, so it is a library feature rather than something to guess:
from robustsignalmaker import lam_frontier, lam_for_coverage
def make(lam):
return BootstrapMaskSelector(task="binary", segment=16, lam=lam,
n_bootstrap=20, seed=0).fit(X, y)
for row in lam_frontier(make, [0.02, 0.05, 0.1, 0.2]):
print(row) # coverage, stability, ambiguity, support
hit = lam_for_coverage(make, target=0.25)
print(hit["lam"], hit["coverage"], hit["target_reached"])
target_reached is False when the bracket cannot deliver the coverage you
asked for. The best point found is still returned, so a miss is visible
rather than renamed a hit.
Marginalising over lenses. Which bands look important depends on whether you look at raw intensity, a derivative, or a smoothed scale. Treat that choice as a nuisance parameter:
from robustsignalmaker import RepresentationEnsembleSelector, default_ensemble
ens = RepresentationEnsembleSelector(
representations=default_ensemble(3, seed=0),
task="binary", segment=16, n_bootstrap=20, seed=0).fit(X, y)
print(ens.selected_segments_)
print("lens agreement:", round(float(ens.representation_agreement()), 3))
Low agreement is informative in itself: it means the answer you would have reported depends on a preprocessing choice you might not have thought of as a choice.
Point-wise instead of banded. Set segment=1; it is the same code path.
Troubleshooting
InsufficientEvidenceErrormeans RSM refused rather than guessed. Usual causes: too few samples for the effective-sample-size floor (n_min_hard=8), an evidence threshold set above what the data carries (rho_min,o_min), or sparse data whose natural scale differs from the defaults. For sparse spectra such as mass spectrometry, declare the scale:min_valid_frac=0.02, rho_min=0.01.- All
pi_are nan. Nothing was assessable. Check that your validity mask has the polarity right (True = observed) and thato_minis not above your observation rate. MissingnessInformativeWarning. The missingness pattern alone predicts the target, so the observation process is confounded with what you are trying to measure. RSM reports this and does not attempt to fix it. Do not silence it: it usually means an instrument or batch effect is standing in for the label.- A verdict of
undecidable. Fewer than two folds were comparable, usually because predictions were refused for lack of evidence. The refusal count is in the result summary. - Everything looks stable and nothing predicts. Stability alone never ranks methods. A perfectly reproducible wrong answer scores 1.0; always read the score beside it.
The RobustMaker family
| Package | Selects | Data |
|---|---|---|
| RobustModelMaker | columns | tabular features |
| RobustPixelMaker | patches | scientific images |
| RobustSignalMaker | points and bands | signals and spectra |
Documentation and examples
The repository carries four guides (user, API reference, interpretation, implementation), six executed example notebooks (NIR spectroscopy, multi-instrument calibration, SELDI-TOF mass spectrometry, Raman mineral identification, wearable sensors, and gold nanoparticle structure functions), and a benchmark suite whose findings, including the negative ones and the predictions that were falsified, are recorded in FINDINGS.md.
Project home: https://github.com/amaxiom/RobustSignalMaker
Licence
MIT. Copyright (c) 2026 Amanda S Barnard.
Metadata
Release files for robustsignalmaker 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| robustsignalmaker-0.2.0.tar.gz | 66.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| robustsignalmaker-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 136.2 kB
Release files / robustsignalmaker-0.2.0.tar.gz
| Download URL | robustsignalmaker-0.2.0.tar.gz |
|---|---|
| Size | 66.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d8b17fd8bd80ff8bfd08e63b5b0a7684e6b1435625b9e504c1b2021043542462
|
|
BLAKE2b-256 checksum How to use checksums |
47e21caa9d4395aed09a73ec3b848f20c5cc5a587deca96c6ca9f1e25f1871e1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.10
|
Release files / robustsignalmaker-0.2.0-py3-none-any.whl
| Download URL | robustsignalmaker-0.2.0-py3-none-any.whl |
|---|---|
| Size | 69.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6939b0eb29ae993f927a073430f968ead47796bb2626e0b5408c3d6c8a89d49d
|
|
BLAKE2b-256 checksum How to use checksums |
caf43b2a9a71c6dc45663e2129cc5191a5bbadeba4a7bab9835c09ec393b9582
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.10
|