confound-controls
Did your classifier learn the signal, or a confound?
For each confound, restrict the negatives to a subset matched 1:1 to the positives on that confound, then recompute AUROC. If the signal survives matching, it is not that confound. If it collapses to chance, it was.
from confound_controls import run_battery, format_battery, bootstrap_auroc
anchor = bootstrap_auroc(df["label"], probs).point # uncontrolled baseline
results = run_battery(
df,
prob_map,
spec={"gc": ["gc"], "expression": ["log_expr"], "joint": ["gc", "log_expr"]},
anchor=anchor,
)
print(format_battery(results, anchor))
Nothing is read from disk and nothing is written to it. The frame, the probabilities, the confound spec and the anchor are all arguments.
The failure this exists to prevent
1:1 matching without replacement from a negative pool the same size as the positive set consumes every negative, whatever the confound values are. The matched comparison is then identical to the unmatched one — and a vacuous control does not announce itself as inconclusive. It returns the original result, which reads as the confound having been ruled out.
This is not hypothetical. It is how the first version of this library's own
test suite behaved: 150 positives against 150 negatives, a score computed from
nothing but GC, matched on GC, verdict confound-robust, AUROC 0.9312.
So evaluate_confound refuses by default:
ValueError: matching consumed the entire negative pool (150 matched from 150
available), so it selected nothing and the control is vacuous -- the matched
comparison equals the unmatched one and will report the confound as ruled out.
Pass require_selective_match=False if you mean it; the result then carries
match_selective=False and format_battery prints [VACUOUS CONTROL].
Verdicts
| verdict | meaning |
|---|---|
confound-robust |
enough of the anchor's above-chance signal survived, and the interval clears chance |
confound-driven |
the interval straddles chance — matching removed the signal |
partial |
signal is real but diminished |
inconclusive |
the interval is too wide to support any of the above (opt in via max_ci_width) |
inverted |
the interval clears chance from BELOW — the control reversed the ranking |
inconclusive has no counterpart in the source this came from, which folded
those cases into confound-driven — reporting an absent measurement as a
finding.
inverted likewise had no counterpart: an interval entirely below chance
excludes chance, so it fell through to partial — reporting a control that
reliably reversed the ranking as a diminished-but-real signal.
Two more controls
Incremental validity — does the new feature add anything beyond the confounds? Fit confound-only and confound+feature models, evaluate both on the same held-out rows, and put a paired bootstrap interval on the AUROC difference.
Paired is the point. Two independent intervals on two AUROCs overlap far more often than the interval on their difference excludes zero, because the pair is computed on the same resampled rows and the shared variance cancels. Comparing two separately-reported AUROCs by eye is exactly the error this prevents.
from confound_controls import incremental_validity
res = incremental_validity(conf_train, conf_test, feat_train, feat_test, y_train, y_test)
res.verdict # adds-signal | harms | no-added-signal | underpowered
feat_train must be out-of-fold. Fitting the feature on the rows the confound
model is evaluated against leaks the label into the augmented model and
manufactures the very lift the control is testing for.
The source's verdict was two-valued ("FM-adds-signal" if lo > 0 else "no-added-signal"), so an interval straddling zero, an interval too wide to
say anything, and an interval lying entirely below zero all reported the
same. The last of those is an augmented model that is reliably worse — a
finding, not an absence of one.
Ablation — destroy the structure of interest, keep the nuisance properties, score with the same model, and see what survives.
from confound_controls import assert_ablation_changed_input, ablation_control
assert_ablation_changed_input(real_inputs, ablated_inputs) # do this first
res = ablation_control(y, p_real, p_ablated)
res.verdict # structure-dependent | structure-independent | partial | inconclusive
That first call is not optional politeness. A broken ablation leaves the
input unchanged, so the scores match exactly, the delta is zero, and the
control reports the model as surviving — the strongest possible result,
obtained by doing nothing. No score-level statistic can catch it, which is why
the check belongs on the input. The source had this instinct
(assert not np.allclose(shuffled, real)); here it is a first-class function
that also handles non-float inputs like sequences, where np.allclose raises.
This is the same shape as the vacuous matched control above: an inert control returns the original answer, which reads as the hypothesis surviving.
What changed from the source
Extracted from an internal matched-negatives script, present byte-identically in a second internal project. The code had been copied between two independent repositories, which is what marked it as worth extracting: duplication across repos is revealed reuse demand rather than a proxy for it.
The matching and the statistics are unchanged. What changed:
ANCHOR = 0.629was a module-level constant — one study's own baseline AUROC compiled into the library, against which every imported use silently scored. Now a required argument, because there is no defensible default for "what does good look like in your problem".- Input paths were derived from
__file__and the confound columns (gc,log_basemean,og_size,n_frac, 64kmer3_*) were hardcoded, so the module ran against one dataset in one repo layout. Both are arguments. - Incomplete matches were silent. When the pool ran out the original
returned fewer negatives than positives and said nothing — and the drop is
not random, it hits the positives in the densest region of confound space.
MatchResult.completenow answers this, and it is enforced by default. - The pass flag counted a hardcoded subset of the study's own confound
names, so a confound added to the spec was scored, printed, and then left out
of the decision it existed to inform.
battery_passescounts every result. recovery()divided byanchor - 0.5without checking it. An anchor at or below chance makes that ratio a non-quantity rather than a large number; it is now refused.- The bootstrap silently skipped single-class resamples. With few positives most resamples are skipped and the interval comes from a handful of replicates. It now raises rather than returning a confident-looking number.
Sequence ablation
For sequence models, the ablation is a knockout: scramble a span, keep its dinucleotide composition, re-score, and ask whether the drop is real and larger than a length-matched control knockout elsewhere.
from confound_controls import (
knockout_span,
sample_control_span,
grouped_delta_ci,
confirm_knockout,
)
ko = knockout_span(seq, start, end, seed=1).require_changed()
ctl = sample_control_span(len(seq), end - start, motif_spans, rng).require_disjoint()
pooled = grouped_delta_ci(real_scores, ko_scores, family_ids)
control = grouped_delta_ci(real_scores, ctl_scores, family_ids)
confirm_knockout(pooled, control) # confirmed | not-confirmed
grouped_delta_ci resamples whole groups, not rows. Promoters from one
paralog family are not independent observations; a row-level bootstrap treats
96 correlated rows as 96 independent ones and reports an interval far narrower
than the data supports. Measured on the bundled fixture: grouped 0.189 wide vs
row-level 0.071.
Three silent fallbacks, all pointing the same way. Each returned a plausible value on failure, and each failure makes the ablation weaker or absent — which reads as the model surviving it:
- the dinucleotide shuffle ended
return seqon failure, so an un-shuffleable span came back unchanged — a knockout that knocked nothing out. The comment called this "rare"; measured on a real 29-mer, knocking out[8:20]leaves it unchanged for 2 of the first 8 seeds. - the control-span sampler tried 50 times to avoid the real motif spans, then returned an overlapping start regardless — a partial real knockout posing as the negative control it is compared against.
- the cluster bootstrap never checked it had enough distinct groups to resample.
All three now report. ShuffleResult.changed, ControlSpan.disjoint, and a
min_groups floor, each with a require_* that raises.
The Altschul-Erikson shuffle itself is ported unchanged — including its
determinism fix. Using set(graph) for the vertex list made rng.choice
consume the PRNG in a different order per process (set iteration depends on the
randomized hash of string keys), so the same (seq, seed) produced different
shuffles run to run. A same-process test cannot catch that; the suite runs four
subprocesses and compares.
Not extracted
aim3_confound_table.py builds the confound table itself from FASTA, DE
results and orthogroup sizes. That is study data-prep rather than a control,
and its sequence-composition code would duplicate what seqbench already
does — so it stays where it is.
The source repos are untouched
Both source projects are active. This is extract-and-leave-source-alone: nothing was refactored there, and neither repo imports this package.
Install
pip install confound-controls
From a checkout, for development:
pip install -e ".[dev]"
pytest -q
Requires numpy, pandas, scipy, scikit-learn.
Metadata
Release files for confound-controls 0.3.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| confound_controls-0.3.2.tar.gz | 45.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| confound_controls-0.3.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 82.0 kB
Release files / confound_controls-0.3.2.tar.gz
| Download URL | confound_controls-0.3.2.tar.gz |
|---|---|
| Size | 45.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6a1349c02674939d15cb4dcca0ca1ad5c0b6d7c47e055640b717eb9a3d5dde11
|
|
BLAKE2b-256 checksum How to use checksums |
62e5fef918c4283e18b17dbbfe32af4d28c183282429c905c736bf89d40b5cb5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.3 {"installer":{"name":"uv","version":"0.11.3","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / confound_controls-0.3.2-py3-none-any.whl
| Download URL | confound_controls-0.3.2-py3-none-any.whl |
|---|---|
| Size | 36.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
95dbf10eb48431a9fccd5505f20020af8d812a8174661fc89eb02a8b8a271261
|
|
BLAKE2b-256 checksum How to use checksums |
2c42e3475ea51870ea655a64f32e0a3e7f8ce6ba8fbceb0a3ad4884976d7be5b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.11.3 {"installer":{"name":"uv","version":"0.11.3","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|