paneldx
Check that your longitudinal dataset is what it claims to be, before you model it.
pip install paneldx
paneldx audit data.csv --time quarter --target risk_score --html report.html
What it does, plainly
Imagine a class attendance sheet where you assume "student #3" is the same child every week. But the sheet is re-sorted by test score each week, so student #3 is a different child every time. Now compute "student #3's progress over the term". You get a number. It means nothing.
Data that tracks the same patients, customers or providers across time has this failure mode constantly, and nothing warns you. Lags, differences, trajectories, sequence models and grouped cross-validation all silently assume the identifier points at one entity. When it does not, nothing errors. The numbers just stop meaning anything.
paneldx tests that assumption against the data itself, and tells you when it
fails. It also catches three related problems that make a model look better than
it is.
When you would use this
- Before training on any panel dataset, especially one assembled from exports, joins, or de-identification passes, where re-sorting can sever an ID from the entity it names.
- When results look too good. An R² of 0.95 on time-series data is more often
a warning than an achievement.
paneldxtells you what a naive baseline scores so you know whether your model beat anything. - When inheriting someone else's data and the entity column arrived without documentation.
- As a CI gate. The command exits non-zero on a disqualifying defect, and also when there is too little evidence to judge, so a pipeline can refuse to train on a broken or unverifiable panel.
- When reviewing a paper or a colleague's analysis and you want to check the panel is sound before reading the results.
Not useful for cross-sectional data with no time dimension, or for panels of fewer than two periods.
The bug this was built from
An earlier physician-analytics pipeline of ours built its entity ID from row position:
physician_id = ((serial_number - 1) % rows_per_quarter) + 1
Each quarter was sorted by platform rank before this ran, so position i was a different doctor every quarter. The ID linked strangers together, and that ID then fed momentum features, a GRU over "trajectories", and the grouped CV splits.
paneldx on that dataset, given no hints:
key: physician_id key: Disease + Opening time
columns explained 2 of 30 (7%) columns explained 13 of 29 (45%)
VERDICT NOT SUPPORTED VERDICT supported by the data
These checks were carried out after the PopNet paper was published in 2025. They were not included in the paper and are not corrected PopNet results.
The search tested all one-column and two-column combinations. Disease + Opening time received the strongest support from the available data. This does
not prove that it is the actual physician identifier, so it is treated as a
candidate key.
That was one of three defects. The other two turned up on the same dataset without being told anything about it:
| Check | What it found |
|---|---|
detect_counters |
Total patients (lag-1 ρ = 0.962), Total visits (0.871), Medical consultation records (0.961). Lifetime totals that barely move, so a target built from them is autocorrelated by construction |
target_leakage |
R² = 0.923 reconstructing the target from its own features, naming inv_rank, log_gifts and log_visits, which were three of the four components the target was averaged from |
persistence_baseline |
Carry-forward MAE 0.0320 against that model's 0.0942. Doing nothing was 2.9x better |
The defects hide each other
A broken key does not only corrupt features. It disarms the check that would have caught it:
| Key used | Persistence MAE | R² | What a researcher concludes |
|---|---|---|---|
| Positional key (as used) | 0.4717 | 0.191 | "carry-forward is useless, my 0.0942 is good" |
Candidate (Disease + Opening time) |
0.0320 | 0.971 | carry-forward beats the model by 2.9x |
Under the positional key the naive forecast looks worthless, so nobody thinks to
compare against it. Under the best-supported candidate key it is close to
unbeatable. This is why paneldx establishes the key before it reports any
baseline, and why running the baseline alone would not have caught it.
Install
pip install paneldx
Python 3.9+, numpy and pandas. Nothing else: no modelling framework, no
template engine.
Usage
from paneldx import audit, to_html
from paneldx import validate_key, discover_keys
from paneldx import detect_counters, target_leakage, persistence_baseline
# Is the entity key supported by the data?
print(validate_key(df, "patient_id", time_col="quarter"))
print(discover_keys(df, time_col="quarter")[0]) # no hints needed
# Are the numbers about to fool you?
print(detect_counters(df, "patient_id", "quarter")) # lifetime totals
print(target_leakage(df, target="risk_score")) # is y inside X?
print(
persistence_baseline(df, "patient_id", "quarter", "risk_score", period_step="QS")
) # declare the cadence
# Or all of it at once
result = audit(df, "quarter", key="patient_id", target="risk_score")
open("report.html", "w").write(to_html(result))
The CLI reads CSV, TSV, Excel, Parquet, Feather and JSON:
paneldx audit panel.xlsx --time t --key site_id patient_id --target outcome
The three traps
Cumulative features. Lifetime totals barely move between periods, so a target built from them is near-perfectly autocorrelated. Models trained on them report excellent R² for restating what they were handed. Difference them into per-period flows.
Target composition. When the target is computed from columns that are also
features, the model is not predicting, it is recovering arithmetic. The metrics
look superb and mean nothing. target_leakage fits a deliberately linear
model: the point is not to predict the target well, but to show no prediction was
ever required.
Missing naive baseline. On autocorrelated panels, carrying the last value forward is often unbeatable. A model reported without it may be losing to a one-line rule, and no reader could tell.
How key validation works
A useful entity key should reveal consistent patterns across periods.
Invariants. Some attributes should not change for the same entity, such as a birth date or account opening date. Under a valid entity key, these values should normally remain stable. Under a broken key, they may change between periods.
Counters. Some values only increase over time. Under a valid entity key, their within-entity changes should normally be non-negative. Under a broken key, they may move in both directions.
paneldx compares this structure with shuffled entity labels. A supported key should perform better than the shuffled reference. A broken key may perform close to it.
The verdict uses the share of columns explained, not only the count. A key created from row order may explain a few columns because those columns were used for sorting. A supported entity key should explain a larger share of the data.
evidence_frac |
Verdict |
|---|---|
| ≥ 0.40 | supported by the data |
| 0.15 to 0.40 | weak, inspect the listed columns by hand |
| < 0.15 | not supported, within-entity quantities are unsafe |
Limitations
- Near-valid keys are hard to separate from perfect ones. The tolerances exist because valid keys may also contain a small number of collisions of entities can score like a clean one. Ties break toward the finer partition, which mitigates but does not eliminate this.
- Leakage detection is linear. A target computed from its features by some
non-linear rule can slip past
target_leakage. A high score is strong evidence; a low one is not a clearance. - Two-column search is O(n²) in columns. On a 24k x 31 table the full pair
search takes about 23 seconds. Pass
--keywhen you already know it. - A passing verdict is not a proof. It says the data is consistent with the key, not that the key is correct. Domain knowledge still wins.
- Needs at least two periods, and enough entities to measure a rate against.
Roadmap
- Panel key discovery and validation
- Cumulative-counter detection
- Target-composition leakage
- Naive-baseline harness
- One-command HTML report
- CLI
- Look-ahead detection in sequence construction
- Faster key search on wide tables
Documentation
Contributing
See CONTRIBUTING.md. False positives are bugs: if paneldx
rejects a key you know is correct, that is worth an issue.
Release notes are in CHANGELOG.md.
Citation
If you use paneldx in your research, please see CITATION.cff.
The DOI for version 0.4.0 will be added after the release is archived on Zenodo. The old DOI belongs to version 0.3.1 and does not refer to the current code.
License
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file paneldx-0.4.0.tar.gz.
File metadata
- Download URL: paneldx-0.4.0.tar.gz
- Upload date:
- Size: 55.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
32c847d994d7fea805be140bace95b71cfc435b16dcd31760640a9800cbcbb75
|
|
| MD5 |
a746a928d8c2222bc8c4801b40052648
|
|
| BLAKE2b-256 |
1802d5a932da8250439e0028072f5de5a4c1476a8c27108ab1472dec7e6797cc
|
Provenance
The following attestation bundles were made for paneldx-0.4.0.tar.gz:
Publisher:
release.yml on stemmatics/paneldx
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
paneldx-0.4.0.tar.gz -
Subject digest:
32c847d994d7fea805be140bace95b71cfc435b16dcd31760640a9800cbcbb75 - Sigstore transparency entry: 2583230223
- Sigstore integration time:
-
Permalink:
stemmatics/paneldx@b0ac25acbc7977ca25d963b5289fa990bcf7747c -
Branch / Tag:
refs/tags/v0.4.0 - Owner: https://github.com/stemmatics
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@b0ac25acbc7977ca25d963b5289fa990bcf7747c -
Trigger Event:
push
-
Statement type:
File details
Details for the file paneldx-0.4.0-py3-none-any.whl.
File metadata
- Download URL: paneldx-0.4.0-py3-none-any.whl
- Upload date:
- Size: 30.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
72facda8358f3de80881a1a2b60a0294bccd72ed8ce99e8997a8e9747cb44245
|
|
| MD5 |
7377342f3f9d65c7f6ff3f646f906bce
|
|
| BLAKE2b-256 |
eb698451f8b5f52c33d5c68a4282057f4a4f48cae84c820304bb2c8d86436535
|
Provenance
The following attestation bundles were made for paneldx-0.4.0-py3-none-any.whl:
Publisher:
release.yml on stemmatics/paneldx
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
paneldx-0.4.0-py3-none-any.whl -
Subject digest:
72facda8358f3de80881a1a2b60a0294bccd72ed8ce99e8997a8e9747cb44245 - Sigstore transparency entry: 2583230244
- Sigstore integration time:
-
Permalink:
stemmatics/paneldx@b0ac25acbc7977ca25d963b5289fa990bcf7747c -
Branch / Tag:
refs/tags/v0.4.0 - Owner: https://github.com/stemmatics
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@b0ac25acbc7977ca25d963b5289fa990bcf7747c -
Trigger Event:
push
-
Statement type: