SleepKit PSG
Prepare heterogeneous polysomnography data with an explicit, reproducible contract.
简体中文 · Migration · Dataset profiles · Output contract
SleepKit PSG converts PSG recordings and sleep-stage annotations into consistent NumPy arrays for research and model development. Dataset assumptions live in validated YAML profiles; discovery, pairing, channel derivation, filtering, resampling, alignment, quality control, and provenance use the same Python API and CLI.
This is the maintained PyPI distribution of the PSGPrep engineering rewrite. Version 2.1 unifies the historical sleep-kit-psg package and PSGPrep 2.0 implementation under the canonical sleep_kit Python namespace.
Core capabilities
- 22 packaged dataset profiles with explicit validation status.
- Deterministic regular-expression pairing and subject grouping; unmatched files remain visible.
- Auditable channel matching, reference subtraction, units, canonical sleep stages, and QC.
- Configurable filtering, resampling, epoch alignment, normalization, and sequence generation.
- Atomic NPZ/NPY writes, resumable records, per-record failure isolation, and validation.
- Processing-contract hashes that prevent incompatible runs from sharing an output directory.
- Optional HMAC pseudonymization of record and subject IDs.
- Typed Python API and matching CLI with stable exit codes and structured JSON reports.
Installation
Python 3.10 or newer is required. The base install includes NumPy, SciPy, MNE, and PyYAML so EDF workflows work immediately:
python -m pip install sleep-kit-psg
Install HDF5 support for DOD and PhysioNet 2018 profiles when needed:
python -m pip install 'sleep-kit-psg[hdf5]'
For development from this repository:
git clone https://github.com/lijinyang439-arch/PSGPrep.git
cd PSGPrep
python -m pip install -e '.[dev]'
Minimal runnable example
The repository includes a deterministic synthetic PSG record. This exercises the same public API, filtering, epoching, output, and validation path without downloading controlled data:
python examples/synthetic_quickstart.py --output ./demo-output
The command exits 0 only when preprocessing and artifact validation both pass. Remove or choose another output directory before rerunning because unrelated or differently configured output is never overwritten silently.
Typical CLI workflow
First inspect the packaged contract and scan pairing without reading signal samples:
sleepkit-psg profiles
sleepkit-psg show-profile shhs1
sleepkit-psg scan \
--profile shhs1 \
--input-root /path/to/shhs1/edfs \
--annotation-root /path/to/shhs1/annotations \
--details
Then process selected canonical channels and validate the result:
sleepkit-psg preprocess \
--profile shhs1 \
--input-root /path/to/shhs1/edfs \
--annotation-root /path/to/shhs1/annotations \
--output-root ./outputs/shhs1 \
--channels C4 E1 \
--target-sfreq 100 \
--workers 4
sleepkit-psg validate --output-root ./outputs/shhs1
Use --require-complete-pairing when any unmatched source must fail the run. Use --fail-fast only with one worker. --log-file appends a log outside the output directory; progress always goes to standard error and the JSON result goes to standard output.
Exit codes are consistent across commands:
| Code | Meaning |
|---|---|
0 |
Command and all selected records/artifacts succeeded |
1 |
Scan was incomplete when required, or a record/artifact failed |
2 |
Invalid arguments, profile, paths, configuration, or I/O setup |
Python API
The high-level API accepts both str and pathlib.Path:
from pathlib import Path
from sleep_kit import preprocess, scan, validate_output
pairing = scan(
"shhs1",
input_root=Path("/path/to/shhs1/edfs"),
annotation_root=Path("/path/to/shhs1/annotations"),
)
print(f"paired={len(pairing.records)} unmatched={len(pairing.unmatched_signals)}")
summary = preprocess(
"shhs1",
input_root=Path("/path/to/shhs1/edfs"),
annotation_root=Path("/path/to/shhs1/annotations"),
output_root=Path("outputs/shhs1"),
channels=("C4", "E1"),
target_sfreq=100,
workers=4,
)
validation = validate_output("outputs/shhs1")
if not summary.succeeded or not validation.valid:
raise RuntimeError((summary.as_dict(), validation.as_dict()))
The main public surface is:
list_profiles() -> tuple[str, ...]load_profile(name_or_path) -> DatasetProfilescan(profile, input_root, annotation_root=None) -> DiscoveryReportpreprocess(profile, input_root, output_root, **options) -> RunSummaryvalidate_output(output_root) -> ValidationReport
run_pipeline(..., options=RunOptions(...)) remains available for advanced and PSGPrep 2.0-compatible usage.
Inputs and supported scope
Signal readers currently cover EDF/BDF/EDF-compatible REC through MNE, explicit NPZ contracts, MATLAB arrays used by PhysioNet 2018, and DOD HDF5. Annotation readers cover NSRR XML, EDF annotations, epoch text/tables, event tables, HDF5 hypnograms, and PhysioNet 2018 files.
Packaged profiles include ABC, CCSHS, CFS, DCSM, DOD, HMC, HomePAP, ISRUC, MASS (mass13), MESA, MNC, MrOS visits 1/2, NCHSDB, PhysioNet 2018, SHHS visits 1/2, Sleep-EDF cassette/telemetry, SOF, STAGES, and WSC. A profile marked migrated-unverified is a starting contract, not a claim that every dataset release or record has been validated. See dataset profiles and the bounded validation report.
Canonical labels are W=0, N1=1, N2=2, N3/N4=3, REM=4, and UNKNOWN=5. The default epoch is 30 seconds, target sampling rate is profile-defined (normally 100 Hz), and normalization defaults to none. Scientific overrides are recorded in provenance.
Output and overwrite behavior
The recommended NPZ layout is:
output-root/
├── records/<record_id>.npz
├── sequences/<record_id>.npz
├── completion/<record_id>.json
├── manifest.jsonl
├── errors.jsonl
├── provenance.json
└── summary.json
Record x is float32 [epoch, channel, sample]; y is int8 [epoch]. Sequence arrays are float32 [sequence, epoch, channel, sample] and int8 [sequence, epoch]. See the complete output contract.
- A completed record is resumed by default and reported as
skipped_existing. overwrite=Trueor--overwritereprocesses only records under the same contract.- A changed profile, version, channel order, sampling rate, normalization, sequence length, format, or ID salt requires a new output directory.
- A non-empty directory without SleepKit provenance is rejected and left untouched.
- Every record failure is written to
errors.jsonl; no exception is silently discarded.
Privacy and reproducibility
Source paths in manifests are relative. For private datasets, enable hash_ids with a high-entropy secret salt:
sleepkit-psg preprocess \
--profile /path/to/private-profile.yaml \
--input-root /path/to/private-signals \
--output-root ./outputs/private \
--hash-ids \
--id-salt-file /secure/path/id-salt.bin
The salt is never copied; only its SHA-256 fingerprint prevents resuming with a different key. Pseudonymization is not anonymization. Split model data by subject_id, not by epoch, sequence, or visit.
FAQ and limitations
Why were zero records processed? scan --details reports whether globs found files, regular expressions extracted IDs, and annotations paired. Point roots at the levels expected by the profile instead of relying on fuzzy filename matching.
Why did a record fail? Inspect errors.jsonl, then its profile's channel candidates, raw label map, sampling rate, and alignment policy. Failures are intentionally explicit.
Can I add a dataset? Copy examples/custom_dataset.yaml, use relative globs and named record_id groups, then add synthetic parser/pairing tests before claiming real-data validation.
Does SleepKit score sleep automatically? No. It prepares recordings and existing sleep-stage annotations for downstream analysis or models.
Are all 22 profiles fully validated? No. Validation status is per profile and bounded. Dataset releases, headers, channel semantics, scoring standards, and unusual records still require study-specific review.
Migration from 1.x and PSGPrep 2.0
Version 2.1 is a breaking product upgrade from sleep-kit-psg 1.x. The canonical import is still sleep_kit, but output defaults, pairing, exceptions, and configuration are stricter. fast_preprocess() and sleepkit-process remain as deprecated compatibility wrappers. The psgprep import and command from PSGPrep 2.0 also remain temporarily available. See MIGRATION.md for exact mappings and behavioral changes.
Development and contribution
python -m pip install -e '.[dev]'
ruff format --no-cache --check .
ruff check --no-cache .
mypy --no-incremental --cache-dir=/tmp/sleepkit-mypy-cache
python -m unittest discover -s tests -v
python scripts/verify_examples.py
python scripts/release_check.py
python -m build
twine check dist/*
check-wheel-contents dist/*.whl
Contributions must not contain recordings, annotations, clinical tables, credentials, private IDs, or local absolute paths. Scientific changes need a profile update, regression test, and clear validation boundary. See CONTRIBUTING.md, SECURITY.md, and CHANGELOG.md.
Citation and license
Use CITATION.cff, cite each source dataset separately, and report the package version plus profile digest stored in provenance.json.
SleepKit PSG is licensed under Apache-2.0. The license covers this software, not any input or generated dataset.
Release files for sleep-kit-psg 2.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sleep_kit_psg-2.1.0.tar.gz | 195.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sleep_kit_psg-2.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 260.4 kB
Release files / sleep_kit_psg-2.1.0.tar.gz
| Download URL | sleep_kit_psg-2.1.0.tar.gz |
|---|---|
| Size | 195.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
68dc09cf4905785ef6914c135bee300febbce02e91c864fbf0c00bd2c14623b5
|
|
BLAKE2b-256 checksum How to use checksums |
de8cedb6488f93e503a083457e4b84e1507358ed8b9c6fe29281516b2d8780fe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.9
|
Release files / sleep_kit_psg-2.1.0-py3-none-any.whl
| Download URL | sleep_kit_psg-2.1.0-py3-none-any.whl |
|---|---|
| Size | 64.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
02300bbc1fd71f433c2f7967663e9637e1a56c77f250b9a0bc5cbc2ba385eab9
|
|
BLAKE2b-256 checksum How to use checksums |
b16a69d5e9482fa3dede2c87e3343deade1d9062d7351ad5af61e57e4445810d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.9
|