Skip to main content

SleepKit PSG: auditable PSG preprocessing for sleep-staging research

SleepKit PSG

Prepare heterogeneous polysomnography data with an explicit, reproducible contract.

PyPI CI Python 3.10+ Apache-2.0

简体中文 · Migration · Dataset profiles · Output contract

SleepKit PSG converts PSG recordings and sleep-stage annotations into consistent NumPy arrays for research and model development. Dataset assumptions live in validated YAML profiles; discovery, pairing, channel derivation, filtering, resampling, alignment, quality control, and provenance use the same Python API and CLI.

This is the maintained PyPI distribution of the PSGPrep engineering rewrite. Version 2.1 unifies the historical sleep-kit-psg package and PSGPrep 2.0 implementation under the canonical sleep_kit Python namespace.

Core capabilities

  • 22 packaged dataset profiles with explicit validation status.
  • Deterministic regular-expression pairing and subject grouping; unmatched files remain visible.
  • Auditable channel matching, reference subtraction, units, canonical sleep stages, and QC.
  • Configurable filtering, resampling, epoch alignment, normalization, and sequence generation.
  • Atomic NPZ/NPY writes, resumable records, per-record failure isolation, and validation.
  • Processing-contract hashes that prevent incompatible runs from sharing an output directory.
  • Optional HMAC pseudonymization of record and subject IDs.
  • Typed Python API and matching CLI with stable exit codes and structured JSON reports.

Installation

Python 3.10 or newer is required. The base install includes NumPy, SciPy, MNE, and PyYAML so EDF workflows work immediately:

python -m pip install sleep-kit-psg

Install HDF5 support for DOD and PhysioNet 2018 profiles when needed:

python -m pip install 'sleep-kit-psg[hdf5]'

For development from this repository:

git clone https://github.com/lijinyang439-arch/PSGPrep.git
cd PSGPrep
python -m pip install -e '.[dev]'

Minimal runnable example

The repository includes a deterministic synthetic PSG record. This exercises the same public API, filtering, epoching, output, and validation path without downloading controlled data:

python examples/synthetic_quickstart.py --output ./demo-output

The command exits 0 only when preprocessing and artifact validation both pass. Remove or choose another output directory before rerunning because unrelated or differently configured output is never overwritten silently.

Typical CLI workflow

First inspect the packaged contract and scan pairing without reading signal samples:

sleepkit-psg profiles
sleepkit-psg show-profile shhs1
sleepkit-psg scan \
  --profile shhs1 \
  --input-root /path/to/shhs1/edfs \
  --annotation-root /path/to/shhs1/annotations \
  --details

Then process selected canonical channels and validate the result:

sleepkit-psg preprocess \
  --profile shhs1 \
  --input-root /path/to/shhs1/edfs \
  --annotation-root /path/to/shhs1/annotations \
  --output-root ./outputs/shhs1 \
  --channels C4 E1 \
  --target-sfreq 100 \
  --workers 4

sleepkit-psg validate --output-root ./outputs/shhs1

Use --require-complete-pairing when any unmatched source must fail the run. Use --fail-fast only with one worker. --log-file appends a log outside the output directory; progress always goes to standard error and the JSON result goes to standard output.

Exit codes are consistent across commands:

Code Meaning
0 Command and all selected records/artifacts succeeded
1 Scan was incomplete when required, or a record/artifact failed
2 Invalid arguments, profile, paths, configuration, or I/O setup

Python API

The high-level API accepts both str and pathlib.Path:

from pathlib import Path

from sleep_kit import preprocess, scan, validate_output

pairing = scan(
    "shhs1",
    input_root=Path("/path/to/shhs1/edfs"),
    annotation_root=Path("/path/to/shhs1/annotations"),
)
print(f"paired={len(pairing.records)} unmatched={len(pairing.unmatched_signals)}")

summary = preprocess(
    "shhs1",
    input_root=Path("/path/to/shhs1/edfs"),
    annotation_root=Path("/path/to/shhs1/annotations"),
    output_root=Path("outputs/shhs1"),
    channels=("C4", "E1"),
    target_sfreq=100,
    workers=4,
)

validation = validate_output("outputs/shhs1")
if not summary.succeeded or not validation.valid:
    raise RuntimeError((summary.as_dict(), validation.as_dict()))

The main public surface is:

  • list_profiles() -> tuple[str, ...]
  • load_profile(name_or_path) -> DatasetProfile
  • scan(profile, input_root, annotation_root=None) -> DiscoveryReport
  • preprocess(profile, input_root, output_root, **options) -> RunSummary
  • validate_output(output_root) -> ValidationReport

run_pipeline(..., options=RunOptions(...)) remains available for advanced and PSGPrep 2.0-compatible usage.

Inputs and supported scope

Signal readers currently cover EDF/BDF/EDF-compatible REC through MNE, explicit NPZ contracts, MATLAB arrays used by PhysioNet 2018, and DOD HDF5. Annotation readers cover NSRR XML, EDF annotations, epoch text/tables, event tables, HDF5 hypnograms, and PhysioNet 2018 files.

Packaged profiles include ABC, CCSHS, CFS, DCSM, DOD, HMC, HomePAP, ISRUC, MASS (mass13), MESA, MNC, MrOS visits 1/2, NCHSDB, PhysioNet 2018, SHHS visits 1/2, Sleep-EDF cassette/telemetry, SOF, STAGES, and WSC. A profile marked migrated-unverified is a starting contract, not a claim that every dataset release or record has been validated. See dataset profiles and the bounded validation report.

Canonical labels are W=0, N1=1, N2=2, N3/N4=3, REM=4, and UNKNOWN=5. The default epoch is 30 seconds, target sampling rate is profile-defined (normally 100 Hz), and normalization defaults to none. Scientific overrides are recorded in provenance.

Output and overwrite behavior

The recommended NPZ layout is:

output-root/
├── records/<record_id>.npz
├── sequences/<record_id>.npz
├── completion/<record_id>.json
├── manifest.jsonl
├── errors.jsonl
├── provenance.json
└── summary.json

Record x is float32 [epoch, channel, sample]; y is int8 [epoch]. Sequence arrays are float32 [sequence, epoch, channel, sample] and int8 [sequence, epoch]. See the complete output contract.

  • A completed record is resumed by default and reported as skipped_existing.
  • overwrite=True or --overwrite reprocesses only records under the same contract.
  • A changed profile, version, channel order, sampling rate, normalization, sequence length, format, or ID salt requires a new output directory.
  • A non-empty directory without SleepKit provenance is rejected and left untouched.
  • Every record failure is written to errors.jsonl; no exception is silently discarded.

Privacy and reproducibility

Source paths in manifests are relative. For private datasets, enable hash_ids with a high-entropy secret salt:

sleepkit-psg preprocess \
  --profile /path/to/private-profile.yaml \
  --input-root /path/to/private-signals \
  --output-root ./outputs/private \
  --hash-ids \
  --id-salt-file /secure/path/id-salt.bin

The salt is never copied; only its SHA-256 fingerprint prevents resuming with a different key. Pseudonymization is not anonymization. Split model data by subject_id, not by epoch, sequence, or visit.

FAQ and limitations

Why were zero records processed? scan --details reports whether globs found files, regular expressions extracted IDs, and annotations paired. Point roots at the levels expected by the profile instead of relying on fuzzy filename matching.

Why did a record fail? Inspect errors.jsonl, then its profile's channel candidates, raw label map, sampling rate, and alignment policy. Failures are intentionally explicit.

Can I add a dataset? Copy examples/custom_dataset.yaml, use relative globs and named record_id groups, then add synthetic parser/pairing tests before claiming real-data validation.

Does SleepKit score sleep automatically? No. It prepares recordings and existing sleep-stage annotations for downstream analysis or models.

Are all 22 profiles fully validated? No. Validation status is per profile and bounded. Dataset releases, headers, channel semantics, scoring standards, and unusual records still require study-specific review.

Migration from 1.x and PSGPrep 2.0

Version 2.1 is a breaking product upgrade from sleep-kit-psg 1.x. The canonical import is still sleep_kit, but output defaults, pairing, exceptions, and configuration are stricter. fast_preprocess() and sleepkit-process remain as deprecated compatibility wrappers. The psgprep import and command from PSGPrep 2.0 also remain temporarily available. See MIGRATION.md for exact mappings and behavioral changes.

Development and contribution

python -m pip install -e '.[dev]'
ruff format --no-cache --check .
ruff check --no-cache .
mypy --no-incremental --cache-dir=/tmp/sleepkit-mypy-cache
python -m unittest discover -s tests -v
python scripts/verify_examples.py
python scripts/release_check.py
python -m build
twine check dist/*
check-wheel-contents dist/*.whl

Contributions must not contain recordings, annotations, clinical tables, credentials, private IDs, or local absolute paths. Scientific changes need a profile update, regression test, and clear validation boundary. See CONTRIBUTING.md, SECURITY.md, and CHANGELOG.md.

Citation and license

Use CITATION.cff, cite each source dataset separately, and report the package version plus profile digest stored in provenance.json.

SleepKit PSG is licensed under Apache-2.0. The license covers this software, not any input or generated dataset.

Release files for sleep-kit-psg 2.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sleep-kit-psg 2.1.0
File Size Uploaded
sleep_kit_psg-2.1.0.tar.gz 195.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sleep-kit-psg 2.1.0
File Interpreter ABI Platform
sleep_kit_psg-2.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 260.4 kB

Release files / sleep_kit_psg-2.1.0.tar.gz

Download URL sleep_kit_psg-2.1.0.tar.gz
Size 195.7 kB
Tags Source
SHA-256 checksum
How to use checksums
68dc09cf4905785ef6914c135bee300febbce02e91c864fbf0c00bd2c14623b5
BLAKE2b-256 checksum
How to use checksums
de8cedb6488f93e503a083457e4b84e1507358ed8b9c6fe29281516b2d8780fe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.9

Release files / sleep_kit_psg-2.1.0-py3-none-any.whl

Download URL sleep_kit_psg-2.1.0-py3-none-any.whl
Size 64.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
02300bbc1fd71f433c2f7967663e9637e1a56c77f250b9a0bc5cbc2ba385eab9
BLAKE2b-256 checksum
How to use checksums
b16a69d5e9482fa3dede2c87e3343deade1d9062d7351ad5af61e57e4445810d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.9

Release history Release notifications | RSS feed

2.1.1

2 release files

This release

2.1.0 This release

2 release files

1.2.2

2 release files

1.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page