Modular hub for lab data acquisition, control, and reproducible persistence
Project description
pyLabHub
pyLabHub is a modular framework for laboratory data acquisition, hardware control, and experiment management.
Its design revolves around three key components: a central hub for data streaming and persistence, adapters that bridge diverse hardware into a unified interface, and connectors that integrate with external programs such as Igor Pro, Python GUIs, or custom clients.
The core principle of pyLabHub is data integrity and reproducibility. It isolates raw experiment data from downstream analysis, ensuring the original record remains uncompromised. At the same time, it provides flexible tools for visualization, real‑time interaction, and automation through scripts that respond dynamically to streaming data. This separation of concerns allows experiments to be both reproducible and openly shareable.
🎯 Target Use Cases
- Lab automation: coordinate instruments and sensors in real‑time experiments
- Real‑time sensing & control: stream measurements while dynamically adjusting hardware
- Data archiving: capture raw experiment data in open, long‑term formats
- Collaborative research: share data and analyses without compromising original records
- Custom interfaces: integrate with Igor Pro, Jupyter notebooks, or new GUIs tailored to specific experiments
✨ Features
- Central Data Hub: low-latency pub/sub bus with persistence
- Adapters: unified drivers for sensors, actuators, DAQs, and instruments
- Connectors: bridges to existing software (Igor Pro, Python, Jupyter, GUIs)
- Persistence: save experiments in scientific formats (HDF5, Parquet, Zarr, CSV)
- Extensible: add new hardware or software integrations with minimal effort
- Data Integrity: isolate raw data from analysis to ensure experiments remain uncompromised
- Reproducibility: keep data open and shareable while preserving the original experimental record
📦 Data Persistence Strategy
pyLabHub will use a hybrid approach to balance efficiency, openness, and reproducibility:
- Zarr for multidimensional array streams (e.g., waveforms, images). Data is chunked by time and channel, compressed (e.g., Blosc + Zstd), and stored in a directory layout that works locally and in cloud object stores.
- Parquet for event logs, metadata, and control records. Appended in time-sliced files that are easy to query with pandas/Arrow/DuckDB.
- HDF5 export provided for compatibility with existing scientific tools, while keeping Zarr/Parquet as the native formats.
This layout ensures raw data is stored separately from analysis, remains uncompromised, and is shareable in open, well‑supported formats.
Example directory layout
runs/
run-2025-09-18-001/
arrays/
raw_adc.zarr/
camera_frames.zarr/
events/
control-20250918-120000.parquet
telemetry-20250918-120000.parquet
manifest.yaml # metadata, checksums, provenance
Run manifest (manifest.yaml) — Draft schema
schema_version: 0.1
run:
id: run-2025-09-18-001
created_at: "2025-09-18T12:00:00Z"
investigator: "Alex Researcher"
organization: "Example University"
award_number: null # e.g., NSF-XXXXXX
description: "NI-DAQ step response test"
identifiers:
dataset_pid: null # DOI/ARK assigned on publish
orcids: [] # contributor ORCIDs
ror: null # institution ROR
acquisition:
timebase:
source_clock_hz: 1e9
hub_clock_offset_s: 0.000123
instruments:
- adapter_id: "hm-ni6321-01"
driver: "ni6321"
version: "0.1.0"
calibration: "cal/ni6321-2025-09-01.yaml"
streams:
- name: raw_adc
kind: array
format: zarr
path: arrays/raw_adc.zarr
dtype: int16
dims: [time, channel]
sample_rate_hz: 200000
channels: 4
units: V
chunking: {time: 2000000, channel: 4}
compression: blosc-zstd
t_start: "2025-09-18T12:00:00.000000Z"
t_end: "2025-09-18T12:10:00.000000Z"
sha256: "<tree-hash-of-zarr>"
- name: control
kind: events
format: parquet
path: events/control-20250918-120000.parquet
rows: 15234
sha256: "<file-hash>"
provenance:
software:
pylabhub: "0.1.0"
python: "3.13.0"
git_commit: "abcdef1"
container_digest: "sha256:..."
conda_lock: "sha256:..."
recipe:
name: "laser_scan_v1"
parameters:
rate_hz: 200000
gain_db: 20
integrity:
manifest_sha256: "<self-hash>"
size_bytes_total: 123456789
checks:
last_verified_at: "2025-09-18T12:15:00Z"
ok: true
sharing:
access: open # open|embargoed|controlled
embargo_until: null
license: CC-BY-4.0
notes: null
Field notes
schema_versionlets the manifest evolve without breaking readers.streamsincludes both array and event stores with paths, shapes, rates, and fixity (sha256).identifiers&sharinghold PIDs/access when published; they can benullduring acquisition.- Timestamps (
t_start,t_end) andtimebasemake time alignment explicit across devices. provenance.softwareandrecipecapture the environment and parameters for reproducibility.
🛠 Roadmap
- MVP Hub: establish the central data hub with pub/sub and basic persistence
- First Adapter: implement a mock hardware adapter and one real device driver (e.g., DAQ)
- First Connector: create a connector for Igor Pro or a simple Python GUI
- Persistence v1: add support for saving runs in open formats (Parquet/Zarr)
- Safety & Control: introduce scheduling, interlocks, and safe‑stop routines
- Cluster Ready: enable multi‑host operation with secure connections
🔏 Additional Features to Integrate/Consider (NSF/NIH‑aligned)
These are post‑MVP enhancements. The current design (raw/analysis separation, Zarr/Parquet layout, run manifests) is intended to make these easy to add later.
- Persistent Identifiers & Rich Metadata: extend
manifest.yamlto include dataset DOI/ARK, award/grant numbers, ORCID for contributors, ROR for institution, data license (e.g., CC‑BY/CC0), instrument calibration, software versions, git commit, environment hash, and units/sampling info. - Repository Export Tooling: CLI to package a run (Zarr/Parquet + manifest) for deposit to discipline‑specific or generalist repositories; include mappings to common schemas (e.g., DataCite/schema.org) and generate checksums/fixity.
- Access Tiers & Privacy Controls: per‑run flags for
open | embargoed | controlled, embargo dates, and optional de‑identification workflow with consent/use‑restriction notes (for human/regulated data). - Provenance Capture: record analysis lineage (inputs → transforms → outputs) using a lightweight model (e.g., W3C PROV/RO‑Crate); store pipeline configs, container/image digests, and exact Python environment (lockfile) alongside outputs.
- Data Integrity & Retention: periodic fixity checks (hash manifests), retention schedules, and optional replication to secondary storage/object store.
- DMP Boilerplate Generator: render a Data Management & Sharing Plan from project settings (formats, repositories, timelines, access level, license) for use in NSF/NIH submissions.
- CITATION & Codemeta: include
CITATION.cfffor software citation and optionalcodemeta.jsonfor machine‑readable project metadata.
⚙️ Configuration Preview (post‑MVP)
A small, human‑readable config.yaml will hold settings that enable NSF/NIH‑aligned exports without changing the storage core. Example:
project:
name: pylabhub-demo
organization: Example University
ror: "" # Research Organization Registry ID (optional)
award_number: "NSF-XXXXXXX" # or NIH grant
contacts:
- name: Alex Researcher
orcid: "0000-0000-0000-0000"
storage:
base_dir: ./runs
formats:
arrays: zarr
events: parquet
compression:
zarr_codec: blosc-zstd
zarr_level: 5
identifiers:
dataset_pid_scheme: doi # doi | ark | none
dataset_pid: null # filled on publish
sharing:
repository:
target: zenodo # or domain-specific repo
community: example # optional
access: open # open | embargoed | controlled
embargo_until: null # e.g., 2026-01-01
license: CC-BY-4.0
provenance:
record_git_commit: true
record_conda_lock: true
record_container_digest: true
analysis_lineage: prov # enable PROV/RO-Crate style tracking
export:
hdf5: true # provide HDF5 export package alongside native store
include_checksums: true
This keeps raw/analysis separation intact while making it straightforward to: assign PIDs, export to repositories, control access tiers/embargo, and capture provenance.
📂 Project Structure
pylabhub/
├── hub/ # central data hub
├── adapters/ # hardware I/O interfaces
├── connectors/ # integration with external software
├── examples/ # sample scripts and configs
├── docs/ # documentation
├── tests/ # unit + integration tests
🚀 Getting Started
Clone the repository:
git clone https://github.com/YOUR-USERNAME/pylabhub.git
cd pylabhub
Install in development mode:
pip install -e .
Run the simple loopback example:
python examples/simple_loopback.py
🤝 Contributing
At this stage, the focus is on building a working framework. Once the core hub, a first adapter, and a connector are stable, contribution guidelines will be added here. Until then, feedback and suggestions are welcome through issues.
📖 Documentation
Work in progress — see the docs/ folder for drafts.
Future docs will cover APIs, schema design, and hardware integration guides.
📜 License
This project is licensed under the BSD 3-Clause License.
© 2025 Quan Qing, Arizona State University
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pylabhub-0.0.1a0.tar.gz.
File metadata
- Download URL: pylabhub-0.0.1a0.tar.gz
- Upload date:
- Size: 16.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5e868078974446de94da84626573c1ac83c994c785f4c47d988b81d715b21d8f
|
|
| MD5 |
6cc97bf04d60d02a5ba4835284288dc3
|
|
| BLAKE2b-256 |
327aadf94fff7ad43f1fdc5638b11f98d87f411fb6442193285c1982c52a69b7
|
Provenance
The following attestation bundles were made for pylabhub-0.0.1a0.tar.gz:
Publisher:
release.yml on Qing-LAB/pylabhub
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pylabhub-0.0.1a0.tar.gz -
Subject digest:
5e868078974446de94da84626573c1ac83c994c785f4c47d988b81d715b21d8f - Sigstore transparency entry: 535289955
- Sigstore integration time:
-
Permalink:
Qing-LAB/pylabhub@bca73e55b3548cbe61509e6eecb0f2ebe60d62db -
Branch / Tag:
refs/tags/v0.0.1a0 - Owner: https://github.com/Qing-LAB
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@bca73e55b3548cbe61509e6eecb0f2ebe60d62db -
Trigger Event:
push
-
Statement type:
File details
Details for the file pylabhub-0.0.1a0-py3-none-any.whl.
File metadata
- Download URL: pylabhub-0.0.1a0-py3-none-any.whl
- Upload date:
- Size: 11.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f74c91c4dbb0df5b049905f84109af327e4d1a623fdefd231e93b7c4ab20b01b
|
|
| MD5 |
1394c6f99ab6c083617cc5c000a8641c
|
|
| BLAKE2b-256 |
72744eb04c4c28729f2aca08b212c40ff19f6cb627e91e6dff8858e950731252
|
Provenance
The following attestation bundles were made for pylabhub-0.0.1a0-py3-none-any.whl:
Publisher:
release.yml on Qing-LAB/pylabhub
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pylabhub-0.0.1a0-py3-none-any.whl -
Subject digest:
f74c91c4dbb0df5b049905f84109af327e4d1a623fdefd231e93b7c4ab20b01b - Sigstore transparency entry: 535290017
- Sigstore integration time:
-
Permalink:
Qing-LAB/pylabhub@bca73e55b3548cbe61509e6eecb0f2ebe60d62db -
Branch / Tag:
refs/tags/v0.0.1a0 - Owner: https://github.com/Qing-LAB
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@bca73e55b3548cbe61509e6eecb0f2ebe60d62db -
Trigger Event:
push
-
Statement type: