NeuroDAGs
An Extensible and Declarative DAG Framework for Reproducible Neuroscience Workflows
M/EEG studies generate many interdependent intermediate derivatives. Recomputing full pipelines is wasteful; reusing valid intermediates is non-trivial. Large-scale studies require reproducible, extensible, and efficient workflows. NeuroDAGs addresses this with a declarative, graph-based framework for scalable and reusable derivative computation.
Docs | Comparison with Snakemake/Pydra | Poster BRaIN Symposium 2026 Montreal
Core Idea
Pipelines are defined as a directed acyclic graph (DAG) of computation nodes that output reusable derivatives, executed for each input file.
Design Principles
- Reproducible, transparent workflows defined declaratively in YAML — version-controllable and LLM-friendly.
- Uniform node abstraction — preprocessing, features, and any custom nodes are treated identically.
- Directory-agnostic — outputs mirror inputs' organization. Derivatives are labeled with a
@DerivativeNamesuffix. - xarray-centered outputs — derivatives stored as language-agnostic, metadata-rich, dimension-aware xarray → NetCDF.
- Graph-based reuse — if a derivative is already computed and
overwrite=False, it is skipped automatically.
Features
- Agnostic to data organization / directory hierarchy
- SLURM / HPC friendly with file-level parallelism via joblib
- Graph-based caching: skip already-computed derivatives
- Extensible node system — add nodes without forking the package
- YAML-based declarative configuration
- Unified CLI:
neurodags run,dry-run,dataframe,dag,view,validate,tui - Built-in Terminal User Interface (TUI) for pipeline management and execution
- Built-in nodes for preprocessing, spectral analysis, entropy, complexity, and data transformations
- Dataframe assembly (wide or long format) from derivative artifacts
- Dry-run mode — inspect planned computations without executing
- Built-in Dash-Plotly explorer for
.fifand.ncfiles
Installation
pip install neurodags
# Or with TUI support
pip install neurodags[tui]
With uv (recommended):
uv add neurodags
# Or with TUI support
uv add neurodags[tui]
Quickstart
See the quickstart example — full synthetic pipeline, no real data required.
CLI Reference
NeuroDAGs installs a unified neurodags command. Global flags (placed before the subcommand) apply everywhere:
neurodags --log-level WARNING run pipeline.yml # suppress INFO output
neurodags --log-file run.jsonl run pipeline.yml # also write logs to JSONL file
All subcommands accept -d/--datasets <path> to override the datasets YAML defined in the pipeline file.
The JSONL log file loads directly as a dataframe:
import pandas as pd
df = pd.read_json("run.jsonl", lines=True)
Validation
neurodags validate pipeline.yml # load config, print datasets / derivatives summary
neurodags validate pipeline.yml -d alt.yml # override datasets
Execution
neurodags run pipeline.yml # run all derivatives in DerivativeList
neurodags run pipeline.yml --derivative CleanedEEG # run a specific derivative
neurodags run pipeline.yml --derivative A --derivative B # run multiple
# parallelism
neurodags run pipeline.yml --n-jobs 4 # 4 workers
neurodags run pipeline.yml --n-jobs -1 # all cores
neurodags run pipeline.yml --n-jobs 4 --joblib-backend loky --joblib-prefer processes
# subset / error control
neurodags run pipeline.yml --max-files-per-dataset 10
neurodags run pipeline.yml --only-index 0 5 12 # process only these file indices
neurodags run pipeline.yml --skip-errors # skip files with a prior .error marker
neurodags run pipeline.yml --raise-on-error # stop on first failure
Dry Run
Inspect the execution plan without running any nodes. Returns a CSV/Parquet describing each file, derivative, and whether the output is already cached.
neurodags dry-run pipeline.yml # all derivatives
neurodags dry-run pipeline.yml --derivative CleanedEEG # one derivative
neurodags dry-run pipeline.yml --output plan.csv # save to CSV
neurodags dry-run pipeline.yml --output plan.parquet # or Parquet
neurodags dry-run pipeline.yml --n-jobs 4 # parallel dry-run
neurodags dry-run pipeline.yml --skip-errors # exclude errored files from plan
Status
Quick summary of done / missing / errored counts per derivative — no CSV needed.
neurodags status pipeline.yml # summary table
neurodags status pipeline.yml --derivative Alpha # filter to one derivative
neurodags status pipeline.yml --list-errors # print errored file paths + .error paths
neurodags status pipeline.yml --list-missing # print missing file paths
neurodags status pipeline.yml --list-errors --list-missing
neurodags status pipeline.yml --n-jobs 4 # parallelize underlying dry-run
neurodags status pipeline.yml --format json # machine-readable JSON
Exit code 0 only when all derivatives are complete (no missing, no errored); 1 otherwise.
Source File Count
neurodags count-inputs pipeline.yml # number of source (input) files the pipeline will process
neurodags count-inputs pipeline.yml --derivative CleanedEEG # count for a specific derivative
Dataframe Assembly
neurodags dataframe pipeline.yml --format wide --output features.csv
neurodags dataframe pipeline.yml --format long --output features.parquet
neurodags dataframe pipeline.yml --include-derivative PowerSpectrum --include-derivative BandPower
neurodags dataframe pipeline.yml --max-files-per-dataset 5
neurodags dataframe pipeline.yml --n-jobs 4 # parallel file-level collection
neurodags dataframe pipeline.yml --n-jobs -1 # all cores
DAG Visualization
neurodags dag pipeline.yml # print Mermaid text
neurodags dag pipeline.yml --html pipeline_dag.html # export to HTML (ELK layout by default)
neurodags dag pipeline.yml --html pipeline_dag.html --open # export and open in browser
neurodags dag pipeline.yml --derivative CleanedEEG --html d.html # node-level DAG for one derivative
neurodags dag pipeline.yml --html pipeline_dag.html --layout dagre # offline fallback (no CDN)
HTML output uses the ELK layout engine by default — orthogonal edge routing with crossing minimisation, significantly cleaner than curved edges for dense pipelines. Use --layout dagre for offline environments.
File Explorer
neurodags view path/to/file.fif # launch Dash-Plotly explorer for MNE .fif files
neurodags view path/to/file.nc # launch Dash-Plotly explorer for xarray NetCDF files
SLURM / HPC Scripts
Generate ready-to-submit SLURM array job scripts:
neurodags slurm-script pipeline.yml # per-file pattern (default)
neurodags slurm-script pipeline.yml --pattern flat # file × derivative flat array
neurodags slurm-script pipeline.yml --pattern chained # chained per-derivative arrays
neurodags slurm-script pipeline.yml --output run_array.sh # write to file
neurodags slurm-script pipeline.yml --derivative CleanedEEG # restrict to specific derivatives
See HPC guide for details on each pattern.
TUI (Terminal User Interface)
Requires pip install neurodags[tui]:
neurodags tui # launch empty TUI, load config interactively
neurodags tui pipeline.yml # launch with config pre-loaded
neurodags tui pipeline.yml -d alt.yml # with datasets override
Development
git clone https://github.com/yjmantilla/neurodags
cd neurodags
uv sync --all-extras --all-groups # creates .venv and installs all deps incl. dev/test/docs
uv run pre-commit install
Key commands (all via uv run):
uv run ruff check src/ # lint (fix: uv run ruff check src/ --fix)
uv run black --check . # format check (fix: uv run black .)
uv run pytest -q # run tests
uv run pytest -s -q --no-cov --pdb # debug a failing test
uv run sphinx-build -b html docs docs/_build/html -W --keep-going # build docs
rm -rf docs/_build # clean docs
No uv? Install it with
pip install uvorcurl -Ls https://astral.sh/uv/install.sh | sh. All commands above work with plainpython/piptoo — swapuv run→ activate.venv,uv sync→pip install -e .[dev,test,docs].
Project Structure
my_project/
├── datasets.yml # Dataset sources and paths
├── pipeline.yml # Derivative definitions and execution list
└── custom_nodes.py # Optional custom node definitions
Quick Example
datasets.yml
my_dataset:
name: MyDataset
file_pattern:
local: data/**/*.vhdr
hpc: /cluster/BIDS/**/*.vhdr
derivatives_path:
local: outputs/
hpc: /cluster/scratch/out
pipeline.yml
datasets: datasets.yml
mount_point: local
new_definitions: custom_nodes.py # optional
DerivativeDefinitions:
CleanedEEG:
nodes:
- id: 0
derivative: SourceFile
- id: 1
node: basic_preprocessing
args:
mne_object: id.0
resample: 256
filter_args:
l_freq: 0.5
h_freq: 110
PowerSpectrum:
for_dataframe: True
nodes:
- id: 0
derivative: CleanedEEG.fif
- id: 1
node: mne_spectrum_array
args:
meeg: id.0
method: multitaper
DerivativeList:
- CleanedEEG
- PowerSpectrum
Python
from neurodags.loaders import load_configuration
from neurodags.orchestrators import run_pipeline
config = load_configuration("pipeline.yml")
# Run all derivatives in "DerivativeList", auto-sorted by dependency order
run_pipeline(config)
# Or run specific ones (also sorted by dependency order)
run_pipeline(config, derivatives=["CleanedEEG"])
CLI
neurodags validate pipeline.yml
# Run all derivatives in DerivativeList (dependency-sorted)
neurodags run pipeline.yml
# Or run specific ones
neurodags run pipeline.yml --derivative CleanedEEG
Custom Nodes
Add nodes without modifying or forking the package:
# custom_nodes.py
from neurodags.nodes import register_node
from neurodags.definitions import Artifact, NodeResult
@register_node
def my_node(data) -> NodeResult:
result = compute(data)
return NodeResult(
artifacts={
".nc": Artifact(
item=result,
writer=lambda path: result.to_netcdf(path),
),
},
)
Key rules:
- A node is a function decorated with
@register_node. - It returns a
NodeResult. - A
NodeResultcontainsartifacts— a dict mapping file extension toArtifact(item, writer).
Dataframe Assembly
from neurodags.orchestrators import build_derivative_dataframe
df = build_derivative_dataframe("pipeline.yml", output_format="wide")
Derivatives marked for_dataframe: True are collected automatically. Supports "wide" (one row per file) and "long" (one row per value) formats.
CLI equivalent:
neurodags dataframe pipeline.yml --format wide --output derivative_dataframe.csv
Parallel Execution
# pipeline.yml
n_jobs: 4 # -1 = all cores, 1 or null = serial
joblib_backend: loky
joblib_prefer: processes
Or via Python:
run_pipeline(config, derivatives=["MyDerivative"], n_jobs=4)
Or via CLI:
neurodags run pipeline.yml --derivative MyDerivative --n-jobs 4
Visualization
neurodags view path/to/file.fif
neurodags view path/to/file.nc
# Alternative module entry point
python -m neurodags.visualization path/to/file.fif
python -m neurodags.visualization path/to/file.nc
Built-in Dash-Plotly explorer with dimension-aware UI — dropdown per axis, plot types: Line, Scatter, Bar, Heatmap.
Inspection (Dry Run)
# All derivatives in DerivativeList
run_pipeline(config, dry_run=True)
# Or a specific one
run_pipeline(config, derivatives=["MyDerivative"], dry_run=True)
Returns a dataframe describing the execution plan without running any nodes. When a node fails, a .error marker file is written with the error message — failed files are retried on the next run. If a retry succeeds, the .error marker is automatically removed.
CLI equivalent:
# All derivatives in DerivativeList
neurodags dry-run pipeline.yml --output dry_run_results.csv
# Or a specific one
neurodags dry-run pipeline.yml --derivative MyDerivative --output dry_run_results.csv
Derivative Flags
| Flag | Default | Description |
|---|---|---|
save |
True |
Persist artifacts to disk. False = compute but don't write. |
overwrite |
False |
Force recompute even if output exists. |
for_dataframe |
False |
Include this derivative in build_derivative_dataframe. |
Custom Node Definitions
Point new_definitions to one or more Python files:
new_definitions:
- custom_nodes/my_nodes.py
- /abs/path/to/other_nodes.py
Relative paths are resolved from the pipeline YAML location.
Documentation
https://yjmantilla.github.io/neurodags/
HDF5 / NetCDF Note
If you encounter RuntimeError: NetCDF: HDF error:
uv run pip install --no-binary=h5py h5py
# or without uv:
pip install --no-binary=h5py h5py
Contributing
See CONTRIBUTING.md.
License
MIT. See LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file neurodags-0.3.0.tar.gz.
File metadata
- Download URL: neurodags-0.3.0.tar.gz
- Upload date:
- Size: 101.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7035a2d072292f114181de786a5124e4da737fdbb921a088d2f19aed0a164e6c
|
|
| MD5 |
c665482462c8d470bd37b3a75466562d
|
|
| BLAKE2b-256 |
4236efac9f9283ac971edd508cab64b0cd6b7efabff25728eb786f3f47fc5c29
|
Provenance
The following attestation bundles were made for neurodags-0.3.0.tar.gz:
Publisher:
publish.yml on yjmantilla/neurodags
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
neurodags-0.3.0.tar.gz -
Subject digest:
7035a2d072292f114181de786a5124e4da737fdbb921a088d2f19aed0a164e6c - Sigstore transparency entry: 2166353173
- Sigstore integration time:
-
Permalink:
yjmantilla/neurodags@d9b7c343c31d96838b373f6d109e35968c9c964f -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/yjmantilla
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@d9b7c343c31d96838b373f6d109e35968c9c964f -
Trigger Event:
push
-
Statement type:
File details
Details for the file neurodags-0.3.0-py3-none-any.whl.
File metadata
- Download URL: neurodags-0.3.0-py3-none-any.whl
- Upload date:
- Size: 116.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f1c57b7f22586f90d9aee4096a4472fa04861c596f3c0582e7b459292785ab27
|
|
| MD5 |
14bffd2389fb0dbe3b1d49c1f75a1617
|
|
| BLAKE2b-256 |
e51812967cf06902ee4ab27a1d5cfcc351830ba77747ed86b041153dd16f8373
|
Provenance
The following attestation bundles were made for neurodags-0.3.0-py3-none-any.whl:
Publisher:
publish.yml on yjmantilla/neurodags
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
neurodags-0.3.0-py3-none-any.whl -
Subject digest:
f1c57b7f22586f90d9aee4096a4472fa04861c596f3c0582e7b459292785ab27 - Sigstore transparency entry: 2166353274
- Sigstore integration time:
-
Permalink:
yjmantilla/neurodags@d9b7c343c31d96838b373f6d109e35968c9c964f -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/yjmantilla
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@d9b7c343c31d96838b373f6d109e35968c9c964f -
Trigger Event:
push
-
Statement type: