nomad-ml-workflows
A NOMAD plugin for supporting machine-learning workflows. It currently provides the Export Entries action, which turns a permission-scoped NOMAD search into a reusable JSON, Parquet, or CSV dataset.
📦 Installation
Install the package with pip:
pip install nomad-ml-workflows
To add the plugin to your NOMAD deployment, see Adding this plugin to NOMAD.
✨ Key functionality
The Export Entries action lets users:
- select entries with a NOMAD search query and ownership scope;
- export complete archives or selected archive paths, optionally resolving references;
- generate JSON, Parquet, or CSV datasets; and
- save the generated dataset to a staging project/upload as a ZIP archive or directory.
JSON preserves the selected archive structure. Parquet provides typed tabular data and preserves nested quantity values as a sharded dataset, while CSV represents nested values as JSON text.
Action input
The action form groups its fields into Search options and Export options. Its underlying input shape is:
{
"user_id": "<injected by NOMAD>",
"upload_id": "<destination-project-id>",
"search_settings": {
"owner": "visible",
"max_entries": 1000,
"query": "{\n \"entry_type\": \"ELNSample\"\n}",
"required": [
{
"type": "include",
"path": "results.method",
"resolve_references": false
}
]
},
"export_settings": {
"file_format": "parquet",
"create_zip_archive": true
}
}
user_id is required by the workflow but is not shown in the action form. The other
fields have these semantics:
| Field | Values and behavior |
|---|---|
upload_id |
ID of the staging project/upload that receives the exported artifacts. |
search_settings.owner |
visible, public, user, shared, or staging; defaults to visible. |
search_settings.max_entries |
Positive requested limit, bounded by max_entries_export_limit; defaults to the smaller of 1,000 and the deployment limit. |
search_settings.query |
NOMAD search query supplied as JSON text. The query can be copied from View API Call in a NOMAD search app. |
search_settings.required |
Archive paths to include. An empty list exports the complete archive. Set resolve_references to include content reached through references below a path. |
export_settings.file_format |
parquet, csv, or json; defaults to parquet. |
export_settings.create_zip_archive |
Defaults to true. Set it to false to publish a project subdirectory instead of a ZIP file. |
The current public model supports include directives only. Exclusion is not yet available through the action form.
Generated output
The generated ZIP archive or directory contains:
| File | Contents |
|---|---|
data.json or data.csv |
The exported archive data in the selected single-file format. This file is omitted when no entries match. |
data.parquet/part-NNNNN.parquet |
The exported Parquet dataset. Each deterministic part uses the same schema and the directory is omitted when no entries match. |
selected_entries.json |
The ordered entry_id and upload_id pairs selected for export. |
metadata.json |
The export metadata under data, together with its JSON schema and an explanatory note. |
The metadata records the original input, search timing, export-limit state, and the
num_entries_available, num_entries_selected, and num_entries_exported counts. It
also records nomad_deployment_api_host, nomad_version, and
nomad_ml_workflows_version. A zero-match export still contains metadata.json and
an empty selected_entries.json.
PyArrow reads the complete Parquet dataset directly from its directory:
import pyarrow.parquet as pq
table = pq.read_table('data.parquet')
⚙️ Configuration
Configure the action entry point in the nomad.yaml of the NOMAD Oasis:
plugins:
entry_points:
options:
nomad_ml_workflows.actions:export_entries:
max_entries_export_limit: 100000
# Deployment cap for one Export Entries action.
read_archives_timeout: 7200
# Start-to-close timeout, in seconds, for reading archives and
# writing the selected output format.
write_tabular_timeout: 7200
# Start-to-close timeout, in seconds, for writing the output
# tabular artifact.
max_write_buffer_size_bytes: 67108864 # 64 MB
# Maximum number of encoded NDJSON input bytes represented by parsed
# rows buffered before writing to the output tabular file. Protects
# against rows with large encoded representations. One oversized row
# may exceed this target.
max_write_buffer_size_rows: 1024
# Maximum number of parsed rows buffered before writing to the output
# tabular file. Protects against many small or sparse rows whose Python
# and Arrow representations are much larger than their NDJSON bytes.
For deployments where PyArrow uses the mimalloc memory-pool backend, the CPU
action worker can be started with immediate memory purging enabled:
MIMALLOC_PURGE_DELAY=0
The tabular exporter requests that unused Arrow memory be released after every
batch. Setting the environment variable MIMALLOC_PURGE_DELAY=0 asks mimalloc
to return unused pages to the operating system immediately. In local measurements,
this produced slightly lower post-export RSS.
This setting is optional and must be present in the CPU worker environment before the worker starts. Immediate purging may trade some allocation performance for lower retained RSS.
🚀 Adding this plugin to NOMAD
NOMAD Oasis
Follow the NOMAD Oasis configuration documentation to add and enable the plugin in an Oasis.
Local NOMAD development installation
Use the dedicated
nomad-distro-dev repository for
an integrated local NOMAD development environment.
🛠️ Development
Clone the repository and create a virtual environment with Python 3.10, 3.11, or 3.12:
git clone https://github.com/FAIRmat-NFDI/nomad-ml-workflows.git
cd nomad-ml-workflows
python3.12 -m venv .pyenv
. .pyenv/bin/activate
python -m pip install --upgrade pip
python -m pip install uv
uv pip install -e '.[dev]'
Run the focused Export Entries tests:
pytest tests/actions/export_entries/test_models.py \
tests/actions/export_entries/test_utils.py
Run linting and formatting checks with Ruff:
ruff check .
ruff format . --check
For interactive test debugging, pass --pdb to pytest. To serve the documentation
locally, use the development dependencies and run:
mkdocs serve
👥 Main contributors
| Name | |
|---|---|
| Sarthak Kapoor | sarthak.kapoor@physik.hu-berlin.de |
📄 License
This project is licensed under the MIT License. See LICENSE.
Metadata
Release files for nomad-ml-workflows 0.0.15
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nomad_ml_workflows-0.0.15.tar.gz | 151.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nomad_ml_workflows-0.0.15-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 202.6 kB
Release files / nomad_ml_workflows-0.0.15.tar.gz
| Download URL | nomad_ml_workflows-0.0.15.tar.gz |
|---|---|
| Size | 151.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
80c58dda7969548fdc27457c40403b8d1fbf84d87ab0d3558616b055fb51570f
|
|
BLAKE2b-256 checksum How to use checksums |
495856f769fe1ee0ab8ca204543798aac3def89cd2633f660e24eda442492ecb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / nomad_ml_workflows-0.0.15-py3-none-any.whl
| Download URL | nomad_ml_workflows-0.0.15-py3-none-any.whl |
|---|---|
| Size | 50.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ac5c76b3ec43aacf5e3ab2e7954e3d76ca7347c2bd1a3d62ffb2998993ec8345
|
|
BLAKE2b-256 checksum How to use checksums |
63d987f1bdde48dac26fdee99e1bf898214184917e14ddebc988c38a603bed29
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|