nomad-ml-workflows
A NOMAD plugin for supporting machine-learning workflows. It currently provides the Export Entries action, which turns a permission-scoped NOMAD search into a reusable JSON, Parquet, or CSV dataset.
📦 Installation
Install the package with pip:
pip install nomad-ml-workflows
To add the plugin to your NOMAD deployment, see Adding this plugin to NOMAD.
✨ Key functionality
The Export Entries action lets users:
- select entries with a NOMAD search query and ownership scope;
- export complete archives or selected archive paths, optionally resolving references;
- generate JSON, Parquet, or CSV datasets; and
- save the generated dataset to a staging project/upload as a ZIP archive or directory.
JSON preserves the selected archive structure. Parquet provides typed tabular data and preserves nested quantity values as a sharded dataset, while CSV represents nested values as JSON text.
Action input
The action form groups its fields into Search options and Export options. Its underlying input shape is:
{
"user_id": "<injected by NOMAD>",
"upload_id": "<destination-project-id>",
"search_settings": {
"owner": "visible",
"max_entries": 1000,
"query": "{\n \"entry_type\": \"ELNSample\"\n}",
"required": [
{
"type": "include",
"path": "results.method",
"resolve_references": false
}
]
},
"export_settings": {
"file_format": "parquet",
"create_zip_archive": true
}
}
user_id is required by the workflow but is not shown in the action form. The other
fields have these semantics:
| Field | Values and behavior |
|---|---|
upload_id |
ID of the staging project/upload that receives the exported artifacts. |
search_settings.owner |
visible, public, user, shared, or staging; defaults to visible. |
search_settings.max_entries |
Positive requested limit, bounded by max_entries_export_limit; defaults to the smaller of 1,000 and the deployment limit. |
search_settings.query |
NOMAD search query supplied as JSON text. The query can be copied from View API Call in a NOMAD search app. |
search_settings.required |
Archive paths to include. An empty list exports the complete archive. Set resolve_references to include content reached through references below a path. |
export_settings.file_format |
parquet, csv, or json; defaults to parquet. |
export_settings.create_zip_archive |
Defaults to true. Set it to false to publish a project subdirectory instead of a ZIP file. |
The current public model supports include directives only. Exclusion is not yet available through the action form.
Generated output
The generated ZIP archive or directory contains:
| File | Contents |
|---|---|
data.json or data.csv |
The exported archive data in the selected single-file format. This file is omitted when no entries match. |
data.parquet/part-NNNNN.parquet |
The exported Parquet dataset. Each deterministic part uses the same schema and the directory is omitted when no entries match. |
selected_entries.json |
The ordered entry_id and upload_id pairs selected for export. |
metadata.json |
The export metadata under data, together with its JSON schema and an explanatory note. |
The metadata records the original input, search timing, export-limit state, and the
num_entries_available, num_entries_selected, and num_entries_exported counts. It
also records nomad_deployment_api_host, nomad_version, and
nomad_ml_workflows_version. A zero-match export still contains metadata.json and
an empty selected_entries.json.
PyArrow reads the complete Parquet dataset directly from its directory:
import pyarrow.parquet as pq
table = pq.read_table('data.parquet')
⚙️ Configuration
Configure the action entry point in the nomad.yaml of the NOMAD Oasis:
plugins:
entry_points:
options:
nomad_ml_workflows.actions:export_entries:
max_entries_export_limit: 100000
# Deployment cap for one Export Entries action.
read_archives_timeout: 7200
# Start-to-close timeout, in seconds, for reading archives and
# writing the selected output format.
write_tabular_timeout: 7200
# Start-to-close timeout, in seconds, for writing the output
# tabular artifact.
max_write_buffer_size_bytes: 67108864 # 64 MB
# Maximum number of encoded NDJSON input bytes represented by parsed
# rows buffered before writing to the output tabular file. Protects
# against rows with large encoded representations. One oversized row
# may exceed this target.
max_write_buffer_size_rows: 1024
# Maximum number of parsed rows buffered before writing to the output
# tabular file. Protects against many small or sparse rows whose Python
# and Arrow representations are much larger than their NDJSON bytes.
🚀 Adding this plugin to NOMAD
NOMAD Oasis
Follow the NOMAD plugin installation documentation to add and enable the plugin in an Oasis.
Local NOMAD development installation
Use the dedicated
nomad-distro-dev repository for
an integrated local NOMAD development environment.
🛠️ Development
Clone the repository and create a virtual environment with Python 3.10, 3.11, or 3.12:
git clone https://github.com/FAIRmat-NFDI/nomad-ml-workflows.git
cd nomad-ml-workflows
python3.12 -m venv .pyenv
. .pyenv/bin/activate
python -m pip install --upgrade pip
python -m pip install uv
uv pip install -e '.[dev]'
Run the focused Export Entries tests:
pytest tests/actions/export_entries/test_models.py \
tests/actions/export_entries/test_utils.py
Run linting and formatting checks with Ruff:
ruff check .
ruff format . --check
For interactive test debugging, pass --pdb to pytest. To serve the documentation
locally, use the development dependencies and run:
mkdocs serve
👥 Main contributors
| Name | |
|---|---|
| Sarthak Kapoor | sarthak.kapoor@physik.hu-berlin.de |
📄 License
This project is licensed under the MIT License. See LICENSE.
Metadata
Release files for nomad-ml-workflows 0.0.13
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nomad_ml_workflows-0.0.13.tar.gz | 142.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nomad_ml_workflows-0.0.13-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 182.4 kB
Release files / nomad_ml_workflows-0.0.13.tar.gz
| Download URL | nomad_ml_workflows-0.0.13.tar.gz |
|---|---|
| Size | 142.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3b53486f0cdd18331d846bab2b6726b04a71a911ac288fb3cf679e46513c26ab
|
|
BLAKE2b-256 checksum How to use checksums |
0b8ec64ec18a71f5b57853e8c14c42ca737c032e817422fcec59cf68fba04e9c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|
Release files / nomad_ml_workflows-0.0.13-py3-none-any.whl
| Download URL | nomad_ml_workflows-0.0.13-py3-none-any.whl |
|---|---|
| Size | 39.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
31aad37ec268baea4564a39a0881267fe09f563a2efe39e81c3f73e531842a57
|
|
BLAKE2b-256 checksum How to use checksums |
161b7b2c267d5381d6d05b25e58064dff7345f94f8a359d8eceb68de43a8bf65
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.25
|