Skip to main content

nomad-ml-workflows

A NOMAD plugin for supporting machine-learning workflows. It currently provides the Export Entries action, which turns a permission-scoped NOMAD search into a reusable JSON, Parquet, or CSV dataset.

📦 Installation

Install the package with pip:

pip install nomad-ml-workflows

To add the plugin to your NOMAD deployment, see Adding this plugin to NOMAD.

✨ Key functionality

The Export Entries action lets users:

  • select entries with a NOMAD search query and ownership scope;
  • export complete archives or selected archive paths, optionally resolving references;
  • generate JSON, Parquet, or CSV datasets; and
  • save the generated dataset to a staging project/upload as a ZIP archive or directory.

JSON preserves the selected archive structure. Parquet provides typed tabular data and preserves nested quantity values as a sharded dataset, while CSV represents nested values as JSON text.

Action input

The action form groups its fields into Search options and Export options. Its underlying input shape is:

{
  "user_id": "<injected by NOMAD>",
  "upload_id": "<destination-project-id>",
  "search_settings": {
    "owner": "visible",
    "max_entries": 1000,
    "query": "{\n  \"entry_type\": \"ELNSample\"\n}",
    "required": [
      {
        "type": "include",
        "path": "results.method",
        "resolve_references": false
      }
    ]
  },
  "export_settings": {
    "file_format": "parquet",
    "create_zip_archive": true
  }
}

user_id is required by the workflow but is not shown in the action form. The other fields have these semantics:

Field Values and behavior
upload_id ID of the staging project/upload that receives the exported artifacts.
search_settings.owner visible, public, user, shared, or staging; defaults to visible.
search_settings.max_entries Positive requested limit, bounded by max_entries_export_limit; defaults to the smaller of 1,000 and the deployment limit.
search_settings.query NOMAD search query supplied as JSON text. The query can be copied from View API Call in a NOMAD search app.
search_settings.required Archive paths to include. An empty list exports the complete archive. Set resolve_references to include content reached through references below a path.
export_settings.file_format parquet, csv, or json; defaults to parquet.
export_settings.create_zip_archive Defaults to true. Set it to false to publish a project subdirectory instead of a ZIP file.

The current public model supports include directives only. Exclusion is not yet available through the action form.

Generated output

The generated ZIP archive or directory contains:

File Contents
data.json or data.csv The exported archive data in the selected single-file format. This file is omitted when no entries match.
data.parquet/part-NNNNN.parquet The exported Parquet dataset. Each deterministic part uses the same schema and the directory is omitted when no entries match.
selected_entries.json The ordered entry_id and upload_id pairs selected for export.
metadata.json The export metadata under data, together with its JSON schema and an explanatory note.

The metadata records the original input, search timing, export-limit state, and the num_entries_available, num_entries_selected, and num_entries_exported counts. It also records nomad_deployment_api_host, nomad_version, and nomad_ml_workflows_version. A zero-match export still contains metadata.json and an empty selected_entries.json.

PyArrow reads the complete Parquet dataset directly from its directory:

import pyarrow.parquet as pq

table = pq.read_table('data.parquet')

⚙️ Configuration

Configure the action entry point in the nomad.yaml of the NOMAD Oasis:

plugins:
  entry_points:
    options:
      nomad_ml_workflows.actions:export_entries:
        max_entries_export_limit: 100000
        # Deployment cap for one Export Entries action.

        read_archives_timeout: 7200
        # Start-to-close timeout, in seconds, for reading archives and
        # writing the selected output format.

        write_tabular_timeout: 7200
        # Start-to-close timeout, in seconds, for writing the output 
        # tabular artifact.

        max_write_buffer_size_bytes: 67108864  # 64 MB
        # Maximum number of encoded NDJSON input bytes represented by parsed
        # rows buffered before writing to the output tabular file. Protects
        # against rows with large encoded representations. One oversized row
        # may exceed this target.

        max_write_buffer_size_rows: 1024
        # Maximum number of parsed rows buffered before writing to the output
        # tabular file. Protects against many small or sparse rows whose Python
        # and Arrow representations are much larger than their NDJSON bytes.

🚀 Adding this plugin to NOMAD

NOMAD Oasis

Follow the NOMAD plugin installation documentation to add and enable the plugin in an Oasis.

Local NOMAD development installation

Use the dedicated nomad-distro-dev repository for an integrated local NOMAD development environment.

🛠️ Development

Clone the repository and create a virtual environment with Python 3.10, 3.11, or 3.12:

git clone https://github.com/FAIRmat-NFDI/nomad-ml-workflows.git
cd nomad-ml-workflows
python3.12 -m venv .pyenv
. .pyenv/bin/activate
python -m pip install --upgrade pip
python -m pip install uv
uv pip install -e '.[dev]'

Run the focused Export Entries tests:

pytest tests/actions/export_entries/test_models.py \
  tests/actions/export_entries/test_utils.py

Run linting and formatting checks with Ruff:

ruff check .
ruff format . --check

For interactive test debugging, pass --pdb to pytest. To serve the documentation locally, use the development dependencies and run:

mkdocs serve

👥 Main contributors

Name Email
Sarthak Kapoor sarthak.kapoor@physik.hu-berlin.de

📄 License

This project is licensed under the MIT License. See LICENSE.

Metadata

Release files for nomad-ml-workflows 0.0.12

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nomad-ml-workflows 0.0.12
File Size Uploaded
nomad_ml_workflows-0.0.12.tar.gz 142.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nomad-ml-workflows 0.0.12
File Interpreter ABI Platform
nomad_ml_workflows-0.0.12-py3-none-any.whl Python 3 none any Details

Total release size: 182.3 kB

Release files / nomad_ml_workflows-0.0.12.tar.gz

Download URL nomad_ml_workflows-0.0.12.tar.gz
Size 142.6 kB
Tags Source
SHA-256 checksum
How to use checksums
4e16a8dfcf1d49b6b9639b103ae1cf849231343758428ce6a3f0b4876ffa9705
BLAKE2b-256 checksum
How to use checksums
cf14e88cb8b57d302049e96b9203e3aa2e2243497138f532ad1f0f8c4648a28c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.25

Release files / nomad_ml_workflows-0.0.12-py3-none-any.whl

Download URL nomad_ml_workflows-0.0.12-py3-none-any.whl
Size 39.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9f4feb08d29650b9c160a19bad13c77ce3d8f1f87e5f536bf6b419831fb94d51
BLAKE2b-256 checksum
How to use checksums
67bd0b606a2eb4d256fc3fde8bc4a20d2416de6ce22a264adb34ca50720f99dd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.25

Release history Release notifications | RSS feed

0.0.15

2 release files

0.0.13

2 release files

This release

0.0.12 This release

2 release files

0.0.11

2 release files

0.0.10

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page