SEAMM Extract Clusters Plug-in
A SEAMM plug-in for extracting molecular clusters (trimers, tetramers, … larger n-mers) from a condensed-phase, typically periodic, configuration as unwrapped, non-periodic structures – e.g. many-body training data and diagnostics for machine-learned force fields.
Free software: BSD-3-Clause
Documentation: https://molssi-seamm.github.io/extract_clusters_step/index.html
Features
Clusters are connected subgraphs of a molecular contact graph (molecules are in contact if their contact atoms are within a cutoff, minimum image), so chains, rings and stars all occur – not just the most compact cluster.
Any cluster size, and several sizes per frame (e.g. 3, 4).
Works on many frames in one step – all configurations of a system holding a trajectory, or any standard SEAMM structure selection – so no loop is needed and the clusters land in one system.
Stratification so the set is flat in a spread coordinate (radius of gyration or largest centroid separation), with bin edges from equal quantiles of a pilot sample or given explicitly; optional balancing over the contact-graph motif.
Clusters are unwrapped across the periodic boundary, centred, non-periodic, with molecules and bonds intact.
Provenance on every configuration: unique names <frame>_<seed>_<m1-m2-...> and #ExtractClusters#scan properties (size, spread, motif, contacts, source molecules) that survive SDF/extxyz export; a clusters.csv summary per step.
Works on any cell (orthorhombic fast path via a periodic KD-tree; exact minimum image otherwise) and on non-periodic sources.
Acknowledgements
This package was created with Cookiecutter and the molssi-seamm/cookiecutter-seamm-plugin project template.
Developed by the Molecular Sciences Software Institute (MolSSI), which receives funding from the National Science Foundation under award ACI-1547580
History
- 2026.9.19 – Centre coordination: stored, and selectable
Every cluster now carries a centre coordination property – the number of contacts of its most-connected molecule within the cluster – and the full degrees sequence, in clusters.csv and summary.json too. For five or more molecules the motif is only labelled by its number of contacts, so this is what distinguishes a 4-star from a 5-chain.
A new “Centre coordination” option accepts only clusters with the given coordination(s), e.g. 3 for star tetramers or 4 for a complete first shell, applied at acceptance so the set need not be post-filtered. Candidates rejected for it are counted and reported.
The random-seed help and the user guide now say that a supplementary run over the same frames must use a different seed.
- 2026.9.18.2 – Bugfix: the dialog failed to open
Opening the step’s dialog failed with “‘LabeledCombobox’ object has no attribute ‘entry’”: the binding that updates the motif list when the cluster sizes change assumed the wrong kind of widget. Fixed, and the dialog is now exercised by a test.
- 2026.9.18.1 – Extract from many structures in one step
The step now takes the standard SEAMM structure selection: the current configuration (the default, as before), all or the last or first configurations of the current system, of all systems, or of systems chosen by name, or a variable holding a list of configurations. Selecting all the configurations of a system that holds a trajectory extracts from every frame in one step, with no loop, and puts all the clusters in one system. Each frame’s clusters are prefixed with its configuration name; clusters.csv and summary.json record the frame; the report aggregates over frames and lists the count per frame. One random stream covers the whole run, so the printed seed reproduces it.
Requires seamm 2026.9.18.1 and molsystem 2026.9.17.2 or later.
- 2026.9.18 – Reproducible seeds and motif selection
The random seed actually used is now always printed and recorded in a new summary.json in the step directory (with the source, bin edges and counts by motif), so a run made with the seed set to “random” can be reproduced by entering the printed value.
A new “Restrict to motifs” option accepts only clusters with the given contact-graph topology, e.g. ring or ring, star. The dialog offers the motifs the requested cluster sizes can produce. Candidates rejected for their motif are counted and reported, since rare motifs use up the attempt budget.
- 2026.9.17.1 – Bugfix: a structure without bonds gave clusters of atoms
Molecules are identified from the bonds, so a configuration read from a format that carries no connectivity (extended XYZ without bond perception) was treated as one atom per molecule and the “clusters” were silently groups of atoms. The step now stops with a clear error pointing at the Read Structure “Perceive bonds” option. Configurations made only of noble-gas atoms or monatomic ions, which legitimately have no bonds, are still accepted.
- 2026.9.17 – Initial release of the Extract Clusters step
Extracts n-molecule clusters (trimers, tetramers, … larger n-mers) from the current, typically periodic, condensed-phase configuration as unwrapped, centred, non-periodic configurations in a new system, e.g. many-body training data and diagnostics for machine-learned force fields.
Clusters are connected subgraphs of a molecular contact graph (molecules are in contact if their contact atoms are within a cutoff, minimum image), so chains, rings and stars all occur; several sizes can be extracted from one frame.
Optional stratification so the set is flat in the radius of gyration or the largest centroid separation, with bin edges from equal quantiles of a pilot sample or given explicitly, and optional balancing over the contact-graph motif.
Provenance on every cluster: unique names <frame>_<seed>_<molecules> and #ExtractClusters#scan properties (size, spread, motif, contacts, bin, source molecules) that survive SDF/extxyz export, plus a clusters.csv per step.
Metadata
Release files for extract-clusters-step 2026.9.19
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| extract_clusters_step-2026.9.19.tar.gz | 76.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| extract_clusters_step-2026.9.19-py2.py3-none-any.whl | Python 3, Python 2 | none | any | Details |
Total release size: 112.4 kB
Release files / extract_clusters_step-2026.9.19.tar.gz
| Download URL | extract_clusters_step-2026.9.19.tar.gz |
|---|---|
| Size | 76.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fccfe045f6d49b4f9d41697092eb943497a8f1343ab09e85655e4e8fe983ec22
|
|
BLAKE2b-256 checksum How to use checksums |
91ddac025f90f638316a07edbea069c34c2331b8b1167b8f3519d2209abcebf4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / extract_clusters_step-2026.9.19-py2.py3-none-any.whl
| Download URL | extract_clusters_step-2026.9.19-py2.py3-none-any.whl |
|---|---|
| Size | 35.7 kB |
| Tags | Python 2 Python 3 |
|
SHA-256 checksum How to use checksums |
86fab4d5cb97df1e57f812ce94d4bc3baf7037421be7f74f295de040c52e0172
|
|
BLAKE2b-256 checksum How to use checksums |
25a772a0c419b498b9156f7cac28ab5e39ee2a344ce11f9f3c0ee81fddfc7e6f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|