Matra-Genoa
Matra-Genoa is a generative material transformer for the efficient generation of novel, symmetry-aware crystal structures. It utilizes an invertible tokenized representation of symmetrized crystals, including free coordinates. It can be conditioned on stability (energy above the convex hull), elemental compositions, space group and Wyckoff positions.
This repo contains the source code to (conditionally) generate crystal structures building on PyTorch, Pymatgen and more.
Resources
- 📄 Paper: arXiv:2501.16051
- 📊 MatraGenoa3M: 3 million generated crystals on figshare. Not relaxed; future updates may include novel relaxed structures.
- 🖥️ Matra-Genoa-MPAS Front End: matra.pierrepauldb.com
Citation
If you use this model or code in your research, please cite:
Pierre-Paul De Breuck, Hashim A. Piracha, Gian-Marco Rignanese, Miguel A. L. Marques
A generative material transformer using Wyckoff representation
arXiv:2501.16051 (2025) – https://arxiv.org/abs/2501.16051
De Breuck, P.-P., Piracha, H.A., Rignanese, G.-M. and Marques, M.A.L. A generative material transformer using Wyckoff representation. npj Comput Mater 12, 60 (2026). https://doi.org/10.1038/s41524-025-01940-8
Quickstart (Recommended)
Get started quickly by setting up a virtual environment and installing the package:
# Create and activate a virtual environment (uv recommended)
uv venv
source .venv/bin/activate
# Install the package
uv pip install matra-genoa
Once installed, you can start generating crystal structures immediately:
from matra_genoa import MatraGenoa
# Initialize the model (downloads default checkpoints automatically)
model = MatraGenoa()
# Generate tokens for 4 structures
sequences = model.generate(n=4, T=0.75, batch_size=16)
# Reconstruct a generated sequence into a pymatgen structure
from matra_genoa.utils import reconstruct
structure = reconstruct(sequences[0])
print(structure)
[!IMPORTANT] Generated crystal structures should be considered "as-is" from the transformer. It is highly recommended to relax these structures using a cheap universal Machine Learning Interatomic Potential (uMLIP) after reconstruction.
GPU Acceleration
Generating crystal structures is significantly faster on a GPU. You can move the model to your preferred device:
import torch
# Check for CUDA (NVIDIA) or MPS (Apple Silicon)
if torch.cuda.is_available():
device = torch.device("cuda")
elif torch.backends.mps.is_available():
device = torch.device("mps")
else:
device = torch.device("cpu")
print(f"Using device: {device}")
model.to(device)
[!TIP] When using a GPU, you can increase the
batch_sizeinmodel.generate(..., batch_size=256)to improve throughput.
[!TIP] Sampling Temperature (
T): UseT ≈ 0.7for structures close to the training distribution (generally more stable). For greater diversity, you can increase this up toT ≈ 2.0, though higher temperatures are more likely to produce invalid or "broken" crystal structures.
Conditioned Generation
You can guide the generation process by providing a starting prompt (conditioning). The model uses specific syntax blocks followed by a STOP token to define constraints.
Available Syntax Blocks:
AMT [value]: Number of distinct chemical elements (e.g.,AMT 2for binaries).EHULL_DISC [EH0/EH1]: Discrete stability target.EH0indicates stability (distance to convex hull < 75 meV/atom).ELMS [elements]: Target chemical elements.STOICH [ratios]: Target stoichiometry.SPACEGROUP S[1-230]: Target space group number.WYCKOFF W[idx]: Target Wyckoff positions.EHULL [value]: Continuous stability target.
Example:
# Target stable binary compounds containing Na and Cl
condition = "EHULL_DISC EH0 STOP AMT 2 STOP ELMS Na Cl STOP"
tokens = model.generate(n=100, condition=condition)
Batch Generation and Decoding
For generating many structures at once, generate_structures() samples and reconstructs sequences into pymatgen Structure objects in one call, keeping only sequences that reconstruct successfully:
structures = model.generate_structures(n=1000, T=0.8, decode_jobs=8, progress=True)
decode_jobs sets how many CPU workers reconstruct structures in parallel. Decoding is always CPU-bound, so for best throughput put the model on a GPU (see GPU Acceleration) for generation while decode_jobs handles decoding on CPU. You can also decode a sequence list you already have with model.decode(sequences, n_jobs=8). See 03_batch_generation_and_decoding.ipynb for more detail.
Tutorials
For more in-depth examples and advanced usage, explore the interactive notebooks:
01_load_model_and_generate.ipynb: Basics of loading, sampling, and structure reconstruction.02_conditioned_generation_and_reconstruction.ipynb: Guided generation with conditioning prompts.03_batch_generation_and_decoding.ipynb: CPU-parallel decoding of large batches withgenerate_structures()/decode().examples/quickstart.py: A ready-to-run Python script.
Detailed Installation
From PyPI
pip install matra-genoa
From GitHub
You can also install the package directly from the source repository:
pip install git+https://github.com/ppdebreuck/matra-genoa.git
To include optional dependencies for development or extra materials helpers:
# For tutorials and development
pip install "matra-genoa[dev,tutorials] @ git+https://github.com/ppdebreuck/matra-genoa.git"
# For additional materials science helpers (smact, etc.)
pip install "matra-genoa[materials] @ git+https://github.com/ppdebreuck/matra-genoa.git"
Local Development
If you have the repository cloned locally, install it in editable mode:
pip install -e ".[dev,tutorials]"
(To generate a source distribution locally for hosting, run python -m build which will create a dist/ directory.)
Checkpoints
The following models are currently available:
Matra-Genoa-MPAS(Default)Matra-Genoa-MP
Checkpoints are automatically downloaded to your local cache:
$MATRA_CACHE_DIRif set- otherwise
~/.cache/matra(on Linux/macOS)
Package Layout
The repository follows a modern src/ layout:
src/matra_genoa/: Core Python package logictests/: Unit teststutorials/: Interactive Jupyter notebooksexamples/: Script-based examplescheckpoints/: Local checkpoint placeholders and cache logic
Metadata
Release files for matra-genoa 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| matra_genoa-0.1.0.tar.gz | 70.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| matra_genoa-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 139.1 kB
Release files / matra_genoa-0.1.0.tar.gz
| Download URL | matra_genoa-0.1.0.tar.gz |
|---|---|
| Size | 70.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4627cc5ef1ac190bb64450bfcfc640f0e732322a50072eb7992122c47ac926a2
|
|
BLAKE2b-256 checksum How to use checksums |
7e9f6aef280fb8760c1dc107131749cc7ecaecbba788f25ef18613a56129030b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.10
|
Release files / matra_genoa-0.1.0-py3-none-any.whl
| Download URL | matra_genoa-0.1.0-py3-none-any.whl |
|---|---|
| Size | 68.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
17e4e9414b9c45dfa45eb06b7c20977149273cbd688bb51ef7dbc6c2e9c3f5c4
|
|
BLAKE2b-256 checksum How to use checksums |
2a0fe7f5eae8e75a547dae161d2b7c06868972a86c1683941099d82ada0df241
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.10
|