Skip to main content

Matra-Genoa

arXiv

Matra-Genoa is a generative material transformer for the efficient generation of novel, symmetry-aware crystal structures. It utilizes an invertible tokenized representation of symmetrized crystals, including free coordinates. It can be conditioned on stability (energy above the convex hull), elemental compositions, space group and Wyckoff positions.

This repo contains the source code to (conditionally) generate crystal structures building on PyTorch, Pymatgen and more.


Resources


Citation

If you use this model or code in your research, please cite:

Pierre-Paul De Breuck, Hashim A. Piracha, Gian-Marco Rignanese, Miguel A. L. Marques
A generative material transformer using Wyckoff representation
arXiv:2501.16051 (2025) – https://arxiv.org/abs/2501.16051

De Breuck, P.-P., Piracha, H.A., Rignanese, G.-M. and Marques, M.A.L. A generative material transformer using Wyckoff representation. npj Comput Mater 12, 60 (2026). https://doi.org/10.1038/s41524-025-01940-8

Quickstart (Recommended)

Get started quickly by setting up a virtual environment and installing the package:

# Create and activate a virtual environment (uv recommended)
uv venv
source .venv/bin/activate

# Install the package
uv pip install matra-genoa

Once installed, you can start generating crystal structures immediately:

from matra_genoa import MatraGenoa

# Initialize the model (downloads default checkpoints automatically)
model = MatraGenoa()

# Generate tokens for 4 structures
sequences = model.generate(n=4, T=0.75, batch_size=16)

# Reconstruct a generated sequence into a pymatgen structure
from matra_genoa.utils import reconstruct
structure = reconstruct(sequences[0])
print(structure)

[!IMPORTANT] Generated crystal structures should be considered "as-is" from the transformer. It is highly recommended to relax these structures using a cheap universal Machine Learning Interatomic Potential (uMLIP) after reconstruction.

GPU Acceleration

Generating crystal structures is significantly faster on a GPU. You can move the model to your preferred device:

import torch

# Check for CUDA (NVIDIA) or MPS (Apple Silicon)
if torch.cuda.is_available():
    device = torch.device("cuda")
elif torch.backends.mps.is_available():
    device = torch.device("mps")
else:
    device = torch.device("cpu")

print(f"Using device: {device}")
model.to(device)

[!TIP] When using a GPU, you can increase the batch_size in model.generate(..., batch_size=256) to improve throughput.

[!TIP] Sampling Temperature (T): Use T ≈ 0.7 for structures close to the training distribution (generally more stable). For greater diversity, you can increase this up to T ≈ 2.0, though higher temperatures are more likely to produce invalid or "broken" crystal structures.

Conditioned Generation

You can guide the generation process by providing a starting prompt (conditioning). The model uses specific syntax blocks followed by a STOP token to define constraints.

Available Syntax Blocks:

  • AMT [value]: Number of distinct chemical elements (e.g., AMT 2 for binaries).
  • EHULL_DISC [EH0/EH1]: Discrete stability target. EH0 indicates stability (distance to convex hull < 75 meV/atom).
  • ELMS [elements]: Target chemical elements.
  • STOICH [ratios]: Target stoichiometry.
  • SPACEGROUP S[1-230]: Target space group number.
  • WYCKOFF W[idx]: Target Wyckoff positions.
  • EHULL [value]: Continuous stability target.

Example:

# Target stable binary compounds containing Na and Cl
condition = "EHULL_DISC EH0 STOP AMT 2 STOP ELMS Na Cl STOP"
tokens = model.generate(n=100, condition=condition)

Batch Generation and Decoding

For generating many structures at once, generate_structures() samples and reconstructs sequences into pymatgen Structure objects in one call, keeping only sequences that reconstruct successfully:

structures = model.generate_structures(n=1000, T=0.8, decode_jobs=8, progress=True)

decode_jobs sets how many CPU workers reconstruct structures in parallel. Decoding is always CPU-bound, so for best throughput put the model on a GPU (see GPU Acceleration) for generation while decode_jobs handles decoding on CPU. You can also decode a sequence list you already have with model.decode(sequences, n_jobs=8). See 03_batch_generation_and_decoding.ipynb for more detail.

Tutorials

For more in-depth examples and advanced usage, explore the interactive notebooks:


Detailed Installation

From PyPI

pip install matra-genoa

From GitHub

You can also install the package directly from the source repository:

pip install git+https://github.com/ppdebreuck/matra-genoa.git

To include optional dependencies for development or extra materials helpers:

# For tutorials and development
pip install "matra-genoa[dev,tutorials] @ git+https://github.com/ppdebreuck/matra-genoa.git"

# For additional materials science helpers (smact, etc.)
pip install "matra-genoa[materials] @ git+https://github.com/ppdebreuck/matra-genoa.git"

Local Development

If you have the repository cloned locally, install it in editable mode:

pip install -e ".[dev,tutorials]"

(To generate a source distribution locally for hosting, run python -m build which will create a dist/ directory.)

Checkpoints

The following models are currently available:

  • Matra-Genoa-MPAS (Default)
  • Matra-Genoa-MP

Checkpoints are automatically downloaded to your local cache:

  • $MATRA_CACHE_DIR if set
  • otherwise ~/.cache/matra (on Linux/macOS)

Package Layout

The repository follows a modern src/ layout:

  • src/matra_genoa/: Core Python package logic
  • tests/: Unit tests
  • tutorials/: Interactive Jupyter notebooks
  • examples/: Script-based examples
  • checkpoints/: Local checkpoint placeholders and cache logic

Metadata

Release files for matra-genoa 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for matra-genoa 0.1.0
File Size Uploaded
matra_genoa-0.1.0.tar.gz 70.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for matra-genoa 0.1.0
File Interpreter ABI Platform
matra_genoa-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 139.1 kB

Release files / matra_genoa-0.1.0.tar.gz

Download URL matra_genoa-0.1.0.tar.gz
Size 70.7 kB
Tags Source
SHA-256 checksum
How to use checksums
4627cc5ef1ac190bb64450bfcfc640f0e732322a50072eb7992122c47ac926a2
BLAKE2b-256 checksum
How to use checksums
7e9f6aef280fb8760c1dc107131749cc7ecaecbba788f25ef18613a56129030b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.10

Release files / matra_genoa-0.1.0-py3-none-any.whl

Download URL matra_genoa-0.1.0-py3-none-any.whl
Size 68.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
17e4e9414b9c45dfa45eb06b7c20977149273cbd688bb51ef7dbc6c2e9c3f5c4
BLAKE2b-256 checksum
How to use checksums
2a0fe7f5eae8e75a547dae161d2b7c06868972a86c1683941099d82ada0df241
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.10

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page