CatTSunami: Accelerating Transition State Energy Calculations with Pre-trained Graph Neural Networks
CatTSunami is a framework for high-throughput enumeration of nudged elastic band (NEB) frame sets. It was built for use with machine learned (ML) models trained on OC20, which were demonstrated to be performant on this auxiliary task. To train your own model or obtain pre-trained checkpoints, please see fairchem-core.
This repository contains the validation dataset, framework for enumeration, and accompanying code to run ML-accelerated NEBs and validate new models. For more information, please read the manuscript paper.
Getting started
Configured for use:
- Install fairchem-core and fairchem-data-oc instructions
- Pip innstall fairchem-applications-cattsunami
- Check out the tutorial notebook
pip install fairchem-applications-cattsunami
Configured for local development:
- Clone the fairchem repo
- Install
fairchem-data-ocandfairchem-core: instructions - Install this repository
pip install -e packages/fairchem-applications-cattsunami - Check out the tutorial notebook
Validation Dataset
The validation dataset is comprised of 932 converged DFT NEB calculations to assess model performance on this important task. There are 3 different reaction classes considered: desorptions, dissociations, and transfers. There were 2827 total DFT NEBS performed including those that failed to converge. Unconverged systems have also been included in ASE All Trajectories below. For more information about the converged dataset see the dataset markdown file.
| Splits | Size of compressed version (in bytes) | Size of uncompressed version (in bytes) | MD5 checksum (download link) |
|---|---|---|---|
| ASE Converged Trajectories | 1.5G | 6.3G | 52af34a93758c82fae951e52af445089 |
| ASE All Trajectories | 6.7G | 30G | f5829eeaf7219c5cd3cfb499b8d951da |
Citing this work
If you use this codebase in your work, please consider citing:
@article{wander2024cattsunami,
title={CatTSunami: Accelerating Transition State Energy Calculations with Pre-trained Graph Neural Networks},
author={Wander, Brook and Shuaibi, Muhammed and Kitchin, John R and Ulissi, Zachary W and Zitnick, C Lawrence},
journal={arXiv preprint arXiv:2405.02078},
year={2024}
}
File Structure and Contents
The tar file contains 3 subdirectories: dissociations, desorptions, and transfers. As the names imply, these directories contain the converged DFT trajectories for each of the reaction classes. Within these directories, the trajectories are named to identify the contents of the file. Here is an example and the anatomy of the name:
desorption_id_83_2409_9_111-4_neb1.0.traj
desorptionindicates the reaction type (dissociation and transfer are the other possibilities)ididentifies that the material belongs to the validation in domain split (ood - out of domain is th e other possibility)83is the task id. This does not provide relavent information2409is the bulk index of the bulk used in the ocdata bulk pickle file9is the reaction index. for each reaction type there is a reaction pickle file in the repository. In this case it is the 9th entry to that pickle file111-4the first 3 numbers are the miller indices (i.e. the (1,1,1) surface), and the last number cooresponds to the shift value. In this case the 4th shift enumerated was the one used.neb1.0the number here indicates the k value used. For the full dataset, 1.0 was used so this does not distiguish any of the trajectories from one another.
The content of these trajectory files is the repeating frame sets. Despite the initial and final frames not being optimized during the NEB, the initial and final frames are saved for every iteration in the trajectory. For the dataset, 10 frames were used - 8 which were optimized over the neb. So the length of the trajectory is the number of iterations (N) * 10. If you wanted to look at the frame set prior to optimization and the optimized frame set, you could get them like this:
from ase.io import read
traj = read("desorption_id_83_2409_9_111-4_neb1.0.traj", ":")
unrelaxed_frames = traj[0:10]
relaxed_frames = traj[-10:]
Use
One more note: We have not prepared an lmdb for this dataset. This is because it is NEB calculations are not supported directly in ocp. You must use the ase native OCP class along with ase infrastructure to run NEB calculations. Here is an example of a use:
from ase.io import read
from ase.optimize import BFGS
from fairchem.core import pretrained_mlip, FAIRChemCalculator
from ase.mep import DyNEB
traj = read("desorption_id_83_2409_9_111-4_neb1.0.traj", ":")
images = traj[0:10]
predictor = pretrained_mlip.get_predict_unit("uma-s-1")
neb = DyNEB(images, k=1)
for image in images:
image.calc = FAIRChemCalculator(predictor, task_name="oc20")
optimizer = BFGS(
neb,
trajectory=f"test_neb.traj",
)
conv = optimizer.run(fmax=0.45, steps=200)
if conv:
neb.climb = True
conv = optimizer.run(fmax=0.05, steps=300)
Release files for fairchem-applications-cattsunami 1.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fairchem_applications_cattsunami-1.1.1.tar.gz | 920.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| fairchem_applications_cattsunami-1.1.1-py2.py3-none-any.whl | Python 2, Python 3 | none | any | Details |
Total release size: 1.8 MB
Release files / fairchem_applications_cattsunami-1.1.1.tar.gz
| Download URL | fairchem_applications_cattsunami-1.1.1.tar.gz |
|---|---|
| Size | 920.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c8cd422b756b3c53b7002fbec7d4b7419fa0aef0a097451596022bb3cfa45e1c
|
|
BLAKE2b-256 checksum How to use checksums |
ecc4b8495d2b09344ecbff78a658b5d7173d004218382e1fc2937d04a25ecfcd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 22, 2025.
Transparency logRelease files / fairchem_applications_cattsunami-1.1.1-py2.py3-none-any.whl
| Download URL | fairchem_applications_cattsunami-1.1.1-py2.py3-none-any.whl |
|---|---|
| Size | 922.5 kB |
| Tags | Python 2 Python 3 |
|
SHA-256 checksum How to use checksums |
a7dd96400cde9a8a65b27de5e990fbc1ff888ff814b4bfd470659c27bb2c1a54
|
|
BLAKE2b-256 checksum How to use checksums |
4a4cab290263e3854a36236496edddd30e811bb4c86713ecc0f6537d6fde8371
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 22, 2025.
Transparency log