IsoGen
IsoGen is a toolbox for predicting isotope distributions from protein, RNA, DNA, neutral-mass, and elemental-formula inputs.
It includes both an absolute FFT-based calculation and a neural network prediction.
Pretrained models are included for both peptides and RNA based on either average mass or sequence. DNA prediction uses the RNA model due to the similarity of their elemental compositions.
The FFT methods are absolute and are limited only by the accuracy of the data you put in. They are a little faster, especially on larger species.
The NN methods are very accurate and can be faster on smaller species. The primary advantage of these is that they can be retrained on non-standard isotope distributions.
Installation
Install a published wheel from PyPI:
python -m pip install pyisogen
IsoGen requires Python 3.13 or newer.
Precompiled native libraries are provided for 64-bit Windows and Linux. Linux requires the FFTW 3 runtime; published Linux wheels bundle it during the manylinux repair step. For other platforms, build the native library from source using CMake.
Usage
From Python:
import isogen
protein = isogen.isodist("ACDEFGHIK", type="PEPTIDE", isolen=64)
rna = isogen.isodist("AUGCAGUACGUA", type="RNA", isolen=64)
dna = isogen.isodist("ATGCAGTACGTA", type="DNA", isolen=64)
glucose_mass_dist = isogen.isodist("C6H12O6", type="ATOM", isolen=32)
The output is a numpy array of shape (isolen, 2) with the first column containing the monoisotopic mass and the second column containing the relative intensity. The isolen parameter controls the number of isotopic peaks returned.
In addition to FFT methods, IsoGen provides neural-network models for peptides and RNA. The default is to use the exact "FFT" model with averagines to calculate the isotope distribution from an average protein, RNA, or DNA mass. Turning on "NN" mode uses the neural-network model to predict the isotope distribution from a peptide or RNA sequence.
The PEPTIDE model is trained on peptide sequences, while the RNA model is trained on RNA sequences. The DNA type uses the RNA model, and
The public ATOM type uses the FFT method; no neural-network formula model is
available.
Custom neural-network models
Use isodist_custom to generate a distribution from a binary model file rather
than one of IsoGen's bundled neural-network models:
from pathlib import Path
import isogen
model_file = Path("models/my_peptide_model_64.bin")
custom = isogen.isodist_custom(
"ACDEFGHIK",
model_file=model_file,
isolen=64,
type="PEPTIDE",
)
The function accepts peptide, RNA, and DNA sequences or numeric neutral masses.
It always uses the neural-network method. The model must have the correct input
size for the selected input and type, and its output size must equal isolen.
Peptide sequence models have 20 inputs, RNA/DNA sequence models have 4 inputs,
and neutral-mass models have 5 inputs. Invalid, unreadable, or incompatible
model files raise ValueError. As with isodist, the result has shape
(isolen, 2), containing neutral masses and relative intensities.
Training custom models
Install the training dependencies before importing the training modules:
python -m pip install -e ".[training]"
Training data is stored in NumPy .npz archives. Sequence models expect a
seqs array and mass models expect a masses array. Every archive also needs
a dists array with shape (number_of_examples, isolen). Each row of dists
is the target relative-intensity distribution for its corresponding sequence
or neutral mass. For example:
import numpy as np
np.savez_compressed(
"peptide_training.npz",
seqs=np.asarray(["ACDE", "PEPTIDE", "MARTY"]),
dists=np.asarray(peptide_target_distributions, dtype=np.float32),
)
np.savez_compressed(
"mass_training.npz",
masses=np.asarray([1_000.0, 5_000.0, 10_000.0]),
dists=np.asarray(mass_target_distributions, dtype=np.float32),
)
Use the engine matching the kind of input the model will receive. The helper below directs generated models to a separate directory instead of overwriting the models installed with IsoGen:
from pathlib import Path
from isogen.isogenmass import IsoGenMassEngine
from isogen.isogenpep import IsoGenPepEngine
from isogen.isogenrna import IsoGenRNAEngine
from isogen.isogenrna_averagine import IsoGenRNAveragineEngine
model_dir = Path("trained_models")
model_dir.mkdir(exist_ok=True)
def set_model_directory(engine):
"""Set the output directory before a model is initialized or loaded."""
engine.model.working_dir = str(model_dir)
for model in engine.models:
model.working_dir = str(model_dir)
# Peptide sequences: 20-element amino-acid composition input.
pep = IsoGenPepEngine(isolen=64)
set_model_directory(pep)
pep.train("peptide_training.npz", epochs=20, forcenew=True)
# RNA sequences: 4-element A/C/G/U composition input. This model is also
# used for DNA inference after IsoGen converts thymine to uracil.
rna = IsoGenRNAEngine(isolen=64)
set_model_directory(rna)
rna.train("rna_training.npz", epochs=20, forcenew=True)
# Peptide-like neutral masses: 5-element mass encoding.
mass = IsoGenMassEngine(isolen=64)
set_model_directory(mass)
mass.train_multiple(
["mass_training.npz"],
inputname="masses",
epochs=20,
forcenew=True,
)
# RNA-like neutral masses: 5-element mass encoding.
rna_mass = IsoGenRNAveragineEngine(isolen=64)
set_model_directory(rna_mass)
rna_mass.train_multiple(
["rna_mass_training.npz"],
inputname="masses",
epochs=20,
forcenew=True,
)
IsoGenPepEngine supports output lengths 16, 64, and 128;
IsoGenRNAEngine supports 64 and 128; IsoGenMassEngine models intended for
isodist_custom support 8, 32, 64, and 128; and
IsoGenRNAveragineEngine supports 32, 64, and 128. The output length used to
construct the engine must match the width of dists and the isolen passed to
isodist_custom.
After training, each engine saves a PyTorch .pth checkpoint and a raw .bin
model in trained_models. The .pth file is used to resume Python training;
pass the .bin file to isodist_custom. The generated filenames are
isogenpep_model_<isolen>.bin, isogenrna_model_<isolen>.bin,
isogenmass_model_<isolen>.bin, and
isogen_rnaveragine_model<isolen>.bin, respectively:
custom = isogen.isodist_custom(
"ACDEFGHIK",
model_file=model_dir / "isogenpep_model_64.bin",
isolen=64,
type="PEPTIDE",
)
Passing forcenew=True starts from newly initialized weights. Use
forcenew=False to resume from a matching .pth checkpoint in the configured
model directory. IsoGenMassEngine.train(...) and
IsoGenRNAveragineEngine.train(...) can also generate standard FFT targets
from random masses when a custom target archive is not needed.
Peptide ions and RNA termini
For peptide fragments, pass the fragment sequence and select its neutral
terminal composition with ion_type. IsoGen supports intact H2O (the
default) and the peptide a, b, c, x, y, and z ion types:
b6 = isogen.isodist("PEPTID", type="PEPTIDE", ion_type="b")
y6 = isogen.isodist("EPTIDE", type="PEPTIDE", ion_type="y")
Supply the N-terminal subsequence for a/b/c ions and the C-terminal subsequence for x/y/z ions. Returned values are neutral masses, not charge-adjusted m/z.
RNA does not currently accept named RNA fragment-ion series through
ion_type. For an intact or manually truncated RNA sequence, configure the
supported terminal chemistry with threeend and fiveend:
rna_5_triphosphate = isogen.isodist(
"AUGC",
type="RNA",
threeend="OH",
fiveend="TP",
)
The available 5' settings are hydroxyl (OH), monophosphate (MP, default),
and triphosphate (TP); the supported explicit 3' setting is hydroxyl (OH,
default). These peptide-ion and RNA-terminal options adjust the mass-axis
origin. The sequence-model intensity vector retains its standard terminal
composition.
From the command line:
isogen dist ACDEFGHIK --type PEPTIDE --isolen 64
isogen dist C6H12O6 --type ATOM --isolen 32
isogen plot
See python -m isogen --help for all options.
The source repository also builds a native development executable named
isogen_test.exe on Windows (isogen_test on Linux). It can be run from the
repository's bin directory with isogen_test.exe -mass 10000, but it is not
installed by the Python wheel. Use the isogen console command for installed
packages.
Documentation
Read the full IsoGen documentation.
The documentation sources are also available in the repository's docs
directory. To preview them locally:
python -m pip install -e ".[docs]"
python -m mkdocs serve
Tests
The test suite uses Pyteomics as an independent mass reference. Pyteomics is only part of the optional test dependencies and is not installed with IsoGen:
python -m pip install -e ".[test]"
python -m pytest
Development and model-training modules have additional dependencies:
python -m pip install -e ".[training]"
License
IsoGen is released under the BSD 3-Clause License. See LICENSE for details.
PLEASE CITE THIS SOFTWARE IN ANY PUBLICATIONS THAT USE IT (publication to follow).
Contact
If you have any questions, please email mtmarty@utexas.edu or open a ticket on GitHub.
CHANGELOG
1.0.2
Added support for custom models with isogen_custom function and new C bindings for custom models.
1.0.1
Small updates to README.md
1.0.0
Initial release. Rewrote significantly from UniDec build using AI tool to improve the release and add in atomic formula support.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pyisogen-1.0.2.tar.gz.
File metadata
- Download URL: pyisogen-1.0.2.tar.gz
- Upload date:
- Size: 24.7 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e50a3ab435d37544ddc494ad112ec7c824e1ffd466b97a2c350a91b12579f552
|
|
| MD5 |
697e196c3e68eac628f5d5d954fe9c7e
|
|
| BLAKE2b-256 |
cf13cb02709e1a14260d5fe793252a791e5ca529d4ffcb298cdc236311453057
|
Provenance
The following attestation bundles were made for pyisogen-1.0.2.tar.gz:
Publisher:
publish.yml on michaelmarty/IsoGen
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pyisogen-1.0.2.tar.gz -
Subject digest:
e50a3ab435d37544ddc494ad112ec7c824e1ffd466b97a2c350a91b12579f552 - Sigstore transparency entry: 2302682475
- Sigstore integration time:
-
Permalink:
michaelmarty/IsoGen@3b24bbfdc9fd6b2c92b440bb013ae73cac6f14f0 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/michaelmarty
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@3b24bbfdc9fd6b2c92b440bb013ae73cac6f14f0 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file pyisogen-1.0.2-py3-none-win_amd64.whl.
File metadata
- Download URL: pyisogen-1.0.2-py3-none-win_amd64.whl
- Upload date:
- Size: 22.2 MB
- Tags: Python 3, Windows x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0d9ed9e98b4bd385a3ae08b9136725b7ccd6a5cf2c5dc69f0883a71ce66b9f0d
|
|
| MD5 |
37fe4b0bbce0ceef0042f13c110e9662
|
|
| BLAKE2b-256 |
11c16fe3f03f8fc13551bf6fb878c3871ef567db113f91ea5dfa0464e39258a1
|
Provenance
The following attestation bundles were made for pyisogen-1.0.2-py3-none-win_amd64.whl:
Publisher:
publish.yml on michaelmarty/IsoGen
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pyisogen-1.0.2-py3-none-win_amd64.whl -
Subject digest:
0d9ed9e98b4bd385a3ae08b9136725b7ccd6a5cf2c5dc69f0883a71ce66b9f0d - Sigstore transparency entry: 2302682530
- Sigstore integration time:
-
Permalink:
michaelmarty/IsoGen@3b24bbfdc9fd6b2c92b440bb013ae73cac6f14f0 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/michaelmarty
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@3b24bbfdc9fd6b2c92b440bb013ae73cac6f14f0 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file pyisogen-1.0.2-py3-none-manylinux_2_31_x86_64.whl.
File metadata
- Download URL: pyisogen-1.0.2-py3-none-manylinux_2_31_x86_64.whl
- Upload date:
- Size: 13.1 MB
- Tags: Python 3, manylinux: glibc 2.31+ x86-64
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
40148db687341134323341cad254e7fdc370ee70c4952de38435ba308fb3707d
|
|
| MD5 |
5144e8c98c1579ceb706b06918079389
|
|
| BLAKE2b-256 |
e6735cd3247354923c26a4155a3d38c0388184edffdd01300f917b96e0442a4b
|
Provenance
The following attestation bundles were made for pyisogen-1.0.2-py3-none-manylinux_2_31_x86_64.whl:
Publisher:
publish.yml on michaelmarty/IsoGen
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pyisogen-1.0.2-py3-none-manylinux_2_31_x86_64.whl -
Subject digest:
40148db687341134323341cad254e7fdc370ee70c4952de38435ba308fb3707d - Sigstore transparency entry: 2302682566
- Sigstore integration time:
-
Permalink:
michaelmarty/IsoGen@3b24bbfdc9fd6b2c92b440bb013ae73cac6f14f0 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/michaelmarty
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@3b24bbfdc9fd6b2c92b440bb013ae73cac6f14f0 -
Trigger Event:
workflow_dispatch
-
Statement type: