Skip to main content

GenBio AI

AIDO.ModelGenerator

License python PyPI version Maintenance Downloads DOI

AIDO.ModelGenerator is a software stack powering the development of an AI-driven Digital Organism by enabling researchers to adapt pretrained models and generate finetuned models for downstream tasks. To read more about AIDO.ModelGenerator's integral role in building the world's first AI-driven Digital Organism, see AIDO.

AIDO.ModelGenerator is open-sourced as an opinionated plug-and-play research framework for cross-disciplinary teams in ML & Bio. It is designed to enable rapid and reproducible prototyping with four kinds of experiments in mind:

  1. Applying pre-trained foundation models to new data
  2. Developing new finetuning and inference tasks for foundation models
  3. Benchmarking foundation models and creating leaderboards
  4. Testing new architectures for finetuning performance

while also scaling with hardware and integrating with larger data pipelines or research workflows.

AIDO.ModelGenerator is built on PyTorch, HuggingFace, and Lightning, and works seamlessly with these ecosystems.

See the AIDO.ModelGenerator documentation for installation, usage, tutorials, and API reference.

Who uses ModelGenerator?

🧬 Biologists

  • Intuitive one-command CLIs for in silico experiments
  • Pre-trained model zoo
  • Broad data compatibility
  • Pipeline-oriented workflows

🤖 ML Researchers

  • Reproducible-by-design experiments
  • Architecture A/B testing
  • Automatic hardware scaling
  • Integration with PyTorch, Lightning, HuggingFace, and WandB

☕ Software Engineers

  • Extensible and modular models, tasks, and data
  • Strict typing and documentation
  • Fail-fast interface design
  • Continuous integration and testing

🤝 Everyone benefits from

  • A collaborative hub and focal point for multidisciplinary work on experiments, models, software, and data
  • Community-driven development
  • Permissive license for academic and non-commercial use

Projects using AIDO.ModelGenerator

Installation

git clone https://github.com/genbio-ai/ModelGenerator.git
cd ModelGenerator
pip install -e .

Source installation is necessary to add new backbones, finetuning tasks, and data transformations, as well as use convenience configs and scripts. If you only need to run inference, reproduce published experiments, or finetune on new data, you can use

pip install modelgenerator
pip install git+https://github.com/genbio-ai/openfold.git@c4aa2fd0d920c06d3fd80b177284a22573528442
pip install git+https://github.com/NVIDIA/dllogger.git@0540a43971f4a8a16693a9de9de73c1072020769

Quick Start

Get embeddings from a pre-trained model

mgen predict --model Embed --model.backbone aido_dna_dummy \
  --data SequencesDataModule --data.path genbio-ai/100m-random-promoters \
  --data.x_col sequence --data.id_col sequence --data.test_split_size 0.0001 \
  --config configs/examples/save_predictions.yaml

Get token probabilities from a pre-trained model

mgen predict --model Inference --model.backbone aido_dna_dummy \
  --data SequencesDataModule --data.path genbio-ai/100m-random-promoters \
  --data.x_col sequence --data.id_col sequence --data.test_split_size 0.0001 \
  --config configs/examples/save_predictions.yaml

Finetune a model

mgen fit --model ConditionalDiffusion --model.backbone aido_dna_dummy \
  --data ConditionalDiffusionDataModule --data.path "genbio-ai/100m-random-promoters" \
  --data.x_col sequence --data.y_col label --data.rename_cols "{sequence: sequences}"

Evaluate a model checkpoint

mgen test --model ConditionalDiffusion --model.backbone aido_dna_dummy \
  --data ConditionalDiffusionDataModule --data.path "genbio-ai/100m-random-promoters" \
  --data.x_col sequence --data.y_col label --data.rename_cols "{sequence: sequences}" \
  --ckpt_path logs/lightning_logs/version_X/checkpoints/<your_model>.ckpt

Save predictions

mgen predict --model ConditionalDiffusion --model.backbone aido_dna_dummy \
  --data ConditionalDiffusionDataModule --data.path "genbio-ai/100m-random-promoters" \
  --data.x_col sequence --data.y_col label --data.rename_cols "{sequence: sequences}" \
  --ckpt_path logs/lightning_logs/version_X/checkpoints/<your_model>.ckpt \
  --config configs/examples/save_predictions.yaml

Configify your experiment

This command

mgen fit --model ConditionalDiffusion --model.backbone aido_dna_dummy \
  --data ConditionalDiffusionDataModule --data.path "genbio-ai/100m-random-promoters" \
  --data.x_col sequence --data.y_col label --data.rename_cols "{sequence: sequences}"

is equivalent to mgen fit --config my_config.yaml with

# my_config.yaml
model:
  class_path: ConditionalDiffusion
  init_args:
    backbone: aido_dna_dummy
data:
  class_path: ConditionalDiffusionDataModule
  init_args:
    path: "genbio-ai/100m-random-promoters"
    x_col: sequence
    y_col: label
    rename_cols:
      sequence: sequences

Use composable configs to customize workflows

mgen fit --model SequenceRegression --data PromoterExpressionRegression \
  --config configs/defaults.yaml \
  --config configs/examples/lora_backbone.yaml \
  --config configs/examples/wandb.yaml

We provide some useful examples in configs/examples. Configs use the LAST value for each attribute. Check the full configuration logged with each experiment in logs/lightning_logs/your-experiment/config.yaml, or if using wandb logs/config.yaml.

Use LoRA for parameter-efficient finetuning

This also avoids saving the full model, only the LoRA weights are saved.

mgen fit --data PromoterExpressionRegression \
  --model SequenceRegression --model.backbone.use_peft true \
  --model.backbone.lora_r 16 \
  --model.backbone.lora_alpha 32 \
  --model.backbone.lora_dropout 0.1

Use continued pretraining for finetuning domain adaptation

First run pretraining objective on finetuning data

# https://arxiv.org/pdf/2310.02980
mgen fit --model MLM --model.backbone aido_dna_dummy \
  --data MLMDataModule --data.path leannmlindsey/GUE \
  --data.x_col sequence --data.y_col label --data.rename_cols "{sequence: sequences}" \
  --data.config_name prom_core_notata

Then finetune using the adapted model

mgen fit --model SequenceClassification --model.strict_loading false \
  --data SequenceClassificationDataModule --data.path leannmlindsey/GUE \
  --data.x_col sequence --data.y_col label --data.rename_cols "{sequence: sequences}" \
  --data.config_name prom_core_notata \
  --ckpt_path logs/lightning_logs/version_X/checkpoints/<your_adapted_model>.ckpt

Make sure to turn off strict_loading to replace the adapter!

Use the head/adapter/decoder that comes with the backbone

mgen fit --model SequenceClassification --data GUEClassification \
  --model.use_legacy_adapter true

Metadata

Release files for modelgenerator 0.1.3.post0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for modelgenerator 0.1.3.post0
File Size Uploaded
modelgenerator-0.1.3.post0.tar.gz 7.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for modelgenerator 0.1.3.post0
File Interpreter ABI Platform
modelgenerator-0.1.3.post0-py3-none-any.whl Python 3 none any Details

Total release size: 10.4 MB

Release files / modelgenerator-0.1.3.post0.tar.gz

Download URL modelgenerator-0.1.3.post0.tar.gz
Size 7.5 MB
Tags Source
SHA-256 checksum
How to use checksums
b5e24f66278596c9f65c44d14d3d80390df72465a7159483054b3f464d2cbb19
BLAKE2b-256 checksum
How to use checksums
653c4f36684dad8b4cc5b3145dea2f05f25fd8a3ae29934f90e30f31f1ca377b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Dec 12, 2025.

Transparency log

Release files / modelgenerator-0.1.3.post0-py3-none-any.whl

Download URL modelgenerator-0.1.3.post0-py3-none-any.whl
Size 3.0 MB
Tags Python 3
SHA-256 checksum
How to use checksums
8000de0fa2259cc764ff10bef2563e6c9c8604a2f9dd503eb1e9c2a9dd9928b7
BLAKE2b-256 checksum
How to use checksums
dfdcffcc77c4f2e37e0fd68a2f5bd6251bebe727a3df525c4e691758f3b28c95
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Dec 12, 2025.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.3.post0 This release

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page