Skip to main content
Yanked

This release has been yanked by its maintainers, and will be ignored by installers, except when explicitly specified.
Consider using release 0.6.2 instead.

Megatron Bridge

CICD NeMo Python 3.10+ GitHub Stars

Recipes | Examples | Contributing

Overview

Megatron Bridge is a PyTorch native library under NeMo Framework that leverages megatron-core to provide state-of-the-art training throughput for top models. It enables researchers and community developers to do both pre and post training using a performant and scalable training loop, with features like model parallelisms and mixed precisions (FP8, BF16, FP4 etc.). Megatron Bridge users can either leverage existing 🤗HuggingFace models or define their custom PyTorch model definitions for end-to-end workflows with flexibility.

🔧 Installation

🐳 NeMo-FW container

Best experience, highest performance and full feature support is guaranteed by the NeMo Framework container. Please fetch the most recent $TAG and run the following command to start a container:

docker run --rm -it -w /workdir -v $(pwd):/workdir \
  --entrypoint bash \
  --gpus all \
  nvcr.io/nvidia/nemo:${TAG}

📦 Bare metal install with TransformerEngine

TransformerEngine is a required dependency for Megatron Bridge. To install on bare metal (without any container), the following system requirements need to be fulfilled:

  • PyTorch >= 2.7
  • CUDA >= 12.8
  • cuDNN >= 9.3

We recommend installing the same versions that are present in the latest NGC PyTorch containers. The versions of these components for each container release can be found in the PyTorch and CUDA container release notes.

Please see these instructions for installing cuDNN for your target platform. You can check if CUDA toolkit and cuDNN are installed with:

dpkg -l | grep 'cuda-toolkit'
dpkg -l | grep 'cudnn.*cuda'

You can then run the following to install Megatron Bridge:

pip install torch setuptools pybind11 wheel_stub  # Required for TE
pip install --no-build-isolation megatron-bridge

uv

For installing Megatron Bridge with uv, please refer to our Contribution guide

⚡ Quickstart

To get started, first install Megatron Bridge or download a NeMo Framework container as described above.

Log in to HuggingFace Hub:

huggingface-cli login --token <your token>

You can then run the following to import a model from HuggingFace and start training with mock data:

from megatron.bridge import AutoBridge

import megatron.bridge.recipes.llama.llama32_1b as llama32_1b
from megatron.bridge.training.gpt_step import forward_step
from megatron.bridge.training.pretrain import pretrain

if __name__ == "__main__":
    # Load Llama from HuggingFace Hub and convert to Megatron
    bridge = AutoBridge.from_hf_pretrained("meta-llama/Llama-3.2-1B")
    model_provider = bridge.to_megatron_provider()

    # Get defaults for other configuration from an existing Llama 3.2 recipe
    cfg = llama32_1b.pretrain_config()
    cfg.model = model_provider
    cfg.train.train_iters = 10

    cfg.dataset.sequence_length = cfg.model.seq_length
    cfg.tokenizer.vocab_size = cfg.model.vocab_size

    pretrain(cfg, forward_step)

You can launch the above script with:

torchrun --nproc-per-node=<num devices> /path/to/script.py

🚀 Key Features

  • Bridge with 🤗Hugging Face: Seamless bidirectional conversion between 🤗Hugging Face and Megatron formats for interoperability (model bridges, auto bridge, conversion examples)
  • Flexible to Customize: Lightweight custom training loop making it easy to configure custom logic in data loading, distributed training, checkpointing, evaluation and logging (training framework, training utilities)
  • Supervised & Parameter-Efficient Finetuning: SFT & PEFT implementation tailored for Megatron-based models that supports LoRA, DoRA, and user-defined PEFT methods (PEFT implementations, finetune module, SFT dataset)
  • SoTA Training Recipes: Pre-configured production-ready training recipes for popular models like Llama 3, with optimized hyperparameters and distributed training configuration (Llama recipes, recipe examples)
  • Performance Optimization: Built-in support for FP8 training, model parallelisms, and memory-efficient techniques to offer high utilization and near linear scalability to thousands of nodes. (mixed precision, communication overlap, optimizer utilities)

Supported Models

Megatron Bridge provides out-of-the-box recipes for a wide range of models, built on top of base model architectures from megatron-core:

Large Language Models

Model Style Sizes Pretrain SFT & LoRA
Llama 3 GPT 8b, 70b ✅ APIs available, recipes upcoming
Llama 3.1 GPT 8b, 70b, 405b ✅ APIs available, recipes upcoming
Llama 3.2 GPT 1b, 3b ✅ APIs available, recipes upcoming

Launching Recipes

All recipes are ready to train out of the box, using mock data by default. For an example of how to override the default configuration through YAML or Hydra-style CLI overrides, please have a look at this script. The script can then be launched with torchrun. For example, with the aforementioned script:

torchrun --nproc-per-node=2 pretrain_llama3_8b.py model.tensor_model_parallel_size=1 <additional overrides ...>

Optionally, Megatron Bridge also supports launching with NeMo-Run. See the following examples for reference on launching with NeMo-Run:

These examples can also be run as is with the Llama 3 8b recipe (with NeMo-Run installed).

Launch Llama 3 8b Pretraining with NeMo-Run's run.Script:

uv run python pretrain_llama3_8b_nemo_run_script.py \
    --nproc-per-node=2 \
    model.pipeline_model_parallel_size=1 \
    train.train_iters=10 # this script passes Hydra-style overrides to the target script

Launch Llama 3 8b Pretraining with NeMo-Run's run.Partial

uv run python pretrain_llama3_8b_nemo_run_partial.py \
    --nproc-per-node=2

Performance Benchmarks

Coming soon ...

Project Structure

Megatron-Bridge/
├── examples/
│   ├── models/                  # Bridge usage examples
│   └── recipes/                 # Training examples
├── src/megatron/bridge/
│   ├── data/                    # Dataloaders and iterators
│   ├── models/                  # HuggingFace bridge infrastructure and model-specific implementations
│   │   ├── llama/               # Llama model providers
│   │   └── .../                 # Other models (gpt, t5, etc.)
│   ├── peft/                    # PEFT transformations and wrappers
│   ├── recipes/                 # Complete training recipes
│   ├── training/                # Training loop components
│   │   ├── tokenizers/          # Tokenizer library
│   │   └── utils/               # Training-specific utilities
│   └── utils/                   # Generic utilities for repo-wide usage
└── tests/                       # Comprehensive test suite

Contributing

We welcome community contributions! Please see our Contributor Guidelines for more information on how to get involved.

Metadata

Release files for megatron-bridge 0.1.6290rc0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for megatron-bridge 0.1.6290rc0
File Size Uploaded
megatron_bridge-0.1.6290rc0.tar.gz 310.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for megatron-bridge 0.1.6290rc0
File Interpreter ABI Platform
megatron_bridge-0.1.6290rc0-py3-none-any.whl Python 3 none any Details

Total release size: 760.9 kB

Release files / megatron_bridge-0.1.6290rc0.tar.gz

Download URL megatron_bridge-0.1.6290rc0.tar.gz
Size 310.6 kB
Tags Source
SHA-256 checksum
How to use checksums
62553d3eb60c0bc8e00e16ee9aed081171a1880a87014c722fa8938c7d7e88ff
BLAKE2b-256 checksum
How to use checksums
44a0f7eea26dcb425577c43c8f725ef460c14eb0666e3059e61e21ed099755c4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.12.3

Release files / megatron_bridge-0.1.6290rc0-py3-none-any.whl

Download URL megatron_bridge-0.1.6290rc0-py3-none-any.whl
Size 450.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
69fb2d6dec3ec875e85fb4a4f902fd2acc675a8298dd4ccceee06d4159528186
BLAKE2b-256 checksum
How to use checksums
b80a888cf2ef26e500c8f4e6fe44c36165782fc5cece8181fa0f43f386c8ef2e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.12.3
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page