Skip to main content

OpenLanguageModel (OLM)

OpenLanguageModel is a PyTorch-native library for building, training, teaching, and researching transformer language models. It is designed for people who want the model architecture to stay visible while the training stack stays manageable.

OLM gives you:

  • readable transformer components in olm.nn
  • implemented model families in olm.models
  • local and Hugging Face dataset streams in olm.data
  • single-device, single-node multi-GPU DDP/FSDP, AMP, checkpointing, callbacks, and automatic trainer selection in olm.train

Website · Docs · Install · Colab Notebooks · API Reference · Examples · Issues

Why OLM

Most language-model libraries either hide the architecture behind configuration, or make you rebuild the whole training path from scratch. OLM sits in the middle: every block is an ordinary torch.nn.Module, but data loading, optimization, mixed precision, single-node multi-GPU training, checkpointing, and logging are already wired into a clean path.

That makes it useful for:

  • students learning how language models are assembled and trained
  • researchers running ablations on attention, norms, feed-forward layers, and residual structure
  • practitioners who want existing PyTorch workflows without a hidden runtime

Llama 3 Block In OLM

Model code in OLM is meant to read like the architecture it represents. For example, the Llama 3 block is built from RMSNorm, grouped-query attention, SwiGLU, and explicit residual structure:

from olm.nn.structure import Block
from olm.nn.structure.combinators import Residual
from olm.nn.attention import GroupedQueryAttention
from olm.nn.feedforward import SwiGLUFFN
from olm.nn.norms import RMSNorm


class Llama3Block(Block):
    def __init__(
        self,
        embed_dim: int,
        intermediate_size: int,
        num_heads: int,
        num_kv_heads: int,
        max_seq_len: int,
        dropout: float,
        rope_theta: float,
    ):
        super().__init__([
            Residual(Block([
                RMSNorm(embed_dim, eps=1e-5),
                GroupedQueryAttention(
                    embed_dim,
                    num_heads,
                    num_kv_heads,
                    max_seq_len,
                    dropout=dropout,
                    rope_theta=rope_theta,
                    use_bias=False,
                ),
            ])),
            Residual(Block([
                RMSNorm(embed_dim, eps=1e-5),
                SwiGLUFFN(
                    embed_dim,
                    hidden_dim=intermediate_size,
                    dropout=dropout,
                    bias=False,
                ),
            ])),
        ])

Source: src/olm/models/meta/llama3.py

Train With The Stack Connected

You can keep the model and optimizer as normal PyTorch objects while OLM handles the training loop details:

import torch

from olm.data.datasets import DataLoader, FineWebEduDataset
from olm.data.tokenization import HFTokenizer
from olm.models.openai import GPT2Model
from olm.train import AutoTrainer
from olm.train.optim import AdamW

tokenizer = HFTokenizer("gpt2")
dataset = FineWebEduDataset(tokenizer, context_length=1024, streaming=True)
loader = DataLoader(dataset, batch_size=8, num_workers=4)

model = GPT2Model(
    vocab_size=tokenizer.vocab_size,
    embed_dim=768,
    num_layers=12,
    num_heads=12,
    max_seq_len=1024,
)

trainer = AutoTrainer(
    model,
    AdamW,
    loader,
    device="auto",
    context_length=1024,
    learning_rate=3e-4,
    grad_accum_steps=8,
)
trainer.train(epochs=1, max_steps=1000)

AutoTrainer chooses between CPU, single-GPU, and single-node multi-GPU DDP/FSDP paths based on the hardware and model. You can still use Trainer, DDPTrainer, or FSDPTrainer directly when you want explicit control.

Implemented Model Families

OLM includes named presets and configurable base classes for common transformer families:

Family Source
GPT-2 src/olm/models/openai/gpt2.py
Llama 2 src/olm/models/meta/llama2.py
Llama 3 / 3.1 / 3.2 src/olm/models/meta/llama3.py
Qwen 2.5 src/olm/models/alibaba/qwen2.py
Phi-3 / Phi-3.5 src/olm/models/microsoft/phi3.py
Phi-4 src/olm/models/microsoft/phi4.py
Gemma 2 src/olm/models/google/gemma2.py
OLMo src/olm/models/allenai/olmo.py
OPT src/olm/models/facebook/opt.py

See docs/api.md for the generated API reference and examples/ for training scripts.

Installation

Use Python 3.10, 3.11, or 3.12.

git clone https://github.com/openlanguagemodel/openlanguagemodel.git
cd openlanguagemodel
pip install -e .

For development:

pip install -e ".[dev]"
pytest tests

Optional extras:

pip install -e ".[wandb]"  # Weights & Biases logging
pip install -e ".[docs]"   # documentation tooling

See docs/installation.md for dependency and release-build details.

Documentation Flow

Project Status

OLM v2.2 is the stabilization and release-readiness pass: tied output embeddings by default, model-family smoke coverage, AutoTrainer, streaming datasets, AMP, checkpointing, single-node DDP/FSDP paths, clearer installation docs, and a stronger generated API reference. Multi-node training remains a v4 roadmap item.

Citation

@software{openlanguagemodel2026,
  title = {OpenLanguageModel},
  author = {Tavish Mankash and Vardhaman Kalloli and Keshava Prasad},
  year = {2026},
  url = {https://github.com/openlanguagemodel/openlanguagemodel}
}

License

MIT. See LICENSE.

Metadata

Release files for openlanguagemodel 2.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for openlanguagemodel 2.2.1
File Size Uploaded
openlanguagemodel-2.2.1.tar.gz 142.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for openlanguagemodel 2.2.1
File Interpreter ABI Platform
openlanguagemodel-2.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 351.6 kB

Release files / openlanguagemodel-2.2.1.tar.gz

Download URL openlanguagemodel-2.2.1.tar.gz
Size 142.4 kB
Tags Source
SHA-256 checksum
How to use checksums
82669507f2520e52b951c00daf90b8c6158d33d0ea95e8660f8a49b9b60c8e5e
BLAKE2b-256 checksum
How to use checksums
8253b1b8d7d872f9bbe59b9c12d41a79cf82f28c811dee1fdb9eafaf0f29d0db
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release files / openlanguagemodel-2.2.1-py3-none-any.whl

Download URL openlanguagemodel-2.2.1-py3-none-any.whl
Size 209.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
30618a7890ae9a4d8804aa97590b691d111cfe37845634660d824d5e634e074d
BLAKE2b-256 checksum
How to use checksums
dd607c456379b1ad8a07fa3be0f72ab6c82ee3acd94a026136ea0b2dde5e7050
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

2.2.1 This release

2 release files

2.2.0

2 release files

2.1

2 release files

2.0

2 release files

1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page