Skip to main content
Archived

This project has been archived by its maintainers, and is no longer receiving any updates.

Language Modeling using Transformers (LMT)

Python PyTorch License Code style: ruff

A PyTorch implementation of transformer-based language models including GPT architecture for pretraining and fine-tuning. This project is designed for educational and research purposes to help users understand how the attention mechanism and Transformer architecture work in Large Language Models (LLMs).

🚀 Features

  • GPT Architecture: Complete implementation of decoder-only transformer models
  • Attention Mechanisms: Multi-head self-attention with causal masking
  • Tokenization: Multiple tokenizer implementations (BPE, Naive)
  • Training Pipeline: Comprehensive trainer with pretraining and fine-tuning support
  • Educational Focus: Well-documented code for learning transformer internals
  • Modern Stack: Built with PyTorch 2.7+, Python 3.11+

📦 Installation

Prerequisites

  • Python 3.11 or 3.12
  • PyTorch 2.7+

Install from PyPI

pip install language-modeling-transformers

Install from GitHub

pip install git+https://github.com/michaelellis003/LMT.git

🏃‍♂️ Quick Start

Basic Model Usage

from lmt import GPT, ModelConfig
from lmt.models.config import ModelConfigPresets
import torch

# Create a small GPT model
config = ModelConfigPresets.small_gpt()
model = GPT(config)

# Generate some text
input_ids = torch.randint(0, config.vocab_size, (1, 10))
with torch.no_grad():
    logits = model(input_ids)
    print(f"Output shape: {logits.shape}")  # (1, 10, vocab_size)

Training a Model

from lmt import Trainer, GPT
from lmt.training import BaseTrainingConfig
from lmt.models.config import ModelConfigPresets

# Configure model and training
model_config = ModelConfigPresets.small_gpt()
training_config = BaseTrainingConfig(
    num_epochs=10,
    batch_size=4,
    learning_rate=1e-4
)

# Initialize model and trainer
model = GPT(model_config)
trainer = Trainer(
    model=model,
    train_loader=your_train_loader,
    val_loader=your_val_loader,
    config=training_config
)

# Start training
trainer.train()

Using the Training Script

# Pretraining
python scripts/train.py --task pretraining --num_epochs 20 --batch_size 4

# Classification fine-tuning
python scripts/train.py --task classification --download_model --learning_rate 1e-5

📚 Documentation

Model Components

  • GPT: Main model class implementing decoder-only transformer
  • TransformerBlock: Individual transformer layer with attention and feed-forward
  • MultiHeadAttention: Multi-head self-attention mechanism
  • CausalAttention: Attention with causal masking for autoregressive generation

Tokenizers

  • BPETokenizer: Byte-Pair Encoding tokenizer
  • NaiveTokenizer: Simple character-level tokenizer
  • BaseTokenizer: Abstract base class for custom tokenizers

Training

  • Trainer: Main training orchestrator with support for pretraining and fine-tuning
  • BaseTrainingConfig: Configuration class for training parameters
  • Custom datasets and dataloaders: Support for various text datasets

🗂️ Project Structure

src/lmt/
├── __init__.py              # Main package exports
├── models/                  # Model architectures
│   ├── gpt/                # GPT implementation
│   ├── config.py           # Model configuration
│   └── utils.py            # Model utilities
├── layers/                  # Neural network layers
│   ├── attention/          # Attention mechanisms
│   └── transformers/       # Transformer blocks
├── tokenizer/              # Tokenization implementations
├── training/               # Training pipeline
└── generate.py             # Text generation utilities

scripts/
├── train.py                # Main training script
└── utils.py                # Training utilities

tests/                      # Comprehensive test suite
notebooks/                  # Educational Jupyter notebooks
docs/                       # Sphinx documentation

📊 Examples and Notebooks

Explore the interactive notebooks in the notebooks/ directory:

  • attention.ipynb: Understanding attention mechanisms
  • pretraining_gpt.ipynb: GPT pretraining walkthrough
  • tokenizer.ipynb: Tokenization techniques

🔧 Configuration

Model Configuration

from lmt.models.config import ModelConfig

config = ModelConfig(
    vocab_size=50257,
    embed_dim=768,
    context_length=1024,
    num_layers=12,
    num_heads=12,
    dropout=0.1
)

Training Configuration

from lmt.training.config import BaseTrainingConfig

training_config = BaseTrainingConfig(
    num_epochs=10,
    batch_size=8,
    learning_rate=3e-4,
    weight_decay=0.1,
    print_every=100,
    eval_every=500
)

📄 License

This project is licensed under the Apache License 2.0. See the LICENSE file for details.

Release files for language-modeling-transformers 0.2.8

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for language-modeling-transformers 0.2.8
File Size Uploaded
language_modeling_transformers-0.2.8.tar.gz 24.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for language-modeling-transformers 0.2.8
File Interpreter ABI Platform
language_modeling_transformers-0.2.8-py3-none-any.whl Python 3 none any Details

Total release size: 62.0 kB

Release files / language_modeling_transformers-0.2.8.tar.gz

Download URL language_modeling_transformers-0.2.8.tar.gz
Size 24.5 kB
Tags Source
SHA-256 checksum
How to use checksums
33af675f9a3930cce48c1deb1ee8fe331fbc683dd9a646bba224989690131c41
BLAKE2b-256 checksum
How to use checksums
8e611d8dff707dd3be702e253b5479315eb446c46ef0f0bcf161a9ffb17e5cde
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.7

Release files / language_modeling_transformers-0.2.8-py3-none-any.whl

Download URL language_modeling_transformers-0.2.8-py3-none-any.whl
Size 37.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a75b3c48c1ea9b8782928af88040b885a0af145120da81bac972ae6dbc4716f2
BLAKE2b-256 checksum
How to use checksums
56f40769d5b5c8637abddce0c0d17c5296ca623c52f6c6f0adc6941b86a49fbb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.7

Release history Release notifications | RSS feed

This release

0.2.8 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page