alignrl
From base model to deployed reasoning agent - every LLM post-training technique, implemented and benchmarked.
What is this?
A Python package implementing the complete LLM post-training pipeline: Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO) with verifiable math rewards, and Direct Preference Optimization (DPO). Includes evaluation benchmarks via lm-evaluation-harness and multi-backend inference serving (Unsloth, vLLM, MLX). Built for learning and demonstration, designed to run on free Colab GPUs with QLoRA and Unsloth for memory-efficient training on Qwen2.5-3B.
Pipeline
graph LR
A[Qwen2.5-3B<br/>Base Model] --> B[SFT<br/>Instruction Following]
B --> C[GRPO<br/>Math Reasoning via RL]
B --> D[DPO<br/>Preference Alignment]
C --> E[Evaluation<br/>GSM8K, MATH, ARC]
D --> E
E --> F[Inference<br/>Unsloth / vLLM / MLX]
Quick Start
# Install
pip install git+https://github.com/sacredvoid/alignrl.git
# Train (SFT as an example)
alignrl train sft -c configs/sft.yaml
# Evaluate
alignrl eval --adapter ./outputs/sft/final --stage sft
# Launch comparison demo
alignrl serve --stages base sft=./outputs/sft/final grpo=./outputs/grpo/final
For GPU training, install with the train and unsloth extras:
pip install "alignrl[train,unsloth] @ git+https://github.com/sacredvoid/alignrl.git"
Notebooks
Each notebook is self-contained and runs end-to-end on a free Colab T4 GPU.
Benchmark Results
All evaluations run on Qwen2.5-3B with QLoRA adapters. Best score per benchmark in bold.
| Benchmark | Metric | Base | SFT | GRPO | DPO |
|---|---|---|---|---|---|
| GSM8K | exact_match | 0.31 | 0.45 | 0.62 | 0.43 |
| MATH | exact_match | 0.12 | 0.18 | 0.29 | 0.17 |
| ARC-Challenge | acc_norm | 0.48 | 0.54 | 0.52 | 0.55 |
Key takeaways:
- GRPO dominates math reasoning - GSM8K jumps from 31% to 62% (2x), MATH from 12% to 29% (2.4x)
- DPO edges out on general reasoning - ARC-Challenge best at 55%, suggesting preference alignment improves broad task quality
- SFT is a strong baseline - consistent improvement across all benchmarks before any RL
Module Reference
| Module | Purpose | Key Class |
|---|---|---|
alignrl.sft |
Supervised Fine-Tuning with QLoRA | SFTRunner |
alignrl.grpo |
RL with Verifiable Math Rewards | GRPORunner |
alignrl.dpo |
Direct Preference Optimization | DPORunner |
alignrl.eval |
Benchmark evaluation harness | EvalRunner |
alignrl.inference |
Multi-backend model serving | ModelServer |
alignrl.rewards |
Math reward verifiers for GRPO | math_verify_reward |
alignrl.demo |
Gradio comparison UI | create_demo |
alignrl.cli |
CLI entry point (train, eval, serve) |
main |
alignrl.config |
Pydantic-validated training configs | BaseTrainConfig |
alignrl.types |
Shared protocols and result types | Trainer, TrainResult, EvalResult |
Architecture
The codebase follows a few core design decisions:
- Pydantic configs - Every training stage uses a typed config class inheriting from
BaseTrainConfig, loadable from YAML files. Validation happens at construction time, not at training time. - Common Trainer protocol -
SFTRunner,GRPORunner, andDPORunnerall implement theTrainerprotocol (train(),save(),load()), making them interchangeable in pipelines and tests. - Lazy imports - Heavy dependencies (torch, transformers, unsloth, vllm, mlx-lm) are imported inside methods, not at module level. The base package installs in seconds with just pydantic and pyyaml.
- Unsloth for speed - All training uses Unsloth's
FastLanguageModelwith gradient checkpointing, cutting VRAM usage roughly in half compared to vanilla transformers. Fits Qwen2.5-3B training on a free Colab T4 (16GB). - Structured results - Training returns
TrainResult, evaluation returnsEvalResult. Both are frozen dataclasses that serialize to JSON for the results dashboard.
Project Structure
alignrl/
configs/ # YAML configs for each training stage
docs/ # GitHub Pages results dashboard
notebooks/ # Colab-ready Jupyter notebooks
results/ # Benchmark JSON (consumed by dashboard)
src/alignrl/ # Package source
tests/ # 49 unit tests (pytest)
pyproject.toml # Hatchling build, optional dependency groups
Tech Stack
| Category | Tools |
|---|---|
| Training | TRL, Unsloth, PEFT, bitsandbytes |
| Evaluation | lm-evaluation-harness |
| Inference | vLLM, MLX-LM, Unsloth |
| Demo | Gradio |
| Config | Pydantic, PyYAML |
| Quality | Ruff, mypy, pytest |
License
Release files for alignrl 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| alignrl-0.3.0.tar.gz | 65.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| alignrl-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 86.9 kB
Release files / alignrl-0.3.0.tar.gz
| Download URL | alignrl-0.3.0.tar.gz |
|---|---|
| Size | 65.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
14324beb61d25525a7e69e7094be2925a8906af1e813c0eb31de23249320a91e
|
|
BLAKE2b-256 checksum How to use checksums |
ff9126dd5187da3ef3c31b52a232887cd4ca4f686993ff627c5830c4d641f368
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Mar 26, 2026.
Transparency logRelease files / alignrl-0.3.0-py3-none-any.whl
| Download URL | alignrl-0.3.0-py3-none-any.whl |
|---|---|
| Size | 21.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b0e3ff837f575ceec12800e04ca13ddd2adafa33771a109582db4d6e284d5bdc
|
|
BLAKE2b-256 checksum How to use checksums |
3d3a7bad2749ffac57dff5516201ef12d17f91c0a9a879e84323b169fc616118
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Mar 26, 2026.
Transparency log