Skip to main content

Training Hub

Training Hub is an algorithm-focused interface for common LLM training, continual learning, and reinforcement learning techniques developed by the Red Hat AI Innovation Team.

PyPI version License Documentation (in progress)

Training Hub quickstart examples

New to Training Hub? Read our comprehensive introduction: Get Started with Language Model Post-Training Using Training Hub

Support Matrix

Algorithm Backends GPU Support Install Extra
Supervised Fine-tuning (SFT) InstructLab-Training Multi-GPU, multi-node base
Continual Learning (OSFT) RHAI Innovation Mini-Trainer Multi-GPU, multi-node base
Low-Rank Adaptation (LoRA) + SFT Unsloth Single-GPU, multi-GPU, multi-node [lora]
LoRA + GRPO (Adapter-Based RLVR) ART + Unsloth, verl Single-GPU (ART), multi-GPU, multi-node (verl) [grpo,lora]
GRPO (Full Fine-Tuning RLVR) verl Multi-GPU, multi-node [grpo]
GEPA (Genetic-Pareto Prompt Optimization) GEPA, MLflow CPU (API-based) [gepa]
Embedding Fine-Tuning SentenceTransformers Single-GPU, multi-GPU, CPU [embedding]

Implemented Algorithms

Supervised Fine-tuning (SFT)

Fine-tune language models on supervised datasets with support for:

  • Single-node and multi-node distributed training
  • Configurable training parameters (epochs, batch size, learning rate, etc.)
  • InstructLab-Training backend integration
from training_hub import sft

result = sft(
    model_path="Qwen/Qwen2.5-1.5B-Instruct",
    data_path="/path/to/data",
    ckpt_output_dir="/path/to/checkpoints",
    num_epochs=3,
    effective_batch_size=8,
    learning_rate=1e-5,
    max_seq_len=256,
    max_tokens_per_gpu=1024,
)

Orthogonal Subspace Fine-Tuning (OSFT)

OSFT allows you to fine-tune models while controlling how much of its existing behavior to preserve. Currently we have support for:

  • Single-node and multi-node distributed training
  • Configurable training parameters (epochs, batch size, learning rate, etc.)
  • RHAI Innovation Mini-Trainer backend integration

Here's a quick and minimal way to get started with OSFT:

from training_hub import osft

result = osft(
    model_path="/path/to/model",
    data_path="/path/to/data.jsonl", 
    ckpt_output_dir="/path/to/outputs",
    unfreeze_rank_ratio=0.25,
    effective_batch_size=16,
    max_tokens_per_gpu=2048,
    max_seq_len=1024,
    learning_rate=5e-6,
)

Low-Rank Adaptation (LoRA) + SFT

Parameter-efficient fine-tuning using LoRA with supervised fine-tuning. Features:

  • Memory-efficient training with significantly reduced VRAM requirements
  • Single-GPU and multi-GPU distributed training support
  • Unsloth backend for 2x faster training and 70% less memory usage
  • Support for QLoRA (4-bit quantization) for even lower memory usage
  • Compatible with messages and Alpaca dataset formats
from training_hub import lora_sft

result = lora_sft(
    model_path="Qwen/Qwen2.5-1.5B-Instruct",
    data_path="/path/to/data.jsonl",
    ckpt_output_dir="/path/to/outputs",
    lora_r=16,
    lora_alpha=32,
    num_epochs=3,
    learning_rate=2e-4
)

LoRA + GRPO (Adapter-Based RLVR)

Train LoRA adapters on tool-calling agents using Group Relative Policy Optimization with reinforcement learning from verifiable rewards. Features:

  • Single-turn and multi-turn tool-call verification with automatic per-turn decomposition
  • Two backends: OpenPipe ART + Unsloth GRPO (single-GPU, fast iteration) and verl (multi-GPU, scales to 70B+)
  • Built-in reward functions for tool-call correctness, or bring your own
  • Zero API cost training using ground-truth trace decomposition
from training_hub import lora_grpo

# Single GPU (ART backend)
result = lora_grpo(
    model_path="Qwen/Qwen3-4B",
    data_path="./tool_call_traces.jsonl",
    ckpt_output_dir="./grpo_output",
    backend="art",
    lora_r=32,
    lora_alpha=64,
    num_iterations=15,
)

# Multi GPU (verl backend)
result = lora_grpo(
    model_path="Qwen/Qwen3-4B",
    data_path="./tool_call_traces.jsonl",
    ckpt_output_dir="./grpo_output",
    backend="verl",
    n_gpus=4,
)

GRPO (Full Fine-Tuning RLVR)

Full-parameter GRPO training via the verl backend. Trains all model weights instead of LoRA adapters. Same data formats and reward functions as LoRA + GRPO.

from training_hub import grpo

result = grpo(
    model_path="Qwen/Qwen3-8B",
    data_path="./tool_call_traces.jsonl",
    ckpt_output_dir="./grpo_full_output",
    n_gpus=8,
    num_iterations=8,
)

GEPA (Genetic-Pareto Prompt Optimization)

Gradient-free prompt optimization using evolutionary search with Pareto-based selection and LLM-driven reflection. GEPA evolves textual prompts to maximize task performance without modifying model weights, so it needs no local GPU — it optimizes prompts by calling an LLM endpoint (hosted API or local vLLM/OpenAI-compatible server via api_base). Features:

  • Genetic-Pareto search with LLM reflection to propose improved prompts
  • Works with any model reachable through LiteLLM (hosted APIs or local endpoints)
  • Two backends: gepa (direct gepa.optimize()) and mlflow (MLflow prompt registry, scorers, and tracking)
from training_hub import gepa

result = gepa(
    seed_candidate={"system_prompt": "You are a helpful assistant. Answer the question."},
    task_lm="openai/gpt-4o-mini",
    data_path="./eval_data.jsonl",
    output_dir="./gepa_output",
    reflection_lm="openai/gpt-4o",
    max_metric_calls=200,
)

Embedding Fine-Tuning

Contrastive fine-tuning of sentence embedding models (e.g. all-MiniLM-L6-v2) so that inputs with the same label cluster together in embedding space. Designed for semantic routing / classification — route a query to one of N specialist lanes by nearest-anchor cosine similarity — but applicable to any task that benefits from tighter embedding clusters (retrieval, deduplication, clustering). Features:

  • Three contrastive losses: batch_all_triplet, batch_hard_triplet, and mnrl (Multiple Negatives Ranking Loss)
  • GROUP_BY_LABEL batch sampling so every batch contains multiple labels with at least two samples per label (required for triplet mining)
  • Auto-converts label datasets to (anchor, positive) pairs for MNRL
  • Custom loss_fn support for extensibility
  • Saves in standard sentence-transformers format
from training_hub import embedding_sft

result = embedding_sft(
    model_path="sentence-transformers/all-MiniLM-L6-v2",
    data_path="routing_train.jsonl",    # {"text": "...", "label": 0}
    ckpt_output_dir="./routing_model",
    loss_type="batch_all_triplet",
    num_epochs=20,
    batch_size=32,
    learning_rate=2e-5,
)

Installation

Basic Installation

This installs the base package, but doesn't install the CUDA-related dependencies which are required for GPU training.

pip install training-hub

Development Installation

git clone https://github.com/Red-Hat-AI-Innovation-Team/training_hub
cd training_hub
pip install -e .

For developers: See the Development Guide for detailed instructions on setting up your development environment, running local documentation, and contributing to Training Hub.

LoRA Support

For LoRA training with optimized dependencies:

pip install training-hub[lora]
# or for development
pip install -e .[lora]

Note: The LoRA extras include Unsloth optimizations and PyTorch-optimized xformers for better performance and compatibility.

GRPO Support

For LoRA + GRPO training (both ART and verl backends):

pip install training-hub[grpo,lora]

Note: When combining [grpo] with [cuda] extras, install them sequentially to avoid dependency solver conflicts:

pip install training-hub[grpo,lora]
pip install training-hub[cuda]

The [grpo] extras constrain torch, vllm, and transformers versions for verl compatibility, which may conflict with versions pulled by [cuda]. Sequential installation lets the solver pick compatible versions.

GEPA Support

For gradient-free prompt optimization (includes the MLflow backend):

pip install training-hub[gepa]
# or for development
pip install -e .[gepa]

Note: GEPA optimizes prompts via LLM API calls and does not require CUDA. To optimize against a local model, run a vLLM (or other OpenAI-compatible) server and pass its URL via the api_base parameter.

Embedding Support

For contrastive embedding fine-tuning (sentence-transformers backend):

pip install training-hub[embedding]
# or for development
pip install -e .[embedding]

Note: Embedding fine-tuning uses sentence-transformers>=5.0. It runs on CPU for small models (e.g. all-MiniLM-L6-v2, 23M params) and accelerates on a single or multi-GPU when CUDA is available.

CUDA Support

For GPU training with CUDA support:

pip install training-hub[cuda] --no-build-isolation
# or for development
pip install -e .[cuda] --no-build-isolation

Note: If you encounter build issues with flash-attn, install the base package first:

# Install base package (provides torch, packaging, wheel, ninja)
pip install training-hub
# Then install with CUDA extras
pip install training-hub[cuda] --no-build-isolation

# For development installation:
pip install -e . && pip install -e .[cuda] --no-build-isolation

If you're using uv, you can use the following commands to install the package:

# Installs training-hub from PyPI
uv pip install training-hub && uv pip install training-hub[cuda] --no-build-isolation

# For development:
git clone https://github.com/Red-Hat-AI-Innovation-Team/training_hub
cd training_hub
uv pip install -e . && uv pip install -e .[cuda] --no-build-isolation

Coding Agent Plugin

Training Hub is available as a plugin for two coding agents, bringing LLM training capabilities directly into your coding workflow.

Claude Code

Via org marketplace (recommended — includes all Red Hat AI plugins):

/plugin marketplace add Red-Hat-AI-Innovation-Team/plugins
/plugin install training-hub@Red-Hat-AI-Innovation-Team/plugins

Via this repo directly:

/plugin marketplace add Red-Hat-AI-Innovation-Team/training_hub
/plugin install training-hub@Red-Hat-AI-Innovation-Team/training_hub

From a local clone:

git clone https://github.com/Red-Hat-AI-Innovation-Team/training_hub.git
/plugin marketplace add /path/to/training_hub
Codex CLI
codex plugin marketplace add Red-Hat-AI-Innovation-Team/plugins

Then install the plugin from the marketplace. See .codex-plugin/INSTALL.md for manual installation.

After Installing

Invoke the setup-guide skill to configure your training algorithm, model, and data.

Skill Description
setup-guide Guided first-time configuration
training-guide Run LLM training or fine-tuning
memory-estimation Estimate GPU memory requirements

Getting Started

For comprehensive tutorials, examples, and documentation, see the examples directory.

Metadata

Release files for training-hub 0.10.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for training-hub 0.10.0
File Size Uploaded
training_hub-0.10.0.tar.gz 1.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for training-hub 0.10.0
File Interpreter ABI Platform
training_hub-0.10.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.6 MB

Release files / training_hub-0.10.0.tar.gz

Download URL training_hub-0.10.0.tar.gz
Size 1.4 MB
Tags Source
SHA-256 checksum
How to use checksums
4a3040a5a526fe99c31cb3fc41542b39a45ec54cc51a26cc4940ce4b1565f3c7
BLAKE2b-256 checksum
How to use checksums
58fb7014c61f7340c63406c19f407ee784bb3a594cfdc2d097bbedda5906f0de
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.

Transparency log

Release files / training_hub-0.10.0-py3-none-any.whl

Download URL training_hub-0.10.0-py3-none-any.whl
Size 120.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1517d6655e956a7ad8103f9e738afb55b2fb371c24cb86bbdf551176665ccb5e
BLAKE2b-256 checksum
How to use checksums
28704ea83e3ef18d65f8a4c9e9217ad3f5ba3bb2fac5cc9d663e517590c1f9c0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.12.9

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page