llm-hpo-optimizer
A general-purpose, multi-objective hyperparameter optimization library for LLM training.
llm-hpo-optimizer provides a clean, pluggable framework for running hyperparameter sweeps against any model. The core optimizer needs only the Python standard library. Optional extras unlock Bayesian (TPE) search via Optuna and a built-in adapter for nanoGPT.
Features
- Model-agnostic: implement one
TrainingAdapterclass to optimize any model - Multi-objective scoring: weighted combination of validation loss, memory, parameter count, and training time
- Pareto front extraction: identifies non-dominated configurations automatically
- Two search strategies: pure random search (zero dependencies) or Bayesian TPE via Optuna
- Reproducible: all randomness is controlled by a single
seed - Zero-dependency core: the optimizer engine works with nothing beyond the standard library
- Typed: ships with a
py.typedmarker; all public APIs are fully annotated
Installation
# Core library — random search, zero extra dependencies
pip install llm-hpo-optimizer
# + Bayesian/TPE search (Optuna)
pip install "llm-hpo-optimizer[optuna]"
# + Built-in nanoGPT adapter
pip install "llm-hpo-optimizer[nanogpt]"
Quick Start
from llm_hpo_optimizer import (
optimize,
SearchSpace,
Categorical,
UniformFloat,
LogUniformFloat,
TrainingAdapter,
TrainingResult,
)
# 1. Implement your adapter
class MyAdapter(TrainingAdapter):
def run_trial(self, hyperparams: dict, seed: int) -> TrainingResult:
# Train your model here …
return TrainingResult(
train_loss=0.5,
val_loss=0.6,
best_val_loss=0.55,
best_val_iter=100,
training_iterations=500,
parameter_count=1_000_000,
model_size_mb=4.0,
peak_memory_mb=512.0,
training_time_seconds=30.0,
samples_seen=100_000,
samples_per_second=3333.0,
device="cpu",
seed=seed,
)
# 2. Define the search space
search_space = SearchSpace({
"learning_rate": LogUniformFloat(1e-5, 1e-2),
"hidden_size": Categorical([64, 128, 256, 512]),
"dropout": UniformFloat(0.0, 0.5),
})
# 3. Run the sweep
result = optimize(
method="random", # or "optuna" (requires llm-hpo-optimizer[optuna])
adapter=MyAdapter(),
search_space=search_space,
baseline_hyperparams={"learning_rate": 1e-3, "hidden_size": 128, "dropout": 0.1},
optimization_iterations=20,
seed=42,
)
print(result["best_loss_config"]) # best hyperparams by validation loss
print(result["best_balanced_config"]) # best hyperparams by multi-objective score
Parameter / Search Space Definition
| Class | Description | Example |
|---|---|---|
Categorical |
Discrete, unordered choices | Categorical([16, 32, 64]) |
UniformFloat |
Continuous uniform distribution | UniformFloat(0.0, 0.5) |
LogUniformFloat |
Log-scale uniform (for learning rate) | LogUniformFloat(1e-5, 1e-2) |
from llm_hpo_optimizer import SearchSpace, Categorical, UniformFloat, LogUniformFloat
space = SearchSpace(
parameters={
"learning_rate": LogUniformFloat(1e-5, 1e-2),
"batch_size": Categorical([16, 32, 64]),
"dropout": UniformFloat(0.0, 0.5),
},
constraints=[lambda cfg: cfg["batch_size"] >= 16], # optional
)
Note: If both
n_headandn_embdappear in your search space, the constraintn_embd % n_head == 0is enforced automatically.
Objective Function / Adapter
Implement one method:
class MyAdapter(TrainingAdapter):
def run_trial(self, hyperparams: dict, seed: int) -> TrainingResult:
# Use hyperparams to configure and train your model.
# Use seed to make training reproducible.
# Catch OOM/runtime errors; return failed=True instead of raising.
return TrainingResult(train_loss=..., val_loss=..., ...)
Running the Optimization
result = optimize(
method="random", # "random" or "optuna"
adapter=MyAdapter(),
search_space=search_space,
baseline_hyperparams={...}, # config for trial 0 (the baseline)
optimization_iterations=20, # optimizer-suggested trials
seed=42, # reproducibility
output_path="results.json", # optional — write JSON to disk
)
Accessing Results
# Best hyperparams by validation loss
print(result["best_loss_config"]) # dict
print(result["best_loss_value"]) # float
# Best hyperparams by weighted multi-objective score
print(result["best_balanced_config"]) # dict
print(result["best_balanced_score"]) # float (lower = better)
# The fixed baseline (trial 0)
print(result["baseline_config"]) # dict
# Pareto-optimal configurations
for record in result["pareto_front"]:
print(record["hyperparameters"], record["metrics"]["best_val_loss"])
# All trials
for record in result["all_records"]:
print(record["experiment_id"], record["metrics"]["best_val_loss"])
Advanced Usage
Bayesian Search (Optuna / TPE)
pip install "llm-hpo-optimizer[optuna]"
result = optimize(
method="optuna",
adapter=MyAdapter(),
search_space=search_space,
baseline_hyperparams=baseline,
optimization_iterations=30,
seed=42,
)
Custom Scoring Weights
result = optimize(
...,
score_weights={
"val_loss": 0.7,
"peak_memory_mb": 0.1,
"parameter_count": 0.1,
"training_time": 0.1,
},
)
Using the Runner Directly
from llm_hpo_optimizer import RandomSearchOptimizer, ExperimentRunner
opt = RandomSearchOptimizer(search_space, seed=42)
runner = ExperimentRunner(
optimizer=opt,
adapter=MyAdapter(),
baseline_hyperparams=baseline,
optimization_iterations=20,
seed=42,
)
records = runner.run()
print(f"Pareto front: {len(runner.pareto)} configs")
Built-in nanoGPT Adapter
pip install "llm-hpo-optimizer[nanogpt]"
# nanoGPT's model.py must be on your Python path:
export PYTHONPATH=/path/to/nanoGPT:$PYTHONPATH
from llm_hpo_optimizer import optimize
result = optimize(
dataset_dir="data/shakespeare_char",
method="optuna",
optimization_iterations=15,
training_iterations_per_trial=1000,
)
print(result["best_balanced_config"])
Reproducibility
All randomness is controlled by the seed parameter:
result_a = optimize(..., seed=42)
result_b = optimize(..., seed=42)
assert result_a["all_records"] == result_b["all_records"] # True
Supported Python Versions
Python 3.9, 3.10, 3.11, 3.12
License
MIT — see LICENSE.
Release files for llm-hpo-optimizer 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_hpo_optimizer-0.1.0.tar.gz | 30.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_hpo_optimizer-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 56.5 kB
Release files / llm_hpo_optimizer-0.1.0.tar.gz
| Download URL | llm_hpo_optimizer-0.1.0.tar.gz |
|---|---|
| Size | 30.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1226d4ee480565b876d24016ad34ff8d66dd10dae6180baec0cfa7d2e663905b
|
|
BLAKE2b-256 checksum How to use checksums |
335170904af72ab39620c8112c528dc8eead5e9b148d694e07d6ebb4c74d2d17
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.11
|
Release files / llm_hpo_optimizer-0.1.0-py3-none-any.whl
| Download URL | llm_hpo_optimizer-0.1.0-py3-none-any.whl |
|---|---|
| Size | 25.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7d5e31f85b535dd1acc50a8476b84fa29902e779769938d94e98f9720e97be8b
|
|
BLAKE2b-256 checksum How to use checksums |
8053ee205af1d99b260e5bdd94743dcb7159516e95d16b6c449669a91fdb9cbe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.10.11
|