Tiny train logger
Project description
🚅🪵 Tiny Train Log 🚅🪵
Structured, queryable logging for ML experiments. Track metrics across machines, merge results into one SQLite database, and analyze everything with plain SQL.
Quick start
pip install tinytrainlog
from tinytrainlog import MetricsLogger
with MetricsLogger("./runs", run_name="lr-sweep-3e4") as log:
log.set_config({"model": "resnet50", "lr": 3e-4, "epochs": 10})
log.add_tags(["sweep", "baseline"])
for epoch in range(10):
train_loss = train(model, loader)
log.log_epoch(epoch=epoch, train_loss=train_loss)
val_loss, val_acc = evaluate(model, val_loader)
log.log_eval(epoch=epoch, val_loss=val_loss, val_acc=val_acc)
torch.save(model.state_dict(), log.checkpoint_path(epoch=epoch))
log.log_test(test_acc=0.94, test_loss=0.21)
Everything lands in a single runs.db SQLite file — query it however you like:
import sqlite3
conn = sqlite3.connect("./runs/runs.db")
# Compare learning rates across runs
conn.execute("""
SELECT c.run_name, c.value AS lr, e.value AS val_acc
FROM config c
JOIN eval e USING (run_name)
WHERE c.key = 'lr' AND e.key = 'val_acc'
ORDER BY CAST(e.value AS REAL) DESC
""").fetchall()
API
Initialization:
# Auto-generated run name (e.g. "bold-falcon")
log = MetricsLogger("./runs")
# Explicit name
log = MetricsLogger("./runs", run_name="lr-sweep-3e4")
# Override machine ID (defaults to hostname)
log = MetricsLogger("./runs", machine_id="gpu-box-1")
Logging:
| Method | Purpose |
|---|---|
set_config(dict) |
Hyperparameters and run metadata (JSON, upserts per key) |
add_tags(list) |
Labels for filtering (e.g. ["ablation", "v2"]) |
log_step(step, **metrics) |
Per-batch metrics (loss, lr, throughput) |
log_epoch(epoch, **metrics) |
Per-epoch metrics |
log_eval(step=, epoch=, **metrics) |
Validation metrics (requires at least one of step/epoch) |
log_test(**metrics) |
Final test results (upserts) |
Checkpoints and paths:
log.run_name # "bold-falcon"
log.run_dir # Path("./runs/bold-falcon")
log.checkpoint_dir # Path("./runs/bold-falcon/checkpoints")
log.checkpoint_path(epoch=5) # Path("./runs/bold-falcon/checkpoints/epoch_5.pt")
log.checkpoint_path(step=100) # Path("./runs/bold-falcon/checkpoints/step_100.pt")
Use run_dir to save extra artifacts (plots, predictions, etc.) alongside the run.
Multi-server merging
Ran experiments on multiple machines? Merge them into one database:
MetricsLogger.merge(target_dir="./all_runs", source_dir="/mnt/gpu-box-1/runs")
MetricsLogger.merge(target_dir="./all_runs", source_dir="/mnt/gpu-box-2/runs")
Deleting a run
Remove a run and all its data (config, metrics, checkpoints):
logger.delete_run("old-run")
Recipes
All data lives in a single SQLite file:
import sqlite3
conn = sqlite3.connect("./runs/runs.db")
List all runs with tags:
SELECT r.name, r.machine_id, r.created_at, GROUP_CONCAT(t.tag)
FROM runs r LEFT JOIN tags t ON t.run_name = r.name
GROUP BY r.name
Best run by test accuracy:
SELECT run_name, value FROM test
WHERE key = 'test_acc' ORDER BY value DESC LIMIT 1
Compare hyperparameters across runs:
SELECT r.name,
MAX(CASE WHEN c.key = 'lr' THEN c.value END) AS lr,
MAX(CASE WHEN c.key = 'model' THEN c.value END) AS model,
t.value AS test_acc
FROM runs r
JOIN config c ON c.run_name = r.name
LEFT JOIN test t ON t.run_name = r.name AND t.key = 'test_acc'
GROUP BY r.name
ORDER BY t.value DESC
Training curve for a run (for plotting):
SELECT step, key, value FROM steps
WHERE run_name = 'lr-sweep-3e4' ORDER BY step
Filter runs by tag:
SELECT run_name FROM tags WHERE tag = 'ablation'
Side-by-side eval comparison:
SELECT a.epoch, a.value AS model_a, b.value AS model_b
FROM eval a JOIN eval b USING (epoch, key)
WHERE a.run_name = 'model-a' AND b.run_name = 'model-b' AND a.key = 'val_acc'
ORDER BY a.epoch
Latest runs:
SELECT name, created_at FROM runs ORDER BY created_at DESC LIMIT 10
Runs from a specific machine:
SELECT name FROM runs WHERE machine_id = 'gpu-box-1'
Pareto frontier (accuracy vs. parameter count):
SELECT r.name, c.value AS param_count, t.value AS test_acc
FROM runs r
JOIN config c ON c.run_name = r.name AND c.key = 'param_count'
JOIN test t ON t.run_name = r.name AND t.key = 'test_acc'
WHERE NOT EXISTS (
SELECT 1 FROM config c2 JOIN test t2 ON t2.run_name = c2.run_name
WHERE c2.key = 'param_count' AND t2.key = 'test_acc'
AND CAST(c2.value AS REAL) <= CAST(c.value AS REAL)
AND t2.value >= t.value
AND (CAST(c2.value AS REAL) < CAST(c.value AS REAL) OR t2.value > t.value)
)
ORDER BY CAST(c.value AS REAL)
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file tinytrainlog-0.1.7.tar.gz.
File metadata
- Download URL: tinytrainlog-0.1.7.tar.gz
- Upload date:
- Size: 5.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bcb2eb9732c6966c58d48d9a7339acbf2b215969d6c27727449fa3570060db99
|
|
| MD5 |
f74960202b1caa1109b5ca89f4114f2f
|
|
| BLAKE2b-256 |
13bd062f7bccdf40288d11bcd4f7da6c9fb29cd2b1faddb108bd1f4b83b4658e
|
Provenance
The following attestation bundles were made for tinytrainlog-0.1.7.tar.gz:
Publisher:
python-publish.yml on jdhouseholder/tinytrainlog
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
tinytrainlog-0.1.7.tar.gz -
Subject digest:
bcb2eb9732c6966c58d48d9a7339acbf2b215969d6c27727449fa3570060db99 - Sigstore transparency entry: 1197023611
- Sigstore integration time:
-
Permalink:
jdhouseholder/tinytrainlog@ee0f2593b502251b1405650248fef65b1f80f543 -
Branch / Tag:
refs/tags/v0.1.7 - Owner: https://github.com/jdhouseholder
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@ee0f2593b502251b1405650248fef65b1f80f543 -
Trigger Event:
release
-
Statement type:
File details
Details for the file tinytrainlog-0.1.7-py3-none-any.whl.
File metadata
- Download URL: tinytrainlog-0.1.7-py3-none-any.whl
- Upload date:
- Size: 7.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7a5bb04c0bf818af274ec14fd85b0876e46e058bbbd1528f57ad85fd8560acdf
|
|
| MD5 |
119770c8d88b9cb65ddb93b9f86892af
|
|
| BLAKE2b-256 |
719684a23844211ea8d7389779cc95df5c2e201cc536825254a9d93e95e33e9a
|
Provenance
The following attestation bundles were made for tinytrainlog-0.1.7-py3-none-any.whl:
Publisher:
python-publish.yml on jdhouseholder/tinytrainlog
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
tinytrainlog-0.1.7-py3-none-any.whl -
Subject digest:
7a5bb04c0bf818af274ec14fd85b0876e46e058bbbd1528f57ad85fd8560acdf - Sigstore transparency entry: 1197023654
- Sigstore integration time:
-
Permalink:
jdhouseholder/tinytrainlog@ee0f2593b502251b1405650248fef65b1f80f543 -
Branch / Tag:
refs/tags/v0.1.7 - Owner: https://github.com/jdhouseholder
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@ee0f2593b502251b1405650248fef65b1f80f543 -
Trigger Event:
release
-
Statement type: