silver-run
Local experiment tracking that is transparent enough to debug with a text editor.
Backend-neutral ML run lifecycle, events, and checkpoints for Silver. A Python package designed for ML researchers who need flexible training orchestration across different frameworks.
vs. MLflow and W&B
Silver gives up hosted dashboards, team collaboration UI, and managed artifact
storage. In exchange, runs are local JSON/JSONL you can inspect with cat or
jq, replay without a server, and keep without vendor lock-in.
Replay the experiment visually
restored = store.load(run.id)
open("run-timeline.svg", "w", encoding="utf-8").write(restored.to_svg())
# Or render from the store and save the returned SVG text:
open("run-timeline.svg", "w", encoding="utf-8").write(store.visualize(run.id))
The timeline uses persisted event timestamps and recorded loss curves, so it is replayable after the process exits. See the visual evidence model.
Installation
pip install silver-run
Quick Start
Durable tracking without a server
from silver_run import (
FileCheckpointStore, LocalRunStore, TrainingRun, TrainingRunOptions,
)
store = LocalRunStore(".silver/runs")
run = TrainingRun(TrainingRunOptions(
metadata={"model": "tabular-v1", "dataset": "customers-2026"},
run_store=store,
checkpoint_store=FileCheckpointStore(".silver/checkpoints"),
))
run.start()
run.emit("epoch", {"epoch": 1, "metrics": {"loss": 0.42}})
run.complete()
restored = store.load(run.id)
print(restored.state, restored.metrics, restored.duration)
Run manifests are atomic JSON; events are append-only JSONL; checkpoint IDs are path-safe; checkpoint payloads are explicit JSON instead of unsafe pickle.
from silver_run import TrainingRun, TrainingBackend, TrainingContext
import asyncio
class MyBackend(TrainingBackend):
async def run(self, context: TrainingContext):
for epoch in range(10):
if context.should_stop():
break
# Your training logic here
context.emit({
"kind": "epoch",
"epoch": epoch,
"metrics": {"loss": 0.5 - epoch * 0.05}
})
await asyncio.sleep(0.1)
async def main():
run = TrainingRun()
backend = MyBackend()
final_state = await run.execute(backend)
print(f"Run finished with state: {final_state.value}")
asyncio.run(main())
Features
- Training Lifecycle Management: Full state machine (created, running, paused, stopped, cancelled, completed, failed)
- Event Logging: Comprehensive event tracking with timestamps for training observability
- Checkpoint Management: Pluggable storage backends for model checkpointing
- Backend-Agnostic: Works with PyTorch, TensorFlow, JAX, or any custom training framework
- Async/Await Support: Modern Python async patterns for concurrent training
- Pause/Resume: Control long-running training jobs with pause and resume functionality
- Type Safety: Full type hints for better IDE support and fewer bugs
Use Cases
PyTorch Training Integration
from silver_run import TrainingRun, TrainingBackend, TrainingContext
import torch
import asyncio
class PyTorchBackend(TrainingBackend):
def __init__(self, model, optimizer, train_loader):
self.model = model
self.optimizer = optimizer
self.train_loader = train_loader
async def run(self, context: TrainingContext):
for epoch in range(10):
if context.should_stop():
break
self.model.train()
total_loss = 0
for batch_idx, (data, target) in enumerate(self.train_loader):
self.optimizer.zero_grad()
output = self.model(data)
loss = torch.nn.functional.cross_entropy(output, target)
loss.backward()
self.optimizer.step()
total_loss += loss.item()
# Emit epoch completion event
context.emit({
"kind": "epoch",
"epoch": epoch,
"metrics": {"loss": total_loss / len(self.train_loader)}
})
# Checkpoint every 5 epochs
if epoch % 5 == 0:
await context.checkpoint({
"epoch": epoch,
"model_state_dict": self.model.state_dict(),
"optimizer_state_dict": self.optimizer.state_dict()
})
async def main():
model = torch.nn.Linear(10, 2)
optimizer = torch.optim.Adam(model.parameters())
train_loader = [...] # Your data loader
run = TrainingRun()
backend = PyTorchBackend(model, optimizer, train_loader)
final_state = await run.execute(backend)
# Review events
for event in run.events():
print(f"{event.kind}: {event.data}")
asyncio.run(main())
Training with Pause/Resume
from silver_run import TrainingRun, TrainingBackend
import asyncio
class LongRunningBackend(TrainingBackend):
async def run(self, context: TrainingContext):
for step in range(1000):
if context.should_stop():
break
# Simulate training step
await asyncio.sleep(0.01)
# Emit progress
if step % 100 == 0:
context.emit({
"kind": "progress",
"step": step,
"total": 1000
})
async def main():
run = TrainingRun()
backend = LongRunningBackend()
# Start training in background
training_task = asyncio.create_task(run.execute(backend))
# Pause after some time
await asyncio.sleep(0.5)
run.pause()
print("Training paused")
# Resume after some time
await asyncio.sleep(0.5)
run.resume()
print("Training resumed")
# Wait for completion
final_state = await training_task
print(f"Training finished: {final_state.value}")
asyncio.run(main())
Custom Checkpoint Storage
from silver_run import TrainingRun, CheckpointStore, Checkpoint
import asyncio
class S3CheckpointStore(CheckpointStore):
def __init__(self, bucket, prefix):
self.bucket = bucket
self.prefix = prefix
self.checkpoints = {}
async def save(self, checkpoint: Checkpoint):
# Save to S3
key = f"{self.prefix}/{checkpoint.id}"
print(f"Saving checkpoint to S3: {key}")
self.checkpoints[checkpoint.id] = checkpoint
async def latest(self):
if not self.checkpoints:
return None
return list(self.checkpoints.values())[-1]
async def get(self, id: str):
return self.checkpoints.get(id)
async def main():
store = S3CheckpointStore("my-bucket", "checkpoints")
run = TrainingRun(options=TrainingRunOptions(checkpoint_store=store))
# Use custom checkpoint store
await run.checkpoint({"model": "state"}, "checkpoint-1")
latest = await run.latest_checkpoint()
print(f"Latest checkpoint: {latest.id}")
asyncio.run(main())
Advanced Usage
Event Filtering and Analysis
from silver_run import TrainingRun
# Filter events by type
def get_epoch_events(run):
return [e for e in run.events() if e.kind == "epoch"]
def get_error_events(run):
return [e for e in run.events() if e.kind == "error"]
# Analyze training progression
def analyze_training(run):
epoch_events = get_epoch_events(run)
losses = [e.data.get("metrics", {}).get("loss") for e in epoch_events]
if losses:
print(f"Initial loss: {losses[0]}")
print(f"Final loss: {losses[-1]}")
print(f"Loss reduction: {losses[0] - losses[-1]}")
Multi-Run Experiments
from silver_run import TrainingRun
import asyncio
async def run_experiment(config):
run = TrainingRun()
backend = MyBackend(config)
return await run.execute(backend)
async def main():
configs = [
{"learning_rate": 0.001},
{"learning_rate": 0.01},
{"learning_rate": 0.1}
]
results = await asyncio.gather(*[
run_experiment(config) for config in configs
])
for config, result in zip(configs, results):
print(f"LR {config['learning_rate']}: {result.value}")
asyncio.run(main())
Requirements
- Python 3.10+
Development
# Install development dependencies
pip install -e ".[dev]"
# Run tests
pytest
# Run tests with coverage
pytest --cov=silver_run --cov-report=html
# Run linting
flake8 src/ tests/
mypy src/
Contributing
Contributions are welcome! Please see CONTRIBUTING.md for guidelines.
License
Apache-2.0 - see LICENSE file for details.
Related Packages
- silver-data - Dataset handling
- silver-diagnostics - ML diagnostics
- silver-adapters - Framework adapters
Evidence you can replay
Run timelines and checkpoints are durable evidence for the decision loop:
persist the SVG returned by store.visualize(...), apply one recommended
change, and compare the next run's metrics and events. This keeps debugging
reproducible instead of relying on screenshots or memory.
Release files for silver-run 1.5.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| silver_run-1.5.1.tar.gz | 3.4 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| silver_run-1.5.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.4 MB
Release files / silver_run-1.5.1.tar.gz
| Download URL | silver_run-1.5.1.tar.gz |
|---|---|
| Size | 3.4 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
62731a9dd4859b704b41b4f32c706b29fc7a28148faad0ac8f12bad0aee54508
|
|
BLAKE2b-256 checksum How to use checksums |
b117caddbb4b43ef222108257d4ea6e3acd28eb41379a71bfe49b04598ead34d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency logRelease files / silver_run-1.5.1-py3-none-any.whl
| Download URL | silver_run-1.5.1-py3-none-any.whl |
|---|---|
| Size | 15.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2f95568f53d92008a44176f9c6209868cbaf21326a9eb64c11fed6be1a169191
|
|
BLAKE2b-256 checksum How to use checksums |
8a930e1b598c34b554ebdec87202e85d89a0512c01c50ee15d1788a6ca19676c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.
Transparency log