Skip to main content

silver-run

Silver local experiment tracking

Local experiment tracking that is transparent enough to debug with a text editor.

Python Version License CI Code Style

Backend-neutral ML run lifecycle, events, and checkpoints for Silver. A Python package designed for ML researchers who need flexible training orchestration across different frameworks.

Replay the experiment visually

Neural-network inputs, hidden layers, activations, gradients, prediction, and training health

restored = store.load(run.id)
open("run-timeline.svg", "w", encoding="utf-8").write(restored.to_svg())

# Or render from the store and save the returned SVG text:
open("run-timeline.svg", "w", encoding="utf-8").write(store.visualize(run.id))

The timeline uses persisted event timestamps and recorded loss curves, so it is replayable after the process exits. See the visual evidence model.

Installation

pip install silver-run

Quick Start

Durable tracking without a server

from silver_run import (
    FileCheckpointStore, LocalRunStore, TrainingRun, TrainingRunOptions,
)

store = LocalRunStore(".silver/runs")
run = TrainingRun(TrainingRunOptions(
    metadata={"model": "tabular-v1", "dataset": "customers-2026"},
    run_store=store,
    checkpoint_store=FileCheckpointStore(".silver/checkpoints"),
))

run.start()
run.emit("epoch", {"epoch": 1, "metrics": {"loss": 0.42}})
run.complete()

restored = store.load(run.id)
print(restored.state, restored.metrics, restored.duration)

Run manifests are atomic JSON; events are append-only JSONL; checkpoint IDs are path-safe; checkpoint payloads are explicit JSON instead of unsafe pickle.

from silver_run import TrainingRun, TrainingBackend, TrainingContext
import asyncio

class MyBackend(TrainingBackend):
    async def run(self, context: TrainingContext):
        for epoch in range(10):
            if context.should_stop():
                break
            # Your training logic here
            context.emit({
                "kind": "epoch",
                "epoch": epoch,
                "metrics": {"loss": 0.5 - epoch * 0.05}
            })
            await asyncio.sleep(0.1)

async def main():
    run = TrainingRun()
    backend = MyBackend()
    final_state = await run.execute(backend)
    print(f"Run finished with state: {final_state.value}")

asyncio.run(main())

Features

  • Training Lifecycle Management: Full state machine (created, running, paused, stopped, cancelled, completed, failed)
  • Event Logging: Comprehensive event tracking with timestamps for training observability
  • Checkpoint Management: Pluggable storage backends for model checkpointing
  • Backend-Agnostic: Works with PyTorch, TensorFlow, JAX, or any custom training framework
  • Async/Await Support: Modern Python async patterns for concurrent training
  • Pause/Resume: Control long-running training jobs with pause and resume functionality
  • Type Safety: Full type hints for better IDE support and fewer bugs

Use Cases

PyTorch Training Integration

from silver_run import TrainingRun, TrainingBackend, TrainingContext
import torch
import asyncio

class PyTorchBackend(TrainingBackend):
    def __init__(self, model, optimizer, train_loader):
        self.model = model
        self.optimizer = optimizer
        self.train_loader = train_loader
    
    async def run(self, context: TrainingContext):
        for epoch in range(10):
            if context.should_stop():
                break
            
            self.model.train()
            total_loss = 0
            
            for batch_idx, (data, target) in enumerate(self.train_loader):
                self.optimizer.zero_grad()
                output = self.model(data)
                loss = torch.nn.functional.cross_entropy(output, target)
                loss.backward()
                self.optimizer.step()
                total_loss += loss.item()
            
            # Emit epoch completion event
            context.emit({
                "kind": "epoch",
                "epoch": epoch,
                "metrics": {"loss": total_loss / len(self.train_loader)}
            })
            
            # Checkpoint every 5 epochs
            if epoch % 5 == 0:
                await context.checkpoint({
                    "epoch": epoch,
                    "model_state_dict": self.model.state_dict(),
                    "optimizer_state_dict": self.optimizer.state_dict()
                })

async def main():
    model = torch.nn.Linear(10, 2)
    optimizer = torch.optim.Adam(model.parameters())
    train_loader = [...]  # Your data loader
    
    run = TrainingRun()
    backend = PyTorchBackend(model, optimizer, train_loader)
    final_state = await run.execute(backend)
    
    # Review events
    for event in run.events():
        print(f"{event.kind}: {event.data}")

asyncio.run(main())

Training with Pause/Resume

from silver_run import TrainingRun, TrainingBackend
import asyncio

class LongRunningBackend(TrainingBackend):
    async def run(self, context: TrainingContext):
        for step in range(1000):
            if context.should_stop():
                break
            
            # Simulate training step
            await asyncio.sleep(0.01)
            
            # Emit progress
            if step % 100 == 0:
                context.emit({
                    "kind": "progress",
                    "step": step,
                    "total": 1000
                })

async def main():
    run = TrainingRun()
    backend = LongRunningBackend()
    
    # Start training in background
    training_task = asyncio.create_task(run.execute(backend))
    
    # Pause after some time
    await asyncio.sleep(0.5)
    run.pause()
    print("Training paused")
    
    # Resume after some time
    await asyncio.sleep(0.5)
    run.resume()
    print("Training resumed")
    
    # Wait for completion
    final_state = await training_task
    print(f"Training finished: {final_state.value}")

asyncio.run(main())

Custom Checkpoint Storage

from silver_run import TrainingRun, CheckpointStore, Checkpoint
import asyncio

class S3CheckpointStore(CheckpointStore):
    def __init__(self, bucket, prefix):
        self.bucket = bucket
        self.prefix = prefix
        self.checkpoints = {}
    
    async def save(self, checkpoint: Checkpoint):
        # Save to S3
        key = f"{self.prefix}/{checkpoint.id}"
        print(f"Saving checkpoint to S3: {key}")
        self.checkpoints[checkpoint.id] = checkpoint
    
    async def latest(self):
        if not self.checkpoints:
            return None
        return list(self.checkpoints.values())[-1]
    
    async def get(self, id: str):
        return self.checkpoints.get(id)

async def main():
    store = S3CheckpointStore("my-bucket", "checkpoints")
    run = TrainingRun(options=TrainingRunOptions(checkpoint_store=store))
    
    # Use custom checkpoint store
    await run.checkpoint({"model": "state"}, "checkpoint-1")
    latest = await run.latest_checkpoint()
    print(f"Latest checkpoint: {latest.id}")

asyncio.run(main())

Advanced Usage

Event Filtering and Analysis

from silver_run import TrainingRun

# Filter events by type
def get_epoch_events(run):
    return [e for e in run.events() if e.kind == "epoch"]

def get_error_events(run):
    return [e for e in run.events() if e.kind == "error"]

# Analyze training progression
def analyze_training(run):
    epoch_events = get_epoch_events(run)
    losses = [e.data.get("metrics", {}).get("loss") for e in epoch_events]
    
    if losses:
        print(f"Initial loss: {losses[0]}")
        print(f"Final loss: {losses[-1]}")
        print(f"Loss reduction: {losses[0] - losses[-1]}")

Multi-Run Experiments

from silver_run import TrainingRun
import asyncio

async def run_experiment(config):
    run = TrainingRun()
    backend = MyBackend(config)
    return await run.execute(backend)

async def main():
    configs = [
        {"learning_rate": 0.001},
        {"learning_rate": 0.01},
        {"learning_rate": 0.1}
    ]
    
    results = await asyncio.gather(*[
        run_experiment(config) for config in configs
    ])
    
    for config, result in zip(configs, results):
        print(f"LR {config['learning_rate']}: {result.value}")

asyncio.run(main())

Requirements

  • Python 3.10+

Development

# Install development dependencies
pip install -e ".[dev]"

# Run tests
pytest

# Run tests with coverage
pytest --cov=silver_run --cov-report=html

# Run linting
flake8 src/ tests/
mypy src/

Contributing

Contributions are welcome! Please see CONTRIBUTING.md for guidelines.

License

Apache-2.0 - see LICENSE file for details.

Evidence you can replay

Run timelines and checkpoints are durable evidence for the decision loop: persist the SVG returned by store.visualize(...), apply one recommended change, and compare the next run's metrics and events. This keeps debugging reproducible instead of relying on screenshots or memory.

Release files for silver-run 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for silver-run 1.3.0
File Size Uploaded
silver_run-1.3.0.tar.gz 3.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for silver-run 1.3.0
File Interpreter ABI Platform
silver_run-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 3.4 MB

Release files / silver_run-1.3.0.tar.gz

Download URL silver_run-1.3.0.tar.gz
Size 3.4 MB
Tags Source
SHA-256 checksum
How to use checksums
af3b90f173ee07741c841ad42f0ebd82ad21fe48344e77f407a01a854a22f14d
BLAKE2b-256 checksum
How to use checksums
42e515aa5450296c7b5b6b72f81f665b284f10d592c671c59ca58f39d1adef95
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release files / silver_run-1.3.0-py3-none-any.whl

Download URL silver_run-1.3.0-py3-none-any.whl
Size 14.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
44815acbf9e146a90f4f81b044af67bd3ed19496d99babb6e105ead2d550a713
BLAKE2b-256 checksum
How to use checksums
f6d0a253fb4f57b261605176fe671002643d67df8cf6161afa94cb0e228a6bf9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release history Release notifications | RSS feed

1.5.1

2 release files

1.5.0

2 release files

1.4.0

2 release files

This release

1.3.0 This release

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page