Skip to main content

silver-run

Silver local experiment tracking

Local experiment tracking that is transparent enough to debug with a text editor.

Python Version License CI Code Style

Backend-neutral ML run lifecycle, events, and checkpoints for Silver. A Python package designed for ML researchers who need flexible training orchestration across different frameworks.

Replay the experiment visually

Neural-network inputs, hidden layers, activations, gradients, prediction, and training health

restored = store.load(run.id)
open("run-timeline.svg", "w", encoding="utf-8").write(restored.to_svg())

# Or render from the store and save the returned SVG text:
open("run-timeline.svg", "w", encoding="utf-8").write(store.visualize(run.id))

The timeline uses persisted event timestamps and recorded loss curves, so it is replayable after the process exits. See the visual evidence model.

Installation

pip install silver-run

Quick Start

Durable tracking without a server

from silver_run import (
    FileCheckpointStore, LocalRunStore, TrainingRun, TrainingRunOptions,
)

store = LocalRunStore(".silver/runs")
run = TrainingRun(TrainingRunOptions(
    metadata={"model": "tabular-v1", "dataset": "customers-2026"},
    run_store=store,
    checkpoint_store=FileCheckpointStore(".silver/checkpoints"),
))

run.start()
run.emit("epoch", {"epoch": 1, "metrics": {"loss": 0.42}})
run.complete()

restored = store.load(run.id)
print(restored.state, restored.metrics, restored.duration)

Run manifests are atomic JSON; events are append-only JSONL; checkpoint IDs are path-safe; checkpoint payloads are explicit JSON instead of unsafe pickle.

from silver_run import TrainingRun, TrainingBackend, TrainingContext
import asyncio

class MyBackend(TrainingBackend):
    async def run(self, context: TrainingContext):
        for epoch in range(10):
            if context.should_stop():
                break
            # Your training logic here
            context.emit({
                "kind": "epoch",
                "epoch": epoch,
                "metrics": {"loss": 0.5 - epoch * 0.05}
            })
            await asyncio.sleep(0.1)

async def main():
    run = TrainingRun()
    backend = MyBackend()
    final_state = await run.execute(backend)
    print(f"Run finished with state: {final_state.value}")

asyncio.run(main())

Features

  • Training Lifecycle Management: Full state machine (created, running, paused, stopped, cancelled, completed, failed)
  • Event Logging: Comprehensive event tracking with timestamps for training observability
  • Checkpoint Management: Pluggable storage backends for model checkpointing
  • Backend-Agnostic: Works with PyTorch, TensorFlow, JAX, or any custom training framework
  • Async/Await Support: Modern Python async patterns for concurrent training
  • Pause/Resume: Control long-running training jobs with pause and resume functionality
  • Type Safety: Full type hints for better IDE support and fewer bugs

Use Cases

PyTorch Training Integration

from silver_run import TrainingRun, TrainingBackend, TrainingContext
import torch
import asyncio

class PyTorchBackend(TrainingBackend):
    def __init__(self, model, optimizer, train_loader):
        self.model = model
        self.optimizer = optimizer
        self.train_loader = train_loader
    
    async def run(self, context: TrainingContext):
        for epoch in range(10):
            if context.should_stop():
                break
            
            self.model.train()
            total_loss = 0
            
            for batch_idx, (data, target) in enumerate(self.train_loader):
                self.optimizer.zero_grad()
                output = self.model(data)
                loss = torch.nn.functional.cross_entropy(output, target)
                loss.backward()
                self.optimizer.step()
                total_loss += loss.item()
            
            # Emit epoch completion event
            context.emit({
                "kind": "epoch",
                "epoch": epoch,
                "metrics": {"loss": total_loss / len(self.train_loader)}
            })
            
            # Checkpoint every 5 epochs
            if epoch % 5 == 0:
                await context.checkpoint({
                    "epoch": epoch,
                    "model_state_dict": self.model.state_dict(),
                    "optimizer_state_dict": self.optimizer.state_dict()
                })

async def main():
    model = torch.nn.Linear(10, 2)
    optimizer = torch.optim.Adam(model.parameters())
    train_loader = [...]  # Your data loader
    
    run = TrainingRun()
    backend = PyTorchBackend(model, optimizer, train_loader)
    final_state = await run.execute(backend)
    
    # Review events
    for event in run.events():
        print(f"{event.kind}: {event.data}")

asyncio.run(main())

Training with Pause/Resume

from silver_run import TrainingRun, TrainingBackend
import asyncio

class LongRunningBackend(TrainingBackend):
    async def run(self, context: TrainingContext):
        for step in range(1000):
            if context.should_stop():
                break
            
            # Simulate training step
            await asyncio.sleep(0.01)
            
            # Emit progress
            if step % 100 == 0:
                context.emit({
                    "kind": "progress",
                    "step": step,
                    "total": 1000
                })

async def main():
    run = TrainingRun()
    backend = LongRunningBackend()
    
    # Start training in background
    training_task = asyncio.create_task(run.execute(backend))
    
    # Pause after some time
    await asyncio.sleep(0.5)
    run.pause()
    print("Training paused")
    
    # Resume after some time
    await asyncio.sleep(0.5)
    run.resume()
    print("Training resumed")
    
    # Wait for completion
    final_state = await training_task
    print(f"Training finished: {final_state.value}")

asyncio.run(main())

Custom Checkpoint Storage

from silver_run import TrainingRun, CheckpointStore, Checkpoint
import asyncio

class S3CheckpointStore(CheckpointStore):
    def __init__(self, bucket, prefix):
        self.bucket = bucket
        self.prefix = prefix
        self.checkpoints = {}
    
    async def save(self, checkpoint: Checkpoint):
        # Save to S3
        key = f"{self.prefix}/{checkpoint.id}"
        print(f"Saving checkpoint to S3: {key}")
        self.checkpoints[checkpoint.id] = checkpoint
    
    async def latest(self):
        if not self.checkpoints:
            return None
        return list(self.checkpoints.values())[-1]
    
    async def get(self, id: str):
        return self.checkpoints.get(id)

async def main():
    store = S3CheckpointStore("my-bucket", "checkpoints")
    run = TrainingRun(options=TrainingRunOptions(checkpoint_store=store))
    
    # Use custom checkpoint store
    await run.checkpoint({"model": "state"}, "checkpoint-1")
    latest = await run.latest_checkpoint()
    print(f"Latest checkpoint: {latest.id}")

asyncio.run(main())

Advanced Usage

Event Filtering and Analysis

from silver_run import TrainingRun

# Filter events by type
def get_epoch_events(run):
    return [e for e in run.events() if e.kind == "epoch"]

def get_error_events(run):
    return [e for e in run.events() if e.kind == "error"]

# Analyze training progression
def analyze_training(run):
    epoch_events = get_epoch_events(run)
    losses = [e.data.get("metrics", {}).get("loss") for e in epoch_events]
    
    if losses:
        print(f"Initial loss: {losses[0]}")
        print(f"Final loss: {losses[-1]}")
        print(f"Loss reduction: {losses[0] - losses[-1]}")

Multi-Run Experiments

from silver_run import TrainingRun
import asyncio

async def run_experiment(config):
    run = TrainingRun()
    backend = MyBackend(config)
    return await run.execute(backend)

async def main():
    configs = [
        {"learning_rate": 0.001},
        {"learning_rate": 0.01},
        {"learning_rate": 0.1}
    ]
    
    results = await asyncio.gather(*[
        run_experiment(config) for config in configs
    ])
    
    for config, result in zip(configs, results):
        print(f"LR {config['learning_rate']}: {result.value}")

asyncio.run(main())

Requirements

  • Python 3.10+

Development

# Install development dependencies
pip install -e ".[dev]"

# Run tests
pytest

# Run tests with coverage
pytest --cov=silver_run --cov-report=html

# Run linting
flake8 src/ tests/
mypy src/

Contributing

Contributions are welcome! Please see CONTRIBUTING.md for guidelines.

License

Apache-2.0 - see LICENSE file for details.

Evidence you can replay

Run timelines and checkpoints are durable evidence for the decision loop: persist the SVG returned by store.visualize(...), apply one recommended change, and compare the next run's metrics and events. This keeps debugging reproducible instead of relying on screenshots or memory.

Release files for silver-run 1.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for silver-run 1.2.0
File Size Uploaded
silver_run-1.2.0.tar.gz 3.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for silver-run 1.2.0
File Interpreter ABI Platform
silver_run-1.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 3.4 MB

Release files / silver_run-1.2.0.tar.gz

Download URL silver_run-1.2.0.tar.gz
Size 3.4 MB
Tags Source
SHA-256 checksum
How to use checksums
6c4a7bfedf9aa0581f760c9f6021f00fb51cb5c10b77b2f65838bae6c4d5c69a
BLAKE2b-256 checksum
How to use checksums
7a5ebdb277244dfcf25a91caee30f4c5ffeb8a8c5d5e971419bb7c3b8b85ffc8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release files / silver_run-1.2.0-py3-none-any.whl

Download URL silver_run-1.2.0-py3-none-any.whl
Size 14.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7ae479123f5f562cd80afa4f9c3b2e1557b908578f2822300aae2a026f1ac61d
BLAKE2b-256 checksum
How to use checksums
3d49f5ab8a66002aaf7a5647b408e79a1c78f051b301d1960ca2476195ae7d16
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 23, 2026.

Transparency log

Release history Release notifications | RSS feed

1.5.1

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.0

2 release files

This release

1.2.0 This release

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page