Skip to main content

silver-run

Silver local experiment tracking

Local experiment tracking that is transparent enough to debug with a text editor.

Python Version License CI Code Style

Backend-neutral ML run lifecycle, events, and checkpoints for Silver. A Python package designed for ML researchers who need flexible training orchestration across different frameworks.

vs. MLflow and W&B

Silver gives up hosted dashboards, team collaboration UI, and managed artifact storage. In exchange, runs are local JSON/JSONL you can inspect with cat or jq, replay without a server, and keep without vendor lock-in.

Replay the experiment visually

Neural-network inputs, hidden layers, activations, gradients, prediction, and training health

restored = store.load(run.id)
open("run-timeline.svg", "w", encoding="utf-8").write(restored.to_svg())

# Or render from the store and save the returned SVG text:
open("run-timeline.svg", "w", encoding="utf-8").write(store.visualize(run.id))

The timeline uses persisted event timestamps and recorded loss curves, so it is replayable after the process exits. See the visual evidence model.

Installation

pip install silver-run

Quick Start

Durable tracking without a server

from silver_run import (
    FileCheckpointStore, LocalRunStore, TrainingRun, TrainingRunOptions,
)

store = LocalRunStore(".silver/runs")
run = TrainingRun(TrainingRunOptions(
    metadata={"model": "tabular-v1", "dataset": "customers-2026"},
    run_store=store,
    checkpoint_store=FileCheckpointStore(".silver/checkpoints"),
))

run.start()
run.emit("epoch", {"epoch": 1, "metrics": {"loss": 0.42}})
run.complete()

restored = store.load(run.id)
print(restored.state, restored.metrics, restored.duration)

Run manifests are atomic JSON; events are append-only JSONL; checkpoint IDs are path-safe; checkpoint payloads are explicit JSON instead of unsafe pickle.

from silver_run import TrainingRun, TrainingBackend, TrainingContext
import asyncio

class MyBackend(TrainingBackend):
    async def run(self, context: TrainingContext):
        for epoch in range(10):
            if context.should_stop():
                break
            # Your training logic here
            context.emit({
                "kind": "epoch",
                "epoch": epoch,
                "metrics": {"loss": 0.5 - epoch * 0.05}
            })
            await asyncio.sleep(0.1)

async def main():
    run = TrainingRun()
    backend = MyBackend()
    final_state = await run.execute(backend)
    print(f"Run finished with state: {final_state.value}")

asyncio.run(main())

Features

  • Training Lifecycle Management: Full state machine (created, running, paused, stopped, cancelled, completed, failed)
  • Event Logging: Comprehensive event tracking with timestamps for training observability
  • Checkpoint Management: Pluggable storage backends for model checkpointing
  • Backend-Agnostic: Works with PyTorch, TensorFlow, JAX, or any custom training framework
  • Async/Await Support: Modern Python async patterns for concurrent training
  • Pause/Resume: Control long-running training jobs with pause and resume functionality
  • Type Safety: Full type hints for better IDE support and fewer bugs

Use Cases

PyTorch Training Integration

from silver_run import TrainingRun, TrainingBackend, TrainingContext
import torch
import asyncio

class PyTorchBackend(TrainingBackend):
    def __init__(self, model, optimizer, train_loader):
        self.model = model
        self.optimizer = optimizer
        self.train_loader = train_loader
    
    async def run(self, context: TrainingContext):
        for epoch in range(10):
            if context.should_stop():
                break
            
            self.model.train()
            total_loss = 0
            
            for batch_idx, (data, target) in enumerate(self.train_loader):
                self.optimizer.zero_grad()
                output = self.model(data)
                loss = torch.nn.functional.cross_entropy(output, target)
                loss.backward()
                self.optimizer.step()
                total_loss += loss.item()
            
            # Emit epoch completion event
            context.emit({
                "kind": "epoch",
                "epoch": epoch,
                "metrics": {"loss": total_loss / len(self.train_loader)}
            })
            
            # Checkpoint every 5 epochs
            if epoch % 5 == 0:
                await context.checkpoint({
                    "epoch": epoch,
                    "model_state_dict": self.model.state_dict(),
                    "optimizer_state_dict": self.optimizer.state_dict()
                })

async def main():
    model = torch.nn.Linear(10, 2)
    optimizer = torch.optim.Adam(model.parameters())
    train_loader = [...]  # Your data loader
    
    run = TrainingRun()
    backend = PyTorchBackend(model, optimizer, train_loader)
    final_state = await run.execute(backend)
    
    # Review events
    for event in run.events():
        print(f"{event.kind}: {event.data}")

asyncio.run(main())

Training with Pause/Resume

from silver_run import TrainingRun, TrainingBackend
import asyncio

class LongRunningBackend(TrainingBackend):
    async def run(self, context: TrainingContext):
        for step in range(1000):
            if context.should_stop():
                break
            
            # Simulate training step
            await asyncio.sleep(0.01)
            
            # Emit progress
            if step % 100 == 0:
                context.emit({
                    "kind": "progress",
                    "step": step,
                    "total": 1000
                })

async def main():
    run = TrainingRun()
    backend = LongRunningBackend()
    
    # Start training in background
    training_task = asyncio.create_task(run.execute(backend))
    
    # Pause after some time
    await asyncio.sleep(0.5)
    run.pause()
    print("Training paused")
    
    # Resume after some time
    await asyncio.sleep(0.5)
    run.resume()
    print("Training resumed")
    
    # Wait for completion
    final_state = await training_task
    print(f"Training finished: {final_state.value}")

asyncio.run(main())

Custom Checkpoint Storage

from silver_run import TrainingRun, CheckpointStore, Checkpoint
import asyncio

class S3CheckpointStore(CheckpointStore):
    def __init__(self, bucket, prefix):
        self.bucket = bucket
        self.prefix = prefix
        self.checkpoints = {}
    
    async def save(self, checkpoint: Checkpoint):
        # Save to S3
        key = f"{self.prefix}/{checkpoint.id}"
        print(f"Saving checkpoint to S3: {key}")
        self.checkpoints[checkpoint.id] = checkpoint
    
    async def latest(self):
        if not self.checkpoints:
            return None
        return list(self.checkpoints.values())[-1]
    
    async def get(self, id: str):
        return self.checkpoints.get(id)

async def main():
    store = S3CheckpointStore("my-bucket", "checkpoints")
    run = TrainingRun(options=TrainingRunOptions(checkpoint_store=store))
    
    # Use custom checkpoint store
    await run.checkpoint({"model": "state"}, "checkpoint-1")
    latest = await run.latest_checkpoint()
    print(f"Latest checkpoint: {latest.id}")

asyncio.run(main())

Advanced Usage

Event Filtering and Analysis

from silver_run import TrainingRun

# Filter events by type
def get_epoch_events(run):
    return [e for e in run.events() if e.kind == "epoch"]

def get_error_events(run):
    return [e for e in run.events() if e.kind == "error"]

# Analyze training progression
def analyze_training(run):
    epoch_events = get_epoch_events(run)
    losses = [e.data.get("metrics", {}).get("loss") for e in epoch_events]
    
    if losses:
        print(f"Initial loss: {losses[0]}")
        print(f"Final loss: {losses[-1]}")
        print(f"Loss reduction: {losses[0] - losses[-1]}")

Multi-Run Experiments

from silver_run import TrainingRun
import asyncio

async def run_experiment(config):
    run = TrainingRun()
    backend = MyBackend(config)
    return await run.execute(backend)

async def main():
    configs = [
        {"learning_rate": 0.001},
        {"learning_rate": 0.01},
        {"learning_rate": 0.1}
    ]
    
    results = await asyncio.gather(*[
        run_experiment(config) for config in configs
    ])
    
    for config, result in zip(configs, results):
        print(f"LR {config['learning_rate']}: {result.value}")

asyncio.run(main())

Requirements

  • Python 3.10+

Development

# Install development dependencies
pip install -e ".[dev]"

# Run tests
pytest

# Run tests with coverage
pytest --cov=silver_run --cov-report=html

# Run linting
flake8 src/ tests/
mypy src/

Contributing

Contributions are welcome! Please see CONTRIBUTING.md for guidelines.

License

Apache-2.0 - see LICENSE file for details.

Related Packages

Evidence you can replay

Run timelines and checkpoints are durable evidence for the decision loop: persist the SVG returned by store.visualize(...), apply one recommended change, and compare the next run's metrics and events. This keeps debugging reproducible instead of relying on screenshots or memory.

Release files for silver-run 1.5.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for silver-run 1.5.1
File Size Uploaded
silver_run-1.5.1.tar.gz 3.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for silver-run 1.5.1
File Interpreter ABI Platform
silver_run-1.5.1-py3-none-any.whl Python 3 none any Details

Total release size: 3.4 MB

Release files / silver_run-1.5.1.tar.gz

Download URL silver_run-1.5.1.tar.gz
Size 3.4 MB
Tags Source
SHA-256 checksum
How to use checksums
62731a9dd4859b704b41b4f32c706b29fc7a28148faad0ac8f12bad0aee54508
BLAKE2b-256 checksum
How to use checksums
b117caddbb4b43ef222108257d4ea6e3acd28eb41379a71bfe49b04598ead34d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release files / silver_run-1.5.1-py3-none-any.whl

Download URL silver_run-1.5.1-py3-none-any.whl
Size 15.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2f95568f53d92008a44176f9c6209868cbaf21326a9eb64c11fed6be1a169191
BLAKE2b-256 checksum
How to use checksums
8a930e1b598c34b554ebdec87202e85d89a0512c01c50ee15d1788a6ca19676c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 24, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.5.1 This release

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page