Skip to main content

deep-learning-core

Reusable deep learning framework core.

deep-learning-core contains the vendor-neutral training framework that can be reused across many experiment repositories. It is intended to be the public base package, while optional integrations such as Azure are layered on through extras and companion extension packages.

Current public release: deep-learning-core==0.0.26. Current development version: 0.0.27.

Compatible companion package floors:

  • deep-learning-azure>=0.0.18,<0.1
  • deep-learning-mlflow>=0.0.11,<0.1
  • deep-learning-robotics>=0.0.1,<0.1 with deep-learning-core>=0.0.27,<0.1
  • deep-learning-wandb>=0.0.12,<0.1

Unreleased: 0.0.27

  • episode managers provide generic RL summaries and selective, complete trajectory capture alongside the existing metric-manager system
  • scalar and same-step vector environments share one collector contract with preserved terminal observations and per-lane episode identity
  • Q-learning, DQN, PPO, and SAC now consume vector collection natively; neural policies perform batched inference and replay/rollout storage preserves the correct algorithm-specific scheduling and boundary semantics
  • replay insertion and PPO rollout/GAE computation operate on real batches, while scalar custom trainers remain compatible through the original hooks
  • extension packages can import and register environments without depending on dl-core import order

What's New in 0.0.26?

  • Gymnasium-compatible environments can now be registered, discovered, and created through a first-class environment contract
  • RLTrainer provides an episode-based lifecycle, deterministic evaluation, RL callback hooks, and resumable algorithm checkpoints alongside EpochTrainer
  • QLearningTrainer adds tabular epsilon-greedy learning for finite discrete Gymnasium environments
  • DQNTrainer adds replay-based discrete control with target networks, Double-DQN targets, and a built-in MLP Q-network
  • PPOTrainer adds clipped on-policy optimization, GAE, and discrete or bounded-continuous actor-critic policies
  • SACTrainer adds replay-based bounded-continuous control with twin critics, Polyak targets, and optional automatic entropy tuning
  • dl-run --validate-only now resolves RL environments, models, optimizers, and callbacks without resetting or stepping an environment
  • dl-core add trainer MyPolicy --base rltrainer scaffolds custom algorithms against the episode-oriented lifecycle
  • the local metric callback records RL episode, update, and evaluation metrics
  • local component loading now cleans up registrations and import paths between projects, while registry lookups prefer the most specific matching prefix
  • configuration validation rejects malformed root and component structures more consistently

Install

Install from PyPI:

pip install deep-learning-core

Install with Azure support:

pip install "deep-learning-core[azure]"

Install with local MLflow support:

pip install "deep-learning-core[mlflow]"

Install with W&B support:

pip install "deep-learning-core[wandb]"

Install with multiple variants:

pip install "deep-learning-core[azure,wandb]"

Install in a uv project:

uv add deep-learning-core

deep-learning-core intentionally ships with the full public runtime dependencies, including torch, torchvision, and opencv-python-headless. The Azure extra pulls in deep-learning-azure, which pins the Azure package versions used by the validated Azure packaging stack. The MLflow extra pulls in deep-learning-mlflow for local MLflow tracking. The W&B extra pulls in deep-learning-wandb and leaves the wandb package itself unpinned.

Package Variants

  • deep-learning-core: local training, local sweeps, local sweep analysis, and the experiment scaffold
  • deep-learning-core[azure]: adds the public dl-azure package for Azure execution and Azure dataset foundations
  • deep-learning-core[mlflow]: adds the public dl-mlflow package for local MLflow integration
  • deep-learning-core[wandb]: adds the public dl-wandb package for Weights & Biases integration
  • deep-learning-robotics: adds fast scalar and vector 2D MAPF environments, metrics, and episode media as a separately installed companion package

The extension packages stay separate so the base package remains reusable and vendor-neutral.

You can also install the companion packages directly when you want a specific integration without using extras:

pip install deep-learning-azure
pip install deep-learning-mlflow
pip install deep-learning-robotics
pip install deep-learning-wandb

Scope

  • Base abstractions and registries
  • Built-in accelerators, callbacks, criterions, metrics, and schedulers
  • The standard trainer and standard dataset flow
  • Built-in augmentations
  • Local execution and sweep orchestration
  • Local sweep analysis from saved artifact summaries
  • Experiment repository scaffolding via dl-init

Out Of Scope

  • Azure ML wiring unless the Azure extra is installed
  • Workspace or datastore conventions
  • Experiment-specific datasets, models, and trainers
  • User-owned configs and private data

Quick Start

uv run dl-core list
uv run dl-init --name my-exp --root-dir .

To initialize the current directory in place, omit --name:

uv run dl-init --root-dir .

The generated experiment repository is the normal consumer entry point. Inside that repository, run uv sync, then run:

uv run dl-run --config configs/base.yaml --validate-only
uv run dl-inspect-dataset --config configs/base.yaml
uv run dl-smoke --config configs/base.yaml
cp configs/base.yaml experiments/debug.yaml
uv run dl-run --config experiments/debug.yaml --validate-only
uv run dl-run --config experiments/debug.yaml
uv run dl-sweep experiments/lr_sweep.yaml --preview
uv run dl-sweep experiments/lr_sweep.yaml --only "*seed_2025*"
uv run dl-sweep experiments/lr_sweep.yaml
uv run dl-analyze --sweep experiments/lr_sweep.yaml
uv run dl-sync --sweep experiments/lr_sweep.yaml --artifacts
uv run dl-analyze --sweep experiments/lr_sweep.yaml --name pareto_eer
uv run dl-analyze --sweep experiments/lr_sweep.yaml --compare latest
uv run dl-analyze --sweep experiments/lr_sweep.yaml --metric test/eer --mode min

New local runs use the flattened artifact layout:

  • artifacts/runs/<run_name>/...
  • artifacts/sweeps/<sweep_name>/<run_name>/...

dl-core does not create a latest symlink for these run directories. Use the concrete run directory names directly.

First Run Workflow

If you are starting from scratch, the minimum path is:

pip install deep-learning-core
uv run dl-init --name my-exp --root-dir .
cd my-exp
uv sync

Then:

  1. open these generated files first:
    • src/datasets/my_exp.py
    • configs/base.yaml
    • scripts/temporary/test_dataset.py
    • scripts/temporary/test_model.py
    • experiments/lr_sweep.yaml
    • AGENTS.md
    • CLAUDE.md
  2. implement the generated dataset wrapper under src/datasets/my_exp.py
  3. adjust configs/base.yaml so it points at the dataset/model/trainer you want and set the shared reproducibility defaults you need (seed and deterministic). Keep concrete single-run configs in experiments/, including debug and baseline runs.
  4. smoke-check the generated helpers:
uv run python scripts/temporary/test_dataset.py
uv run python scripts/temporary/test_model.py
  1. start with:
uv run dl-run --config configs/base.yaml --validate-only
uv run dl-inspect-dataset --config configs/base.yaml
uv run dl-smoke --config configs/base.yaml
cp configs/base.yaml experiments/debug.yaml
uv run dl-run --config experiments/debug.yaml --validate-only
uv run dl-run --config experiments/debug.yaml

Once that works, move on to:

uv run dl-sweep experiments/lr_sweep.yaml --preview
uv run dl-sweep experiments/lr_sweep.yaml
uv run dl-analyze --sweep experiments/lr_sweep.yaml
uv run dl-analyze --sweep experiments/lr_sweep.yaml \
  --metric test/eer --mode min \
  --metric test/accuracy --mode max \
  --rank-method rank-sum

dl-sweep --preview prints the expanded run matrix without creating configs or starting runs. Use --export sweep_preview.csv or --export sweep_preview.json when you want to save that expansion for review. Use --only and --skip with glob patterns when you want to execute or preview only a subset of generated run names.

dl-inspect-dataset preserves the configured split behavior, but forces single-process loading so you can quickly verify split sizes and inspect one collated batch without starting a trainer.

dl-sync --sweep ... --artifacts syncs tracked run outputs into the local repo. Backends that already write local artifacts simply refresh the tracker paths. Remote-backed integrations can download the run bundle and patch sweep_tracking.json with the resolved local artifact paths.

dl-analyze defaults to ranking by test/accuracy with max. You can make that explicit or override it with one or more --metric / --mode pairs and choose lexicographic, rank-sum, or pareto ranking.

For Azure-backed sweeps, dl-analyze fetches only the requested metric histories instead of downloading every tracked metric history. Those fetched histories are cached in experiments/<sweep_name>/analysis_cache.json. Use --force to ignore and refresh that cache. Reports are written under experiments/<sweep_name>/analysis/ as v1.md, v2.md, and so on unless you pass --name. A matching JSON report is always written next to each Markdown report, and --compare latest or --compare v1 compares the current ranking against a saved report.

EMA Checkpoints

When EMA is enabled with save_in_checkpoint: true, each checkpoint stores:

  • models_state_dict: normal model weights for training resume
  • ema_state_dict: EMA bookkeeping and shadow-parameter state for trainer-side resume
  • ema_models_state_dict: a full drop-in model state dict with EMA parameters and the original model buffers preserved

That means evaluator-side code can load:

  • checkpoint["models_state_dict"]["main"] for normal weights
  • checkpoint["ema_models_state_dict"]["main"] for EMA weights

without needing to reconstruct EMA state manually.

Post-Training Checkpoint Hooks

After a successful training loop, BaseTrainer.run() calls select_checkpoint() and passes the returned path into post_training(checkpoint_path). This hook runs before run-analysis artifacts are persisted and before tracking callbacks upload finalized artifacts.

The default select_checkpoint() implementation keeps the existing checkpoint callback behavior authoritative: it returns final best.pth when present, falls back to final latest.pth, and returns None if no checkpoint exists. Override select_checkpoint() when a project needs custom single- or multi-metric model selection, and override post_training() for completed-run evaluation, export, or report generation.

If Azure support is installed, uv run dl-init --with-azure will also scaffold Azure-ready config placeholders and azure-config.json.

If local MLflow support is installed, uv run dl-init --with-mlflow will also scaffold an mlflow callback block and local tracking defaults.

If W&B support is installed, uv run dl-init --with-wandb will also scaffold a wandb callback block, W&B tracking defaults, and .env.example.

Companion Packages

Scaffold Commands

Each dl-core add ... command creates the new module and updates the matching local package __init__.py export list under src/.

uv run dl-core describe ... now also shows a minimal YAML snippet for common config-backed component types such as datasets, models, callbacks, optimizers, and trainers.

Common local component scaffolds:

uv run dl-core add model MyResNet
uv run dl-core add trainer MyTrainer
uv run dl-core add trainer MyPolicy --base rltrainer
uv run dl-core add callback MyMetrics
uv run dl-core add metric_manager MyManager
uv run dl-core add episode_manager MyEpisodeManager
uv run dl-core add sampler MySampler
uv run dl-core add optimizer MyOptimizer
uv run dl-core add scheduler MyScheduler
uv run dl-core add criterion MyLoss
uv run dl-core add augmentation MyAugmentation
uv run dl-core add metric MyMetric
uv run dl-core add executor MyExecutor

Default-base scaffolds for augmentations, metrics, metric managers, criterions, models, and executors now start with ready-to-edit method stubs instead of empty wrapper subclasses.

Sweep scaffolds are supported too:

uv run dl-core add sweep DebugSweep
uv run dl-core add sweep AzureEval --tracking azure_mlflow
uv run dl-core add sweep MlflowBaseline --tracking mlflow
uv run dl-core add sweep WandbAblation --tracking wandb

Generated sweep files:

  • live under experiments/
  • extend ../configs/base_sweep.yaml
  • include runnable defaults in fixed
  • start with grid: {}
  • default the tracker experiment destination to the repository root name unless tracking.experiment_name overrides it
  • let the tracker derive sweep grouping from the filename unless tracking.sweep_name overrides it

Project-specific criterions, optimizers, and schedulers can still be added later with uv run dl-core add ... when they are actually needed.

You can inspect registered components and built-in base classes directly from the CLI:

uv run dl-core list
uv run dl-core list sampler
uv run dl-core list metric_manager --json
uv run dl-core describe dataset my_dataset --root-dir .
uv run dl-core describe model my_resnet --root-dir .
uv run dl-core describe class dl_core.core.FrameWrapper
uv run dl-core describe class dl_azure.datasets.AzureComputeMultiFrameWrapper
uv run dl-core describe dataset my_dataset --root-dir . --json

The built-in sampler list now includes label, which balances samples by any metadata key using either undersample or oversample.

Example sampler config:

dataset:
  sampler:
    label:
      key: attack
      mode: undersample

The describe command shows:

  • resolved class and registered names
  • constructor signature
  • inheritance chain
  • docstring
  • declared properties
  • class-level attributes
  • public methods defined on the class

It does not discover instance attributes created dynamically inside __init__ without constructing the class.

Scaffolds can target a specific base when you need one:

uv run dl-core add dataset MyDataset
uv run dl-core add dataset FrameSet --base frame
uv run dl-core add dataset TextSet --base text_sequence
uv run dl-core add dataset ActSet --base adaptive_computation
uv run dl-core add callback EpochLogger --base metric_logger
uv run dl-core add metric_manager PadMetrics --base standard
uv run dl-core add optimizer AdamwWrapper --base adamw
uv run dl-core add scheduler CosineWrapper --base cosine

Built-in callbacks include dataset_refresh, which rebuilds selected dataset splits at epoch boundaries. Example:

callbacks:
  dataset_refresh:
    refresh_frequency: 1
    splits: [train]

When dl-azure is importable, the dataset scaffold also exposes Azure bases:

uv run dl-core add dataset AzureFrames --base azure_compute_frame
uv run dl-core add dataset AzureSeq --base azure_compute_multiframe
uv run dl-core add dataset AzureStream --base azure_streaming
uv run dl-core add dataset AzureStreamSeq --base azure_streaming_multiframe

Plain deep-learning-core currently exposes dataset bases for:

  • BaseWrapper
  • FrameWrapper
  • TextSequenceWrapper
  • AdaptiveComputationDataset

TextSequenceWrapper adds sequence-aware batch padding for tokenized inputs. AdaptiveComputationDataset adds per-class sample stream helpers for adaptive-time computation trainers. Multiframe dataset bases are still provided through dl-azure.

Releases

  • Publish is the production workflow for PyPI.
  • Trusted publishing is configured through GitHub Actions environments rather than long-lived API tokens.
  • The publish action may upload digital attestations alongside the package. That is expected behavior from pypa/gh-action-pypi-publish.
  • Package metadata keeps runtime dependencies unpinned, so the consuming environment resolves the latest compatible public releases.

Documentation

License

MIT. See LICENSE.

Development Validation

uv run --extra dev pytest
uv run --extra dev ruff check src tests
uv run python -m compileall src/dl_core
uv build --no-sources

Release files for deep-learning-core 0.0.27

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for deep-learning-core 0.0.27
File Size Uploaded
deep_learning_core-0.0.27.tar.gz 557.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for deep-learning-core 0.0.27
File Interpreter ABI Platform
deep_learning_core-0.0.27-py3-none-any.whl Python 3 none any Details

Total release size: 856.0 kB

Release files / deep_learning_core-0.0.27.tar.gz

Download URL deep_learning_core-0.0.27.tar.gz
Size 557.9 kB
Tags Source
SHA-256 checksum
How to use checksums
47507a7c962cac9f6c9926785b056fc47c9c8fd9f4260d1297af070adfddd2b5
BLAKE2b-256 checksum
How to use checksums
cdbb360a2e2e41da56d2c55384953340919defcc311865c6bc24c43fdf3c70af
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.

Transparency log

Release files / deep_learning_core-0.0.27-py3-none-any.whl

Download URL deep_learning_core-0.0.27-py3-none-any.whl
Size 298.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4cd4b2a764513b74f54ec41fb30b0a7886e83fc0f0fadb6ef863f0ff03f97fba
BLAKE2b-256 checksum
How to use checksums
e5435ee78643635f9c66a727ccab40591ba6c995984cfc14d97e2d1395c55346
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 25, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.6

2 release files

0.1.5

2 release files

0.1.0

2 release files

0.0.35

2 release files

0.0.34

2 release files

0.0.33

2 release files

0.0.32

2 release files

0.0.31

2 release files

0.0.30

2 release files

0.0.29

2 release files

0.0.28

2 release files

This release

0.0.27 This release

2 release files

0.0.26

2 release files

0.0.25

2 release files

0.0.24

2 release files

0.0.23

2 release files

0.0.20

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page