Skip to main content

miniVERL — lower verl experiment semantics onto one CUDA GPU

CI Build PyPI Python License

PyPI · Stable docs · Development docs · 中文

miniVERL runs a validated subset of verl experiment semantics on one NVIDIA GPU. Give it a resolved verl-shaped config and Parquet prompts; its versioned compiler produces a reviewable local plan, executes actor/reference/teacher/ reward roles in phases, and publishes portable PEFT and data artifacts.

The current development line covers PPO/GAE, GRPO, Dr.GRPO, RLOO and REINFORCE++ against official verl v0.9.0 (483b8a00). PPO uses an independent trainable critic with its own optimizer and checkpoint state; actor KL, entropy regularization, grouped rollouts, task rewards and a pinned sequence-classifier reward role share the same provenance model. The established verl v0.8.0 OPD profiles remain available for direct GKD and sampled-k1 distillation.

PyPI v0.13.0 is stable; main is development.

Start in 60 seconds

Install the CUDA-enabled PyTorch build that matches your machine, then:

python -m pip install "miniverl[train,cuda]"
miniverl data sample --task-rewards --rows 8 --out data/rl-prompts.parquet
miniverl import-verl --profile verl-rl-v0.9-single-gpu-v2 \
  --config examples/verl-rl-v0.9-single-gpu-ppo.yaml --out local-ppo.yaml
miniverl validate local-ppo.yaml
miniverl train local-ppo.yaml --dry-run

The importer writes local-ppo.import-report.json beside the native recipe. It accounts for every accepted source field and records how distributed resource settings lower to one process and one device. Remove --dry-run to load the pinned Qwen3-0.6B actor and execute the two-iteration example.

The [train,cuda] extra installs the ML and quantization stack; select PyTorch's CUDA build separately with the PyTorch installer. The single-GPU guide covers memory planning and the maintainer-measured RTX 4080 environment.

What a run gives you

  • A field-by-field compiler report. Experiment fields retain their meaning; physical distribution fields receive an explicit one-GPU lowering.
  • Policy-bound trajectories. Group/sample identity, generated-token spans, behavior log-probabilities and policy version travel together.
  • Auditable rewards and objectives. Reward components, group advantages, reference KL, clipping, entropy and update metrics are structured data.
  • Exact recovery. Transactional manifests and checkpoints restore policy, optimizer, cursor, RNG and reward identities before another rollout.
  • Portable outputs. PEFT adapters, safetensors, Parquet, resolved config and typed provenance can move into a larger workflow.

How it works

A resolved verl config compiles into a validated single-GPU plan; actor, critic, reference, teacher and reward roles run in phases and produce portable artifacts plus a readiness report.

One GPU is treated as a temporal scheduler for logical roles. RL generates complete prompt groups, scores outcomes, evaluates optional reference and reward-model roles, computes the pinned advantage estimator, and updates the actor. PPO adds a separate critic phase with clipped value updates. OPD uses the same trajectory and checkpoint foundations, with teacher targets supplying the learning signal.

Choose your path

Goal First command Primary artifact Next step
Run local RL miniverl import-verl --profile verl-rl-v0.9-single-gpu-v2 --config verl-ppo.yaml --out local.yaml native recipe + compatibility report RL quickstart
Run local OPD miniverl plan --profile verl-opd-v0.8-single-gpu-v1 --config verl-opd.yaml --out plan.json immutable execution plan OPD quickstart
Fit your GPU miniverl plan --config verl-opd.yaml --probe measured placement plan Hardware planning
Hand off artifacts miniverl export-verl --run runs/my-run --target-verl v0.9.0 --out scaleout actor, critic, Parquet + config bundle Compatibility contract

Native recipes also support SFT, DPO and offline KD, plus calculator, JSON navigation, read-only SQLite and custom tool environments.

Current capability matrix

Experiment surface Local status Contract
PPO / GAE semantically conformant independent critic, clipped value loss, actor/critic exact resume
GRPO / Dr.GRPO semantically conformant verl v0.9 group statistics and vanilla clipped policy loss
RLOO / REINFORCE++ semantically conformant verl v0.9 advantage and masking rules
Grouped n > 1 rollouts supported complete groups, stable sample seeds and behavior-policy identity
Task rewards and trained RM supported built-ins, environment verifier, trusted Python API, or pinned HF sequence classifier
Actor KL and entropy semantically conformant sampled-token reference KL and entropy regularization
Direct GKD / sampled-k1 OPD supported pinned verl v0.8 profiles with teacher targets
Ray, FSDP/FSDP2, Megatron, TP/PP/DP > 1 distributed-only these change physical scale, not the local objective

The generated v0.9 PPO compatibility record binds every example field to its source value, local target, classification and compiler-rule digest. miniverl import-verl accepts a resolved documented subset, rejects unknown algorithm-changing fields, and never substitutes a dataset or reward implementation.

Measured systems evidence

The published Qwen3-0.6B/1.7B OPD workload consumed 32 distinct prompts and completed eight current-policy updates at 3.1914 GiB peak reserved VRAM on an RTX 4080. A matched interruption after update four reproduced byte-identical trajectories, adapter and optimizer tensors. The SmolLM2-360M/1.7B workload completed the same shape at 1.4961 GiB. System records contain configs, hashes and phase timings.

Rollout Runtime v2 measured hf_cached and managed vLLM across 24 RTX 4080 cells spanning response length, sampling mode and n=1/4. See the runtime report. Evidence for the new RL family is published through the exact-wheel release qualification rather than presented as a task-quality comparison.

Research record

The scientific reports preserve their original outcomes and frozen inputs: calculator protocol, RecoveryBench, Alignment Lab, and the External Alignment Gate. They are scoped studies, separate from runtime qualification.

Compatibility boundary

miniVERL targets one local process and one NVIDIA CUDA GPU. The compiler distinguishes exact or conformant semantics, local physical lowering, distributed-only settings, feasible work not yet implemented, and technically unsupported inputs. The compatibility policy is the complete matrix; limitations collects architecture, measurement, security and generalization boundaries in one place. miniVERL is an independent Apache-2.0 project.

Development

git clone https://github.com/DaoyuanLi2816/mini-verl.git
cd mini-verl
python -m pip install -e ".[dev,train]"
pytest -q -m "not gpu and not network"

See CONTRIBUTING.md, CHANGELOG.md, CITATION.cff, the guide for verl users, SECURITY.md and the Apache-2.0 license.

Release files for miniverl 0.13.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for miniverl 0.13.0
File Size Uploaded
miniverl-0.13.0.tar.gz 1.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for miniverl 0.13.0
File Interpreter ABI Platform
miniverl-0.13.0-py3-none-any.whl Python 3 none any Details

Total release size: 2.3 MB

Release files / miniverl-0.13.0.tar.gz

Download URL miniverl-0.13.0.tar.gz
Size 1.7 MB
Tags Source
SHA-256 checksum
How to use checksums
886c2813a31fd58e84370bdf68bab765b8e92f8800cf54cc7f19ad84d9f32c91
BLAKE2b-256 checksum
How to use checksums
215220afdb49f509577502d83d86439bad82361cf13cf1082e503abb84fcb4a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.

Transparency log

Release files / miniverl-0.13.0-py3-none-any.whl

Download URL miniverl-0.13.0-py3-none-any.whl
Size 623.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1284d47c3b981ce1c74bc431254ca49e123ae144b2b49a728d42ff4efdff3a9e
BLAKE2b-256 checksum
How to use checksums
0dd1fc77a43a8b7114f852b0bfd8cf59a1ec77f70cbf136f4733b8009d7907df
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.16.0

2 release files

0.15.0

2 release files

This release

0.13.0 This release

2 release files

0.11.0

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page