Skip to main content

Game Learning Runtime

English | 简体中文

PyPI CI Python Rust License: MIT

Game Learning Runtime (GLR) is an agent-first, learner-neutral control plane for game learning. It gives an agent one stable JSON CLI and strict contracts to start an authorized game bridge, train while recording review video, query prior runs and world knowledge, research guides, revise bounded training and reward plans, and reproduce a verified model bundle in a new game instance.

Define observations, actions, masks, rewards, events, and episode boundaries once. The same adapter remains reusable with TorchRL, custom PPO or IMPALA, behavior cloning, offline datasets, evaluation, and automated QA.

A game-learning runtime designed to be operated by agents.

GLR is for games and test environments you own or are authorized to instrument. It does not include anti-cheat bypasses, stealth injection, or game-specific reverse-engineering code.

See GLR running

GLR collector running against the bundled synthetic counter adapter

This GIF is generated from a real local ContractEnvironment + SyncCollector run against GLR's explicitly synthetic counter adapter. It demonstrates the public contract without exposing a game account, machine path, process/window identity, or proprietary runtime data. See the showcase provenance and capture policy before contributing footage from a live adapter.

One boundary, many consumers

Game / simulator
      │
      ▼
Runtime adapter (C#, C++, Rust, Python, official API, ...)
      │
      ▼
GLR protocol + environment contract
      │
      ├── TorchRL
      ├── custom PPO / IMPALA
      ├── BC / DAgger / offline learning
      ├── recorder / replay
      └── evaluation / automated QA

Game adapters never import PPO, IMPALA, BC, or TorchRL. Learning code does not need to know whether the runtime is Unity, Unreal, Source, native, or a test simulator. GLR standardizes the data and lifecycle boundary, not one language, engine, transport, or algorithm.

What GLR standardizes

Contract Included today
Environment reset, truthful live attach, step, close, termination, truncation, episode and step identity
Data Recursive tensor specs, hybrid/parameterized actions, masks, events, rewards, immutable transitions
Bridge Capability negotiation, reset/step fencing, metadata deny-by-default, transport-neutral driver ports
Runtime Host Rust glr-hostd, bounded glr.host.v1 stdio, Python HostBridgeDriver, synthetic process smoke
Provider SDKs .NET Standard 2.0 C# contract for Unity/BepInEx and header-only C++20 contract for Unreal/native providers
Runtime integration Backward-compatible glr.runtime-integration.v2 profiles for source plugins, authorized loader plugins, and external attachments
Training config Strict glr.training.v1 knowledge sources, lifecycle policy, bridge requirements, auditable weighted rewards
Training safety Episode shaping budgets, mandatory terminal outcomes, failed-return ceilings, BC provenance gates, and checksummed demonstration artifacts
Collection Fixed-length or terminal-bounded unrolls for PPO/IMPALA plus glr.transition.v1 JSONL for BC/offline use
Integrations Optional Gymnasium, TorchRL 0.13, and model-neutral PyTorch BC/PPO/GAE/V-trace objectives
Validation Fail-closed contract wrapper and privacy-safe synthetic conformance profiles
Agent control plane Standalone Rust glr JSON CLI, strict project roles, bounded research/plan/train/evaluate goals, SQLite run queries, spatial knowledge transfer, managed binary/Skill updates
Review and supervised capture Concurrent project-owned H.264 capture with checksummed episode/step-to-frame index
Run review reports Offline interactive glr.run-report.v1 HTML with metrics, event timeline, route traces, progression, explicit PvP results, and checksummed media links
Agent workflow Separate glr-adapter-builder and glr-cli Skills for adapter construction versus operation
Model reproduction glr.model-bundle.v1 copies config, source/lock inputs, seeds, versions, weights, metrics, and SHA-256 provenance

Game-specific adapters, authenticated target-bound local provider transport, distributed actor transport, production trainers, and reference policies remain roadmap work.

Quick start

Install the matching standalone archive from the latest GitHub Release. Verify glr-{version}-{rust-target}.zip against SHA256SUMS, then put glr and glr-hostd on PATH. The archive also carries the glr-cli and glr-adapter-builder Skills.

The CLI is the primary deployment, integration, training, query, and playback entrypoint. It does not require Python:

glr --version
glr --project . --json doctor
glr --json update --check

Install the Python SDK only when a trainer, adapter, or learner imports it:

uv add game-learning-runtime

Choose optional integrations only where they are needed:

uv add "game-learning-runtime[torchrl]"
uv add "game-learning-runtime[torch]"
uv add "game-learning-runtime[gymnasium,torchrl]"

Agent-operated project

Once an authorized project provides glr-project.json and its fixed bridge roles, an agent can use the same machine-readable interface for the complete workflow:

glr --project . --json doctor
glr --project . --json runtime start
glr --project . --context config/contexts/ranked.toml --json train
glr --project . --json task list
glr --project . --json task run season --set profile=example/default
glr --project . --json goal run --goal goals/reach-destination.json
glr --project . --json query entities --world forest --kind shrine
glr --project . --json query routes --world forest --to-entity shrine.forest-1
glr --project . --json query research --tag navigation
glr --project . --json report build run-0123456789abcdef
glr --project . --json play --bundle artifacts/model-bundle

The project manifest owns exact executable paths, environment identity, data locations, and runtime/trainer/player/researcher/planner/evaluator/recorder roles. GLR validates and orchestrates those roles; it does not embed a game-specific launcher, scraper, or learning algorithm.

Projects can add strict, fixed-argv workflows in glr.toml. Use runner = "vx" for Python training tasks so VX owns runtime and virtual environment resolution while GLR owns validation, dependency ordering, timeout, logs, and execution receipts. See Extend GLR with declarative VX tasks.

Select invocation-specific configuration with a strict project-relative glr.run-context.v1 TOML file. GLR freezes the context and its declared inputs, passes the receipt to every configured role, and persists it with the run. Domain concepts such as a season or ruleset remain labels or VX task policy; they do not become core CLI commands. See Bind an invocation run context.

glr update --check only inspects the latest stable release. After an explicit update request, glr update verifies the exact target archive and SHA256SUMS, then updates glr, its sibling glr-hostd, and the project-owned Skills. It never changes game code, trainer dependencies, models, datasets, or project configuration. The check uses GitHub's public latest-release asset link without consuming anonymous REST API quota. Use --no-skills for binary-only maintenance or --skills-dir for an explicitly selected Skills directory. The former --yes form remains accepted for compatibility.

Library integration

Collect a learner-neutral unroll:

from game_learning_runtime import ContractEnvironment, SyncCollector
from game_learning_runtime.examples import CounterEnvironment, always_increment

environment = ContractEnvironment(CounterEnvironment(target=3))
collector = SyncCollector(environment, actor_id="local-actor")
unroll = collector.collect(always_increment, steps=16, policy_version=0)

print(len(unroll.transitions), unroll.total_reward)

For an authorized already-running game, advertise live-attach and select that lifecycle explicitly:

environment = ContractEnvironment(authorized_live_adapter)
collector = SyncCollector(environment, start_mode="attach")
unroll = collector.collect(policy, steps=128, stop_on_done=True)

Attach starts a fresh logical GLR episode at step zero. It never claims that the physical game world was reset or seeded.

From a user goal to verified replay

user goal + hard budgets
          │
          ▼
allowed rules / text / video + prior runs / spatial knowledge
          │
          ▼
research ──► plan ──► train + indexed video ──► authoritative evaluation
                ▲                                      │
                └──────── bounded revision ────────────┘
                                                       │
                                                       ▼
                                   queryable evidence + model bundle
                                                       │
                                                       ▼
                                      verified replay in a new instance

Training can run a project-owned small-window recorder concurrently and bind its H.264 output to episode/step IDs for human review and later supervised-data selection. Goal runs gather allowed official rules, text guides, video tutorials, and runtime traces through the configured researcher; adjust declarative reward plans between bounded trials; and stop only on matching authoritative runtime metrics. Guides and transferred knowledge remain advisory until fresh runtime evidence verifies them. See Operate GLR as an agent-first control plane.

Unity and Unreal integration lanes

GLR keeps one learner-facing contract while making the runtime boundary explicit:

  • Source or official extension SDK: run an engine plugin with semantic state, native actions, main-thread dispatch, controllable time, and a truthful physical reset.
  • Authorized binary-only runtime: attach externally through an official API, telemetry, replay, or bounded rendered observation/input seam. It defaults to real-time attach, exact target binding, input lease cleanup, and verified post-state.
  • Authorized mod-loader runtime: host a reviewed bounded-command adapter in BepInEx or UE4SS. It keeps truthful real-time attach, semantic observations, game-thread dispatch, exact loader/version provenance, and an empty-deny action vocabulary until game-specific handlers are reviewed.
from game_learning_runtime import EngineFamily, RuntimeIntegrationProfile

profile = RuntimeIntegrationProfile.for_source(EngineFamily.UNITY)
environment = profile.connect(authorized_driver)

Generate a Unity or Unreal adapter lane with the repository-owned Skill, then replace its synthetic semantics while keeping the contract tests green. See the engine runtime integration guide.

For no-source games that explicitly permit mods, GLR can generate a BepInEx 5 LTS Unity Mono host or a UE4SS 3.x Lua host:

vx python .agents/skills/glr-adapter-builder/scripts/scaffold_adapter.py `
  --output adapters/example_loader `
  --package example_loader `
  --environment-id example.loader-v1 `
  --engine unity `
  --access loader `
  --loader bepinex `
  --loader-version v5.4.23.5

The generated deployment command stages a checksummed payload; it never scans for or modifies a game installation. See the loader-plugin integration guide.

Runtime Host and engine providers

The implemented Runtime Host centralizes strict framing and lifecycle fencing without trying to replace engine bootstraps:

TorchRL / PPO / IMPALA / BC
  -> BridgeEnvironment -> HostBridgeDriver
  -> glr-hostd (Rust)
  -> C# Unity provider / C++ Unreal provider
  -> official plugin, BepInEx, UE4SS, or official mod SDK

Run the real cross-process conformance path and compile both provider contracts:

vx just host-smoke
vx just provider-sdk-check

glr-hostd currently ships only synthetic-counter over serialized stdio. It has a 1 MiB hard frame bound and never retries a mutating action, but it does not yet claim authenticated or target-bound IPC and cannot yet connect a live external C#/C++ provider. See the Runtime Host and provider SDK guide for the exact current boundary and Unity/Unreal implementation path.

Optional DeepSeek Harness

DeepSeekHarnessProvider is a separate, disabled-by-default control-plane provider for structured analysis, completion, or planning tasks. It does not alter learner-facing environment contracts, discover credentials, or grant runtime action authority. Explicit handlers receive bounded JSON tasks with permissions, deadlines, and idempotency keys; failures and timeouts are cached, and ordered events/state snapshots can be recovered through the optional LocalHarnessOrchestrator. See the DeepSeek Harness guide and configuration sample.

Reproducible local development

GLR pins Python, uv, just, rustup, Rust, and .NET SDK inputs. Local development and GLR's GitHub Actions execute the same recipes:

vx setup
vx just check
vx just ci

Knowledge and rewards as data

from game_learning_runtime import (
    EpisodeRewardGuard,
    RewardSignal,
    load_reward_safety_config,
    load_training_config,
)

config = load_training_config("training.json")
guard = EpisodeRewardGuard(config, load_reward_safety_config("reward-safety.json"))
reward = guard.compose([RewardSignal(name="progress", source="runtime", value=0.25)])
print(reward.total, reward.contributions)

Runtime telemetry should be authoritative; web guides and strategy priors should be advisory. Reward terms require authoritative sources by default. Adapters may also use the learner-neutral runtime evidence contracts to record settled route edges, stall/oscillation telemetry, modal navigation boundaries, and trajectory/recording lineage. These records never control the adapter, widen action masks, or replace authoritative reward and terminal evidence. Configuration is data only: GLR does not evaluate reward expressions as code. KnowledgeInjector additionally validates bounded glr.knowledge-snapshot.v1 payloads and selects stage/tag-relevant acquire, engage, upgrade, and avoid advice into an immutable learner context. The learner owns encoding; the context never gains action or reward authority. The episode guard caps positive shaping per step and episode, requires an authoritative terminal outcome, and prevents a failed episode from retaining a positive return. DemonstrationGate separately rejects policy self-imitation, failed episodes, and unknown provenance from BC by default. See training safety.

Build an adapter with the Agent Skill

The repository-owned glr-adapter-builder Skill gives a new agent a bounded workflow for:

  1. researching current game mechanics with source provenance;
  2. separating physical reset from truthful live attach;
  3. scaffolding a synthetic trainable seam;
  4. defining knowledge, reward budgets, and BC provenance policy;
  5. implementing fenced observations/actions through a runtime bridge; and
  6. validating conformance before a bounded authorized runtime trace.

The Skill never turns web strategy into runtime authority and never treats a synthetic test as live-game acceptance.

Give the Skill to your agent

The fastest path is to clone this repository and start Codex from its root. Codex discovers repository skills under .agents/skills automatically. Invoke the workflow explicitly in your prompt:

$glr-adapter-builder Create an authorized Unity adapter with source access. Scaffold a trainable environment, research manifest, reward configuration, and contract tests.

For an authorized binary-only runtime, say external access. For an authorized mod-enabled runtime, name BepInEx or UE4SS and the exact compatible upstream tag. The Skill chooses truthful attach, denies unknown actions, and refuses source-only capability claims.

To use the Skill from another repository, ask Codex's built-in installer to install it from GitHub:

$skill-installer install https://github.com/loonghao/GameLearningRuntime/tree/main/.agents/skills/glr-adapter-builder

Start a new agent turn after installation, then invoke $glr-adapter-builder. Pin the GitHub URL to a release tag or commit SHA when you need a reproducible team setup. Agents that implement the open Agent Skills standard can instead place the same glr-adapter-builder directory under the target repository's .agents/skills/ directory. See the official Codex Skills documentation.

The generated lane includes the environment skeleton, training.json, reward-safety.json, demonstration-policy.json, runtime-integration.json, a provenance-aware research manifest, tests, Agent instructions, a model-bundle smoke trainer, vx.toml, and a justfile.

Distribute the skills as an Agent Plugin

For plugin-capable agents, this repository also ships the self-contained game-learning-runtime-skills plugin. Its .codex-plugin/plugin.json follows the Agent Plugin manifest contract and its skills/ payload contains the same glr-adapter-builder skill as the repository source. Copy or archive the plugin directory without changing its internal layout; a compatible host should resolve each skill relative to the plugin's skills/ directory. Pin the repository to a release tag or commit SHA when sharing it with a team.

The repo-local .agents/plugins/marketplace.json exposes the plugin as game-learning-runtime-skills for hosts that support Agent Plugin marketplaces. Point that host at the repository marketplace, then install the plugin by that name; other Agent Skills-compatible hosts can consume the same plugin directory directly.

For Codex CLI, the equivalent commands are:

codex plugin marketplace add loonghao/GameLearningRuntime
codex plugin add game-learning-runtime-skills@game-learning-runtime

Before publishing a change, verify that the distributable payload has not drifted from the repository-owned skills:

vx uv run python scripts/package_agent_plugin.py --check

Maintainers can intentionally refresh the payload after editing a source skill with --sync, then rerun the check and the normal vx run check gates. The skill's bundled scripts and references are resolved from the installed skill root, so user-level plugin installs do not depend on a .agents/skills path in the consuming project.

Loader lanes also include bounded host source and a deployment manifest. From that generated directory, run:

vx setup
vx run check
vx run train
vx run reproduce

train emits a synthetic BC smoke model plus a self-contained checksummed reproduction environment. Replace the learner while preserving the glr.model-bundle.v1 gate; see reproducible model bundles.

For projects whose bridge already exists, use the separate glr-cli Skill. It teaches agents to configure and operate runtime, training capture, goal loops, history queries, knowledge transfer, verified playback, and explicitly authorized managed updates without changing adapter internals. Both Skills ship in every standalone GLR archive.

Distribute the skills as an Agent Plugin

For plugin-capable agents, this repository also ships the self-contained game-learning-runtime-skills plugin. Its .codex-plugin/plugin.json follows the Agent Plugin manifest contract and its skills/ payload contains the repository's glr-adapter-builder and glr-cli Skills. Copy or archive the plugin directory without changing its internal layout; a compatible host should resolve each Skill relative to the plugin's skills/ directory. Pin the repository to a release tag or commit SHA when sharing it with a team.

The repo-local .agents/plugins/marketplace.json exposes the plugin as game-learning-runtime-skills for hosts that support Agent Plugin marketplaces. Point that host at the repository marketplace, then install the plugin by that name; other Agent Skills-compatible hosts can consume the same plugin directory directly.

For Codex CLI, the equivalent commands are:

codex plugin marketplace add loonghao/GameLearningRuntime
codex plugin add game-learning-runtime-skills@game-learning-runtime

Before publishing a change, verify that the distributable payload has not drifted from the repository-owned Skills:

vx uv run python scripts/package_agent_plugin.py --check

Maintainers can intentionally refresh the payload after editing a source Skill with --sync, then rerun the check and the normal vx run check gates. The Skills' bundled scripts and references are resolved from each installed Skill root, so user-level plugin installs do not depend on a .agents/skills path in the consuming project.

TorchRL and custom learners

Use the optional TorchRL adapter:

from game_learning_runtime.examples import CounterEnvironment
from game_learning_runtime.integrations.torchrl import TorchRLEnvironment

env = TorchRLEnvironment(CounterEnvironment())
rollout = env.rollout(max_steps=32)

Or reuse the masked PPO objective in a custom PyTorch learner:

from game_learning_runtime.integrations.torch_objectives import ppo_loss

terms = ppo_loss(
    policy_logits=logits,
    actions=actions,
    old_log_prob=old_log_prob,
    advantages=advantages,
    values=values,
    value_targets=value_targets,
    action_mask=action_mask,
)
terms.loss.backward()

Hybrid and multi-head learners can instead call ppo_loss_from_log_prob with their summed per-head log-probability and entropy tensors. GLR shares the PPO mathematics without owning the project's distributions, policy network, action encoding, or optimizer.

Reuse the CI workflow

Any uv-managed Python repository can call GLR's public reusable workflow:

jobs:
  quality:
    uses: loonghao/GameLearningRuntime/.github/workflows/reusable-python-ci.yml@v0.17.0 # x-release-please-version
    with:
      python-versions: '["3.10", "3.12"]'
      sync-args: "--frozen --all-groups"
      lint-command: "uv run ruff check . && uv run mypy"
      test-command: "uv run pytest"

Pin a release tag or commit SHA in production. Release Please keeps the example tag synchronized with package releases. The reusable workflow receives no deployment secrets and only checks out and tests the calling repository.

Releases

Conventional Commits on main create or update a Release Please pull request. Merging that reviewed PR creates the tag and GitHub Release, verifies and builds the tagged source, attaches provenance, publishes the Python distributions to PyPI through Trusted Publishing, and attaches checksummed unified GLR archives for Linux, Windows, Intel macOS, and Apple Silicon. Each archive contains the standalone Rust CLI, matching Runtime Host, install manifest, and both GLR Skills; the release also includes the C# provider package. See the release runbook.

Documentation

See CONTRIBUTING.md for the development contract and SECURITY.md for private vulnerability reporting. GLR is licensed under the MIT License.

Goal-driven QA

Run bounded checks against a human-readable objective and get a self-contained report grouped by local date:

$env:PYTHONPATH = "src"
python -m game_learning_runtime.qa "inspect the whole game for bugs" `
  --project . `
  --check smoke python -c "print('adapter smoke ok')"

Each run is written to .glr-qa/YYYY-MM-DD/<UTC-time>/ with result.json and index.html. Checks are intentionally command based so an adapter can attach its own deterministic training, replay, or live-host probe while GLR keeps the goal, evidence, timeout, and report contract stable.

Goal-driven QA

Run bounded checks against a human-readable objective and get a self-contained report grouped by local date:

$env:PYTHONPATH = "src"
python -m game_learning_runtime.qa "inspect the whole game for bugs" --project . --check smoke python -c "print('adapter smoke ok')"

Each run is written to .glr-qa/YYYY-MM-DD/<UTC-time>/ with result.json and index.html. Checks are command based so an adapter can attach deterministic training, replay, or live-host probes while GLR keeps the goal, evidence, timeout, and report contract stable.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

game_learning_runtime-0.17.0.tar.gz (794.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

game_learning_runtime-0.17.0-py3-none-any.whl (201.9 kB view details)

Uploaded Python 3

File details

Details for the file game_learning_runtime-0.17.0.tar.gz.

File metadata

  • Download URL: game_learning_runtime-0.17.0.tar.gz
  • Upload date:
  • Size: 794.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for game_learning_runtime-0.17.0.tar.gz
Algorithm Hash digest
SHA256 2c09b9172773d498d44bfb1a16dcf4a06dcafd72e539bfb166fef28e623a5181
MD5 f388d158932e3e67452c91fac0869aea
BLAKE2b-256 8ae495ead0308c201225720f4fd4f8bc28351d89d458b3ddf9e7376b14422f3c

See more details on using hashes here.

Provenance

The following attestation bundles were made for game_learning_runtime-0.17.0.tar.gz:

Publisher: release.yml on loonghao/GameLearningRuntime

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file game_learning_runtime-0.17.0-py3-none-any.whl.

File metadata

File hashes

Hashes for game_learning_runtime-0.17.0-py3-none-any.whl
Algorithm Hash digest
SHA256 28c3e4cc1700dc971f50acd1b3b97e2d2a1201391d6632c5672c5debca807d1c
MD5 ff33f3ab8660f7ce24e375e68d17addb
BLAKE2b-256 75aef1506dae0837bae6685ac4e75eb1e7cec9660d35f82e8146f149796ef6e4

See more details on using hashes here.

Provenance

The following attestation bundles were made for game_learning_runtime-0.17.0-py3-none-any.whl:

Publisher: release.yml on loonghao/GameLearningRuntime

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.17.0 This release

2 files

0.16.0

2 files

0.15.0

2 files

0.14.0

2 files

0.13.2

2 files

0.13.1

2 files

0.13.0

2 files

0.12.1

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.1

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page