Skip to main content

TopoGym

CI PyPI License Python pre-commit Ruff Discord

Gridworld environments with certified topology, for exploration research.

TopoGym is a Gymnasium environment library where the shape of every world — its chambers (sort of like rooms), decoys (filled rooms, large icebergs, or other blatant & large obstructions), and identifications (going in a circle or going in a circle while twisting in space) — is known exactly: computed from the free-space cell complex by GUDHI and cross-checked against the analytic expectation at generation time. Everything is deterministic up to seeds, end to end. We provide benchmarks for reinforcement learning researchers to test how good their agents are at exploring complex environment shapes.


EnvironmentalIceShip

ClownChase

SpaceWarp

DontFall

SearchRescue

BankRobber

Nested3-50

TopTorus-50

Full gallery and per-environment documentation: docs/envs/ · docs/environments/.

Environments

One benchmark, TopoGym-v1, in three slices under a universal interface (egocentric Discrete(3) turn-left / turn-right / forward actions with an occluded egocentric view by default — the rendered agent is a MiniGrid-style arrow, so its heading is always visible; actions="fourway" opts into Discrete(4) screen-direction actions with the universal (x, y) + 16-slot texture vector):

slice families axis status
GridWorld2D Dilution, Chambers2, ChamberCount, Decoys, Shape{Sq,Ci,Tr,St}, Nested, GiveUp, Bottleneck, Maze world size, chamber/decoy count, shape, nesting depth, corridor length, braiding 🟢 Live; Beta
Texture IceShip, EnvironmentalIceShip, Ladders, BankRobber, DontFall, SpaceWarp, ClownChase, SearchRescue semantic local signals — and exactly where they fail 🟢 Live; Beta
Top TopPlane, TopCylinder, TopMobius, TopTorus, TopKlein, TopRP2 global topology with zero local signal 🟢 Live; Beta

Every id is stable: gym.make("TopoGym/{Family}-{size}-v0", seed=n). Details per family: docs/environments/.

Install

pip install topogym              # deps: gymnasium, numpy, gudhi
pip install "topogym[play]"      # + pygame, for keyboard play

Development: git clone, then pip install -e ".[testing,play,assets]".

Quick start

import gymnasium as gym
import topogym  # registers the TopoGym/* ids

env = gym.make("TopoGym/Decoys4-50-v0", seed=3)
obs, info = env.reset(seed=0)
info["topology"]["betti_z2"]         # [1, 4, 0] — doors walkable
info["topology"]["betti_z2_sealed"]  # [2, 5, 0] — doors count as walls

Episodes truncate after a pre-determined horizon — the larger of 1.2 * max(W, H) and 3x the turn-aware optimal route, so the goal is always reachable with room to wander; the goal pays +1 terminal reward by default (reward_mode="sparse") and sits inside a designated chamber. reward_mode="none" for pure exploration, "coverage", "deceptive"; goal=False removes the goal; p_slip=0.1 for sticky-action noise; complex="rips" swaps the homology backend to a Vietoris–Rips complex on the quotient metric.

Compose custom worlds with the fluent spec API:

from topogym.spec import Torus

env = Torus(15).holes(3).chambers(1).compile(seed=7)

Measure what an agent actually discovered — from its own trajectory:

from topogym.tda import ExplorationTracker
from topogym.stats import StatsRecorder

env = StatsRecorder(gym.make("TopoGym/Nested3-50-v0", seed=1))
tracker = ExplorationTracker(env)
tracker.reset(seed=0)
# ... run your policy ...
tracker.summary()      # discovery-time persistence: real vs transient loops
env.episodes           # per-episode rows: return, coverage, chamber entries

Archive-style (Go-Explore) resets are built in:

env = gym.make("TopoGym/Maze-100-v0", seed=1, teleport=True)
env.reset(options={"teleport": (12, 40)})  # any previously visited cell

Benchmarks

benchmark what it tests manifest splits RND+PPO ICM+PPO Go-Explore status
TopoGym-v1 topological navigation against decoys, chambers, distractions, and orientation in 2D space croissant.json · docs/manifest.csv tune · train · val · test · size-extrapolation · family-holdout TBD TBD TBD 🟠 in development

The random floor is measured under exactly that protocol: across all 189 hold-out instances (50 episodes each, 9,450 episodes) a uniform-random policy reaches the goal 0% of the time and uncovers 11.0% of the reachable space. Nothing in this benchmark falls out of undirected exploration, and coverage — not steps-to-goal — is what separates methods until one of them solves something.

Baselines report median steps to find the goal, with a 95% bootstrap confidence interval, over the hold-out split. Full metrics, per-slice breakdowns, and the discovery-curve figures live in BENCHMARKS.md.

Every baseline consumes the splits the same way — hyperparameters on tune, gradients on train, early stopping on val, and test read once at the end — enforced by Baseline.run() rather than left to each algorithm. The algorithms themselves are Ray RLlib's; TopoGym does not reimplement PPO. A variant such as RND or ICM subclasses PPOBaseline and overrides one hook, and an algorithm that never uses PPO (Go-Explore explores randomly by default) implements the same small interface.

--group decides what one policy is trained on, and therefore what is being measured. family (the default) trains a policy per family across its sizes and seeds, in the spirit of Procgen's train-on-levels, test-on-held-out-levels design; unit is the strictest per-world version; all asks instead for a single general explorer across every family at once.

pip install topogym[benchmarks]
python scripts/benchmarks/run_baselines_gridworld_v1_benchmark.py \
    --baselines random,ppo --group family --num-env-runners 16
python scripts/benchmarks/run_baselines_gridworld_v1_benchmark.py --smoke   # pipeline check

Environment stepping is the bottleneck — the policy is a small MLP over a 49-dimensional vector — so throughput comes from --num-env-runners and --envs-per-runner, not from an accelerator. --gpus-per-learner is there for CUDA machines; Apple MPS is not a Ray GPU resource.

Published artefacts land in benchmarks/ and are committed; Ray logs, checkpoints, and per-step traces land in runs/ and are not.

All three slices are in every split — GridWorld2D, Texture, and Top — across 63 family-size units. The splits differ only in which seeds they draw, never in which environments they contain: every unit appears in all four, so tune, train, val, and test are samples of the same task rather than different ones.

units instances per split
GridWorld2D 49 294 train · 147 each eval
Texture 8 48 train · 24 each eval
Top 6 36 train · 18 each eval

Seeds come from disjoint bands — tune 1000+, train 2000+, val 3000+, test 4000+, with the canonical seed 0 in none of them — and each instance carries size-scaled placement jitter, so no two are the same world. Every row records its canonical config, certified topology, turn-aware optimal route, and horizon, making a split's difficulty distribution auditable rather than asserted. Every split, and the extrapolation views, are published in croissant.json as their own Croissant record sets.

GridWorld2D dominates by unit count, so report per slice rather than pooling: a single mean over all instances is mostly a GridWorld2D score. Scenario mechanics stay live at benchmark defaults — including ClownChase's depleting reward trickle toward the wrong target, which is deception the benchmark is meant to contain.

import csv, gymnasium as gym, topogym

with open("docs/splits/train.csv") as f:
    for row in csv.DictReader(f):
        env = gym.make(row["template_id"], seed=int(row["seed"]),
                       placement_jitter=int(row["placement_jitter"]),
                       size=int(row["size"]))
        obs, info = env.reset(seed=0)
        # ... train; row["optimal_actions"] is the turn-aware optimum

Regenerate with python scripts/benchmarks/generate_splits.py; browse any split visually with python scripts/browse.py --all --split test -n 4.

Play any environment yourself

python scripts/play.py --list
python scripts/play.py TopoGym/SpaceWarp-v0

Arrow keys move; Tab reveals hidden structure; r resets; Backspace regenerates the layout. Rendering dims everything outside the agent's current line of sight (reveal mode shows all). Set TOPOGYM_DEBUG=1 to stream everything the env computes each step to the console, and TOPOGYM_OVERLAY=1 (alias OVERLAY_ENABLED=1) for the live H1 overlay: every step, the known region's holes are drawn on the grid — representative cycles in yellow, enclosed-wall rims in green (a yellow cycle with no green rim is a transient belief), with a legend and live H1 count top-right.

Determinism, certification, and stats

  • Determinism up to seeds is a guarantee, not an accident: (config, seed) fixes the layout and its metadata byte-for-byte — including everything computed through GUDHI — and (env, reset seed, actions) fixes the episode, p_slip included. Iteration orders are sorted so nothing depends on interpreter hash state; a cross-process test enforces it.
  • Certified metadata on every env (info["topology"]): Betti numbers in both door conventions, Euler characteristic, orientability, genus, bottleneck descriptors, the full generator configuration, and the canonical config string (TG-GridWorld2D-S50-C1-D4-...) as the run-log key. topogym.registry.manifest() emits the validity manifest.
  • Stats built in: info tracks within-episode coverage, lifetime (cross-episode) coverage, chamber entries, and return; StatsRecorder accumulates pandas-ready rows.

Learning from the topology w/ a map

Agents can choose to consume TopoGym's topology through VisitedComplex: feed it the states you have visited and read back the shape of what you know: a map of the holes you have found and the loops enclosing them. The representative cycles are closed walks through archive-restorable states, so an agent can treat them as places to return to, frontiers to push, or features to encode. The certified metadata stays the answer key for scoring; this is the signal.

Actions are named constants — env.step(FORWARD) says what env.step(2) only implies:

from topogym import TURN_LEFT, TURN_RIGHT, FORWARD    # Discrete(3)
from topogym import MOVE_UP, MOVE_DOWN, MOVE_LEFT, MOVE_RIGHT  # fourway

Actions are named constants — env.step(FORWARD) says what env.step(2) only implies:

from topogym import TURN_LEFT, TURN_RIGHT, FORWARD    # Discrete(3)
from topogym import MOVE_UP, MOVE_DOWN, MOVE_LEFT, MOVE_RIGHT  # fourway
from topogym.tda import VisitedComplex

vc = VisitedComplex.from_env(env)   # seeded with lifetime visits
vc.add(new_cells)                   # feed states as you explore
vc.betti()                          # (b0, b1) over the chosen ring
vc.representatives()                # a closed loop of cells per hole
vc.rims(observed=seen)              # where each loop can still tighten

Backends: cubical (movement-consistent on the env's own grid), vr (Vietoris–Rips at any epsilon, over cells or your encoder's vectors), and witness (de Silva–Carlsson landmarks, with the admit/evict policy yours to override). Coefficients: any prime or Z.

Cost — lazy and cached but not incremental, so query once an episode rather than once a step. Measured over F₂ on a dense square archive, calling in this order and timing each with the previous already cached: add fills the archive, then the build (triggered by the first query), then betti(), then representatives(), then rims(). add is negligible throughout (0.03s at 100k).

vr, ε = 1.5 — the general-purpose choice, and the one to assume for non-voxel spaces:

cells build betti representatives
1k 0.02s 0.01s 0.07s
10k 0.20s 0.45s 2.7s
50k 1.13s 3.09s 47s

cubical — for grid environments, where it matches movement:

cells build betti representatives rims
1k 0.02s 0.01s 0.02s ~0
10k 0.26s 0.23s 0.81s ~0
50k 1.48s 1.53s 12.5s ~0
100k 3.00s 5.30s 44.1s 0.01s

Builds and rims are linear and betti near-linear in both backends; representatives is the superlinear one — comfortable to ~20k cells, expensive past 50k. Costs are sequential, so cycles from a 100k-cell cubical archive cost the build plus the extraction (~47s), while a 50-grid archive is ~2.5k cells, where it is hundredths of a second. Use witness to hold a large point cloud at a fixed landmark budget. torsion() runs an integer Smith normal form and is an offline diagnostic, not an online signal.

Env Step Profile

How fast the environment steps under random actions, and how that scales across gymnasium.vector.AsyncVectorEnv:

Step throughput vs parallelism

A 50×50 world with actions="egocentric". obs_mode is the observation — a separate axis from the action space:

obs_mode what the agent sees
local (default) the occluded 7×7 egocentric patch of symbolic codes
dict that patch, plus per-cell textures and absolute position
vector position plus the texture block of the current cell
global the whole grid, unoccluded

Measured on an M5 Mac (18 cores) with 2–10% user/system usage.

python scripts/benchmarks/profiles/step_throughput.py

Documentation

  • docs/specs/topo_gym_overview.pdf — the detailed environment specification: world model, registry, generator schema, modes, reward semantics, complex backends, and the Texture/Top constructions. The authority on the benchmark.
  • docs/environments/ — per-environment pages (spaces, rewards, registered configurations).
  • docs/reference.md — library internals: the cell complex, the generator, TDA, the metrics interface, and VisitedComplex — the incremental visited-state topology structure (cubical / Vietoris–Rips / witness backends, F_p or Z coefficients, representative cycles) for building custom topological agents.
  • croissant.json + docs/manifest.csv — MLCommons Croissant metadata over the pinned registry (one record per environment id with its canonical config and certified topology), auto-generated by scripts/generate_croissant.py.

Contributing 🤝

Citation

If you use TopoGym in your research, please cite:

@software{carlson2026topogym,
  author  = {Carlson, Jason},
  title   = {TopoGym: Environments and Benchmarks for Topological
             Exploration in Reinforcement Learning},
  year    = {2026},
  url     = {https://github.com/jcarlson212/TopoGym},
  version = {0.2.0}
}

MIT. See also CITATION.cff.

Release files for topogym 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for topogym 0.2.0
File Size Uploaded
topogym-0.2.0.tar.gz 201.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for topogym 0.2.0
File Interpreter ABI Platform
topogym-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 432.1 kB

Release files / topogym-0.2.0.tar.gz

Download URL topogym-0.2.0.tar.gz
Size 201.1 kB
Tags Source
SHA-256 checksum
How to use checksums
7ce58b48ac8a3b22746fad41e541db148c91ed76980f1de400c43378bfcb5000
BLAKE2b-256 checksum
How to use checksums
13a6d0f7fb003940e6da9cb6399b902e64d632ea3472bfd3a87b107a781e5e40
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 11, 2026.

Transparency log

Release files / topogym-0.2.0-py3-none-any.whl

Download URL topogym-0.2.0-py3-none-any.whl
Size 231.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c7f9da0ed20c3135e8a8d355824b465860781e1cee3456d5c593bb84439a4763
BLAKE2b-256 checksum
How to use checksums
f566e31fb50ca7927f58034783f25ebbad81a985751b28f5b271a6107cec8ca1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 11, 2026.

Transparency log

Release history Release notifications | RSS feed

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page