TopoGym
Gridworld environments with certified topology, for exploration research.
TopoGym is a Gymnasium environment library where the shape of every world — its chambers (sort of like rooms), decoys (filled rooms, large icebergs, or other blatant & large obstructions), and identifications (going in a circle or going in a circle while twisting in space) — is known exactly: computed from the free-space cell complex by GUDHI and cross-checked against the analytic expectation at generation time. Everything is deterministic up to seeds, end to end. We provide benchmarks for reinforcement learning researchers to test how good their agents are at exploring complex environment shapes.
EnvironmentalIceShip |
ClownChase |
SpaceWarp |
DontFall |
SearchRescue |
BankRobber |
Nested3-50 |
TopTorus-50 |
Full gallery and per-environment documentation:
docs/envs/ ·
docs/environments/.
Environments
One benchmark, TopoGym-v1, in three slices under a universal
interface (egocentric Discrete(3) turn-left / turn-right / forward
actions with an occluded egocentric view by default — the rendered
agent is a MiniGrid-style arrow, so its heading is always visible;
actions="fourway" opts into Discrete(4) screen-direction actions
with the universal (x, y) + 16-slot texture vector):
| slice | families | axis | status |
|---|---|---|---|
| GridWorld2D | Dilution, Chambers2, ChamberCount, Decoys, Shape{Sq,Ci,Tr,St}, Nested, GiveUp, Bottleneck, Maze |
world size, chamber/decoy count, shape, nesting depth, corridor length, braiding | 🟢 Live; Beta |
| Texture | IceShip, EnvironmentalIceShip, Ladders, BankRobber, DontFall, SpaceWarp, ClownChase, SearchRescue |
semantic local signals — and exactly where they fail | 🟢 Live; Beta |
| Top | TopPlane, TopCylinder, TopMobius, TopTorus, TopKlein, TopRP2 |
global topology with zero local signal | 🟢 Live; Beta |
Every id is stable: gym.make("TopoGym/{Family}-{size}-v0", seed=n).
Details per family: docs/environments/.
Install
pip install topogym # deps: gymnasium, numpy, gudhi
pip install "topogym[play]" # + pygame, for keyboard play
Development: git clone, then pip install -e ".[testing,play,assets]".
Quick start
import gymnasium as gym
import topogym # registers the TopoGym/* ids
env = gym.make("TopoGym/Decoys4-50-v0", seed=3)
obs, info = env.reset(seed=0)
info["topology"]["betti_z2"] # [1, 4, 0] — doors walkable
info["topology"]["betti_z2_sealed"] # [2, 5, 0] — doors count as walls
Episodes truncate after a pre-determined horizon — the larger of
1.2 * max(W, H) and 3x the turn-aware optimal route, so the goal is
always reachable with room to wander; the
goal pays +1 terminal reward by default (reward_mode="sparse") and
sits inside a designated chamber. reward_mode="none" for pure
exploration, "coverage", "deceptive"; goal=False removes the goal;
p_slip=0.1 for sticky-action noise; complex="rips" swaps the
homology backend to a Vietoris–Rips complex on the quotient metric.
Compose custom worlds with the fluent spec API:
from topogym.spec import Torus
env = Torus(15).holes(3).chambers(1).compile(seed=7)
Measure what an agent actually discovered — from its own trajectory:
from topogym.tda import ExplorationTracker
from topogym.stats import StatsRecorder
env = StatsRecorder(gym.make("TopoGym/Nested3-50-v0", seed=1))
tracker = ExplorationTracker(env)
tracker.reset(seed=0)
# ... run your policy ...
tracker.summary() # discovery-time persistence: real vs transient loops
env.episodes # per-episode rows: return, coverage, chamber entries
Archive-style (Go-Explore) resets are built in:
env = gym.make("TopoGym/Maze-100-v0", seed=1, teleport=True)
env.reset(options={"teleport": (12, 40)}) # any previously visited cell
Benchmarks
| benchmark | what it tests | manifest | splits | RND+PPO | ICM+PPO | Go-Explore | status |
|---|---|---|---|---|---|---|---|
| TopoGym-v1 | topological navigation against decoys, chambers, distractions, and orientation in 2D space | croissant.json · docs/manifest.csv |
tune · train · val · test · size-extrapolation · family-holdout |
TBD | TBD | TBD | 🟠 in development |
The random floor is measured under exactly that protocol: across all 189 hold-out instances (50 episodes each, 9,450 episodes) a uniform-random policy reaches the goal 0% of the time and uncovers 11.0% of the reachable space. Nothing in this benchmark falls out of undirected exploration, and coverage — not steps-to-goal — is what separates methods until one of them solves something.
Baselines report median steps to find the goal, with a 95% bootstrap confidence interval, over the hold-out split. Full metrics, per-slice breakdowns, and the discovery-curve figures live in BENCHMARKS.md.
Every baseline consumes the splits the same way — hyperparameters on
tune, gradients on train, early stopping on val, and test read
once at the end — enforced by Baseline.run() rather than left to each
algorithm. The algorithms themselves are Ray RLlib's; TopoGym does not
reimplement PPO. A variant such as RND or ICM subclasses PPOBaseline
and overrides one hook, and an algorithm that never uses PPO (Go-Explore
explores randomly by default) implements the same small interface.
--group decides what one policy is trained on, and therefore what is
being measured. family (the default) trains a policy per family
across its sizes and seeds, in the spirit of Procgen's train-on-levels,
test-on-held-out-levels design; unit is the strictest per-world
version; all asks instead for a single general explorer across every
family at once.
pip install topogym[benchmarks]
python scripts/benchmarks/run_baselines_gridworld_v1_benchmark.py \
--baselines random,ppo --group family --num-env-runners 16
python scripts/benchmarks/run_baselines_gridworld_v1_benchmark.py --smoke # pipeline check
Environment stepping is the bottleneck — the policy is a small MLP over
a 49-dimensional vector — so throughput comes from --num-env-runners
and --envs-per-runner, not from an accelerator. --gpus-per-learner
is there for CUDA machines; Apple MPS is not a Ray GPU resource.
Published artefacts land in benchmarks/ and are
committed; Ray logs, checkpoints, and per-step traces land in runs/
and are not.
All three slices are in every split — GridWorld2D, Texture, and Top — across 63 family-size units. The splits differ only in which seeds they draw, never in which environments they contain: every unit appears in all four, so tune, train, val, and test are samples of the same task rather than different ones.
| units | instances per split | |
|---|---|---|
| GridWorld2D | 49 | 294 train · 147 each eval |
| Texture | 8 | 48 train · 24 each eval |
| Top | 6 | 36 train · 18 each eval |
Seeds come from disjoint bands — tune 1000+, train 2000+, val 3000+,
test 4000+, with the canonical seed 0 in none of them — and each
instance carries size-scaled placement jitter, so no two are the same
world. Every row records its canonical config, certified topology,
turn-aware optimal route, and horizon, making a split's difficulty
distribution auditable rather than asserted. Every split, and the
extrapolation views, are published in croissant.json as their own
Croissant record sets.
GridWorld2D dominates by unit count, so report per slice rather than pooling: a single mean over all instances is mostly a GridWorld2D score. Scenario mechanics stay live at benchmark defaults — including ClownChase's depleting reward trickle toward the wrong target, which is deception the benchmark is meant to contain.
import csv, gymnasium as gym, topogym
with open("docs/splits/train.csv") as f:
for row in csv.DictReader(f):
env = gym.make(row["template_id"], seed=int(row["seed"]),
placement_jitter=int(row["placement_jitter"]),
size=int(row["size"]))
obs, info = env.reset(seed=0)
# ... train; row["optimal_actions"] is the turn-aware optimum
Regenerate with python scripts/benchmarks/generate_splits.py; browse any split
visually with python scripts/browse.py --all --split test -n 4.
Play any environment yourself
python scripts/play.py --list
python scripts/play.py TopoGym/SpaceWarp-v0
Arrow keys move; Tab reveals hidden structure; r resets;
Backspace regenerates the layout. Rendering dims everything outside
the agent's current line of sight (reveal mode shows all). Set
TOPOGYM_DEBUG=1 to stream everything the env computes each step to
the console, and TOPOGYM_OVERLAY=1 (alias OVERLAY_ENABLED=1) for
the live H1 overlay: every step, the known region's holes are drawn on
the grid — representative cycles in yellow, enclosed-wall rims in
green (a yellow cycle with no green rim is a transient belief), with a
legend and live H1 count top-right.
Determinism, certification, and stats
- Determinism up to seeds is a guarantee, not an accident:
(config, seed) fixes the layout and its metadata byte-for-byte —
including everything computed through GUDHI — and (env, reset seed,
actions) fixes the episode,
p_slipincluded. Iteration orders are sorted so nothing depends on interpreter hash state; a cross-process test enforces it. - Certified metadata on every env (
info["topology"]): Betti numbers in both door conventions, Euler characteristic, orientability, genus, bottleneck descriptors, the full generator configuration, and the canonical config string (TG-GridWorld2D-S50-C1-D4-...) as the run-log key.topogym.registry.manifest()emits the validity manifest. - Stats built in:
infotracks within-episode coverage, lifetime (cross-episode) coverage, chamber entries, and return;StatsRecorderaccumulates pandas-ready rows.
Learning from the topology w/ a map
Agents can choose to consume TopoGym's topology through
VisitedComplex:
feed it the states you have visited and read back the shape of what
you know: a map of the holes you have found and the loops enclosing
them. The representative cycles are closed walks through
archive-restorable states, so an agent can treat them as places to
return to, frontiers to push, or features to encode. The certified
metadata stays the answer key for scoring; this is the signal.
Actions are named constants — env.step(FORWARD) says what
env.step(2) only implies:
from topogym import TURN_LEFT, TURN_RIGHT, FORWARD # Discrete(3)
from topogym import MOVE_UP, MOVE_DOWN, MOVE_LEFT, MOVE_RIGHT # fourway
Actions are named constants — env.step(FORWARD) says what
env.step(2) only implies:
from topogym import TURN_LEFT, TURN_RIGHT, FORWARD # Discrete(3)
from topogym import MOVE_UP, MOVE_DOWN, MOVE_LEFT, MOVE_RIGHT # fourway
from topogym.tda import VisitedComplex
vc = VisitedComplex.from_env(env) # seeded with lifetime visits
vc.add(new_cells) # feed states as you explore
vc.betti() # (b0, b1) over the chosen ring
vc.representatives() # a closed loop of cells per hole
vc.rims(observed=seen) # where each loop can still tighten
Backends: cubical (movement-consistent on the env's own grid), vr
(Vietoris–Rips at any epsilon, over cells or your encoder's
vectors), and witness (de Silva–Carlsson landmarks, with the
admit/evict policy yours to override). Coefficients: any prime or Z.
Cost — lazy and cached but not incremental, so query once an
episode rather than once a step. Measured over F₂ on a dense square
archive, calling in this order and timing each with the previous
already cached: add fills the archive, then the build (triggered by
the first query), then betti(), then representatives(), then
rims(). add is negligible throughout (0.03s at 100k).
vr, ε = 1.5 — the general-purpose choice, and the one to assume
for non-voxel spaces:
| cells | build | betti | representatives |
|---|---|---|---|
| 1k | 0.02s | 0.01s | 0.07s |
| 10k | 0.20s | 0.45s | 2.7s |
| 50k | 1.13s | 3.09s | 47s |
cubical — for grid environments, where it matches movement:
| cells | build | betti | representatives | rims |
|---|---|---|---|---|
| 1k | 0.02s | 0.01s | 0.02s | ~0 |
| 10k | 0.26s | 0.23s | 0.81s | ~0 |
| 50k | 1.48s | 1.53s | 12.5s | ~0 |
| 100k | 3.00s | 5.30s | 44.1s | 0.01s |
Builds and rims are linear and betti near-linear in both backends;
representatives is the superlinear one — comfortable to ~20k cells,
expensive past 50k. Costs are sequential, so cycles from a 100k-cell
cubical archive cost the build plus the extraction (~47s), while a
50-grid archive is ~2.5k cells, where it is hundredths of a second.
Use witness to hold a large point cloud at a fixed landmark budget.
torsion() runs an integer Smith normal form and is an offline
diagnostic, not an online signal.
Env Step Profile
How fast the environment steps under random actions, and how that
scales across gymnasium.vector.AsyncVectorEnv:
A 50×50 world with actions="egocentric". obs_mode is the
observation — a separate axis from the action space:
obs_mode |
what the agent sees |
|---|---|
local (default) |
the occluded 7×7 egocentric patch of symbolic codes |
dict |
that patch, plus per-cell textures and absolute position |
vector |
position plus the texture block of the current cell |
global |
the whole grid, unoccluded |
Measured on an M5 Mac (18 cores) with 2–10% user/system usage.
python scripts/benchmarks/profiles/step_throughput.py
Documentation
docs/specs/topo_gym_overview.pdf— the detailed environment specification: world model, registry, generator schema, modes, reward semantics, complex backends, and the Texture/Top constructions. The authority on the benchmark.- docs/environments/ — per-environment pages (spaces, rewards, registered configurations).
- docs/reference.md — library internals: the cell
complex, the generator, TDA, the metrics interface, and
VisitedComplex— the incremental visited-state topology structure (cubical / Vietoris–Rips / witness backends, F_p or Z coefficients, representative cycles) for building custom topological agents. croissant.json+docs/manifest.csv— MLCommons Croissant metadata over the pinned registry (one record per environment id with its canonical config and certified topology), auto-generated byscripts/generate_croissant.py.
Contributing 🤝
- Discord: join us.
- Add an environment without writing code:
scripts/new_env.py— walkthrough in docs/contributing_environments.md. - Extend the framework (new families, shapes, mechanics): CONTRIBUTING.md. All new topology ships with certified tests — the homology engine is the referee.
Citation
If you use TopoGym in your research, please cite:
@software{carlson2026topogym,
author = {Carlson, Jason},
title = {TopoGym: Environments and Benchmarks for Topological
Exploration in Reinforcement Learning},
year = {2026},
url = {https://github.com/jcarlson212/TopoGym},
version = {0.2.0}
}
MIT. See also CITATION.cff.
Release files for topogym 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| topogym-0.2.0.tar.gz | 201.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| topogym-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 432.1 kB
Release files / topogym-0.2.0.tar.gz
| Download URL | topogym-0.2.0.tar.gz |
|---|---|
| Size | 201.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7ce58b48ac8a3b22746fad41e541db148c91ed76980f1de400c43378bfcb5000
|
|
BLAKE2b-256 checksum How to use checksums |
13a6d0f7fb003940e6da9cb6399b902e64d632ea3472bfd3a87b107a781e5e40
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 11, 2026.
Transparency logRelease files / topogym-0.2.0-py3-none-any.whl
| Download URL | topogym-0.2.0-py3-none-any.whl |
|---|---|
| Size | 231.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c7f9da0ed20c3135e8a8d355824b465860781e1cee3456d5c593bb84439a4763
|
|
BLAKE2b-256 checksum How to use checksums |
f566e31fb50ca7927f58034783f25ebbad81a985751b28f5b271a6107cec8ca1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 11, 2026.
Transparency log