TopoGym
Gridworld environments with certified topology, for exploration research.
TopoGym is a Gymnasium environment library where the shape of every world — its chambers (sort of like rooms), decoys (filled rooms, large icebergs, or other blatant & large obstructions), and identifications (going in a circle or going in a circle while twisting in space) — is known exactly: computed from the free-space cell complex by GUDHI and cross-checked against the analytic expectation at generation time. Everything is deterministic up to seeds, end to end. We provide benchmarks for reinforcement learning researchers to test how good their agents are at exploring complex environment shapes.
EnvironmentalIceShip |
ClownChase |
SpaceWarp |
DontFall |
SearchRescue |
BankRobber |
Nested3-50 |
TopTorus-50 |
Full gallery and per-environment documentation: docs/envs/ · docs/environments/.
Environments
One benchmark, TopoGym-v1, in three slices under a universal interface (egocentric Discrete(3) turn-left / turn-right / forward actions with an occluded egocentric view by default — the rendered agent is a MiniGrid-style arrow, so its heading is always visible; actions="fourway" opts into Discrete(4) screen-direction actions with the universal (x, y) + 16-slot texture vector):
| slice | families | axis | status |
|---|---|---|---|
| GridWorld2D | Dilution, Chambers2, ChamberCount, Decoys, Shape{Sq,Ci,Tr,St}, Nested, GiveUp, Bottleneck, Maze |
world size, chamber/decoy count, shape, nesting depth, corridor length, braiding | 🟢 Live; Beta |
| Texture | IceShip, EnvironmentalIceShip, Ladders, BankRobber, DontFall, SpaceWarp, ClownChase, SearchRescue |
semantic local signals — and exactly where they fail | 🟢 Live; Beta |
| Top | TopPlane, TopCylinder, TopMobius, TopTorus, TopKlein, TopRP2 |
global topology with zero local signal | 🟢 Live; Beta |
Every id is stable: gym.make("TopoGym/{Family}-{size}-v0", seed=n). Details per family: docs/environments/.
Actions
Both spaces are named, so a policy says what it does rather than passing bare integers. The members are IntEnums, so they go straight to env.step:
from topogym import ActionMode, EgocentricAction, FourwayAction
env.step(EgocentricAction.FORWARD) # default Discrete(3): TURN_LEFT, TURN_RIGHT, FORWARD
env.step(FourwayAction.UP) # actions="fourway": UP, DOWN, LEFT, RIGHT (screen directions)
ActionMode.FOURWAY.actions # -> FourwayAction, for code generic over the mode
The bare names — FORWARD, TURN_LEFT, MOVE_UP, … — remain importable from topogym and are defined from the enums, so FORWARD == EgocentricAction.FORWARD holds by construction.
Install
pip install topogym # deps: gymnasium, numpy, gudhi
pip install "topogym[play]" # + pygame, for keyboard play
Development: git clone, then pip install -e ".[testing,play,assets]".
Quick start
import gymnasium as gym
import topogym # registers the TopoGym/* ids
env = gym.make("TopoGym/Decoys4-50-v0", seed=3)
obs, info = env.reset(seed=0)
info["topology"]["betti_z2"] # [1, 4, 0] — doors walkable
info["topology"]["betti_z2_sealed"] # [2, 5, 0] — doors count as walls
Episodes truncate after a pre-determined horizon — the larger of 1.2 * max(W, H) and 3x the turn-aware optimal route, so the goal is always reachable with room to wander; the goal pays +1 terminal reward by default (reward_mode="sparse") and sits inside a designated chamber. reward_mode="none" for pure exploration, "coverage", "deceptive"; goal=False removes the goal; p_slip=0.1 for sticky-action noise; complex="rips" swaps the homology backend to a Vietoris–Rips complex on the quotient metric.
Compose custom worlds with the fluent spec API:
from topogym.spec import Torus
env = Torus(15).holes(3).chambers(1).compile(seed=7)
Measure what an agent actually discovered — from its own trajectory:
from topogym.tda import ExplorationTracker
from topogym.stats import StatsRecorder
env = StatsRecorder(gym.make("TopoGym/Nested3-50-v0", seed=1))
tracker = ExplorationTracker(env)
tracker.reset(seed=0)
# ... run your policy ...
tracker.summary() # discovery-time persistence: real vs transient loops
env.episodes # per-episode rows: return, coverage, chamber entries
Archive-style (Go-Explore) resets are built in:
env = gym.make("TopoGym/Maze-100-v0", seed=1, teleport=True)
env.reset(options={"teleport": (12, 40)}) # any previously visited cell
Benchmarks
| benchmark | what it tests | manifest | splits | RND+PPO | ICM+PPO | Go-Explore | TopoExplore | status |
|---|---|---|---|---|---|---|---|---|
| TopoGym-v1 | topological navigation against decoys, chambers, distractions, and orientation in 2D space | croissant.json · docs/manifest.csv |
tune · train · val · test · size-extrapolation · family-holdout |
7 / 189 | 15 / 189 | 152 / 189 | 169 / 189 | ✅ published |
Worlds whose goal each method reached, of the 189 hold-out instances. Go-Explore and TopoExplore are read on the training side, where their archive is live; RND and ICM are read on the frozen evaluation, which discards an archive by construction and so scores every archive method zero. Paired world by world, every TopoExplore arm beats Go-Explore (19 worlds won against 2 lost for the strongest, sign test p < 0.001). The numbers, the per-arm sign tests and the frozen-evaluation table are in BENCHMARKS.md.
The random floor: across all 189 hold-out instances (50 episodes each, 9,450 episodes) a uniform-random policy reaches the goal 0% of the time and uncovers 11.0% of the reachable space. Nothing in this benchmark falls out of undirected exploration, and coverage — not steps-to-goal — is what separates methods until one of them solves something.
The published numbers are per-world exploration results: hyperparameters are chosen on the tune split, then every test world is its own experiment — one million environment steps of learning in it, followed by a frozen evaluation. They measure how much of a world a method uncovers given a budget in it, not whether a trained policy transfers; every method learns in the world it is scored on, under the same step budget. Full metrics, per-slice breakdowns, and the discovery-curve figures live in BENCHMARKS.md.
The transfer question — hyperparameters on tune, gradients on train, early stopping on val, and test read once at the end — remains fully posed by the published splits and enforced in-repo by Baseline.run(); we publish the splits so that benchmark can be run rather than exercising it ourselves. The algorithms themselves are Ray RLlib's; TopoGym does not reimplement PPO. A variant such as RND or ICM subclasses PPOBaseline and overrides one hook, and an algorithm that never uses PPO (Go-Explore explores randomly by default) implements the same small interface.
A framing worth keeping in mind when reading the numbers: the current baselines are deliberately unintelligent users of powerful primitives. The archive methods pair an exploration primitive that is strong on its own — remember every cell reached, restart from the frontier — with a policy that is essentially random; the curiosity methods (RND, ICM) learn a policy but explore without any archive at all. Neither half is the ceiling. Pairing them — a learned policy on top of an archive, trained with shared experience across the train split — should do far better than either: ICM + Go-Explore is the obvious first hybrid, and a PPO policy that learns how to use the archive (when to return to a frontier, which frontier to extend) is the more interesting one. Such a hybrid would not be compute-efficient on a handful of worlds — the pure primitive wins any single-world race — but it is the version that could learn to explore in general, and with transfer, amortize its training into a compute efficiency of its own. The interface above is small precisely so that these combinations are cheap to try.
--group decides what one policy is trained on, and therefore what is being measured. family (the default) trains a policy per family across its sizes and seeds, in the spirit of Procgen's train-on-levels, test-on-held-out-levels design; unit is the strictest per-world version; all asks instead for a single general explorer across every family at once.
pip install topogym[benchmarks]
python scripts/benchmarks/run_baselines_gridworld_v1_benchmark.py \
--baselines random,ppo --group family --num-env-runners 16
python scripts/benchmarks/run_baselines_gridworld_v1_benchmark.py --smoke # pipeline check
Environment stepping is the bottleneck — the policy is a small MLP over a 49-dimensional vector — so throughput comes from --num-env-runners and --envs-per-runner, not from an accelerator. --gpus-per-learner is there for CUDA machines; Apple MPS is not a Ray GPU resource.
Published artefacts land in benchmarks/ and are committed; Ray logs, checkpoints, and per-step traces land in runs/ and are not.
All three slices are in every split — GridWorld2D, Texture, and Top — across 63 family-size units. The splits differ only in which seeds they draw, never in which environments they contain: every unit appears in all four, so tune, train, val, and test are samples of the same task rather than different ones.
| units | instances per split | |
|---|---|---|
| GridWorld2D | 49 | 294 train · 147 each eval |
| Texture | 8 | 48 train · 24 each eval |
| Top | 6 | 36 train · 18 each eval |
Seeds come from disjoint bands — tune 1000+, train 2000+, val 3000+, test 4000+, with the canonical seed 0 in none of them — and each instance carries size-scaled placement jitter, so no two are the same world. Every row records its canonical config, certified topology, turn-aware optimal route, and horizon, making a split's difficulty distribution auditable rather than asserted. Every split, and the extrapolation views, are published in croissant.json as their own Croissant record sets.
GridWorld2D dominates by unit count, so report per slice rather than pooling: a single mean over all instances is mostly a GridWorld2D score. Scenario mechanics stay live at benchmark defaults — including ClownChase's depleting reward trickle toward the wrong target, which is deception the benchmark is meant to contain.
import csv, gymnasium as gym, topogym
with open("docs/splits/train.csv") as f:
for row in csv.DictReader(f):
env = gym.make(row["template_id"], seed=int(row["seed"]),
placement_jitter=int(row["placement_jitter"]),
size=int(row["size"]))
obs, info = env.reset(seed=0)
# ... train; row["optimal_actions"] is the turn-aware optimum
Regenerate with python scripts/benchmarks/generate_splits.py; browse any split visually with python scripts/browse.py --all --split test -n 4.
Play any environment yourself
python scripts/play.py --list
python scripts/play.py TopoGym/SpaceWarp-v0
Arrow keys move; Tab reveals hidden structure; r resets; Backspace regenerates the layout. Rendering dims everything outside the agent's current line of sight (reveal mode shows all). Set TOPOGYM_DEBUG=1 to stream everything the env computes each step to the console, and TOPOGYM_OVERLAY=1 (alias OVERLAY_ENABLED=1) for the live H1 overlay: every step, the known region's holes are drawn on the grid — representative cycles in yellow, enclosed-wall rims in green (a yellow cycle with no green rim is a transient belief), with a legend and live H1 count top-right.
Determinism, certification, and stats
- Determinism up to seeds is a guarantee, not an accident:
(config, seed) fixes the layout and its metadata byte-for-byte —
including everything computed through GUDHI — and (env, reset seed,
actions) fixes the episode,
p_slipincluded. Iteration orders are sorted so nothing depends on interpreter hash state; a cross-process test enforces it. - Certified metadata on every env (
info["topology"]): Betti numbers in both door conventions, Euler characteristic, orientability, genus, bottleneck descriptors, the full generator configuration, and the canonical config string (TG-GridWorld2D-S50-C1-D4-...) as the run-log key.topogym.registry.manifest()emits the validity manifest. - Stats built in:
infotracks within-episode coverage, lifetime (cross-episode) coverage, chamber entries, and return;StatsRecorderaccumulates pandas-ready rows.
Learning from the topology w/ a map
Agents can choose to consume TopoGym's topology through VisitedComplex: feed it the states you have visited and read back the shape of what you know: a map of the holes you have found and the loops enclosing them. The representative cycles are closed walks through archive-restorable states, so an agent can treat them as places to return to, frontiers to push, or features to encode. The certified metadata stays the answer key for scoring; this is the signal.
Actions are named constants — env.step(FORWARD) says what env.step(2) only implies:
from topogym import TURN_LEFT, TURN_RIGHT, FORWARD # Discrete(3)
from topogym import MOVE_UP, MOVE_DOWN, MOVE_LEFT, MOVE_RIGHT # fourway
Actions are named constants — env.step(FORWARD) says what env.step(2) only implies:
from topogym import TURN_LEFT, TURN_RIGHT, FORWARD # Discrete(3)
from topogym import MOVE_UP, MOVE_DOWN, MOVE_LEFT, MOVE_RIGHT # fourway
from topogym.tda import VisitedComplex
vc = VisitedComplex.from_env(env) # seeded with lifetime visits
vc.add(new_cells) # feed states as you explore
vc.betti() # (b0, b1) over the chosen ring
vc.representatives() # a closed loop of cells per hole
vc.rims(observed=seen) # where each loop can still tighten
Backends: cubical (movement-consistent on the env's own grid), vr (Vietoris–Rips at any epsilon, over cells or your encoder's vectors), and witness (de Silva–Carlsson landmarks, with the admit/evict policy yours to override). Coefficients: any prime or Z.
Cost — lazy and cached but not incremental, so query once an episode rather than once a step. Measured over F₂ on a dense square archive, calling in this order and timing each with the previous already cached: add fills the archive, then the build (triggered by the first query), then betti(), then representatives(), then rims(). add is negligible throughout (0.03s at 100k).
vr, ε = 1.5 — the general-purpose choice, and the one to assume for non-voxel spaces:
| cells | build | betti | representatives |
|---|---|---|---|
| 1k | 0.02s | 0.01s | 0.07s |
| 10k | 0.20s | 0.45s | 2.7s |
| 50k | 1.13s | 3.09s | 47s |
cubical — for grid environments, where it matches movement:
| cells | build | betti | representatives | rims |
|---|---|---|---|---|
| 1k | 0.02s | 0.01s | 0.02s | ~0 |
| 10k | 0.26s | 0.23s | 0.81s | ~0 |
| 50k | 1.48s | 1.53s | 12.5s | ~0 |
| 100k | 3.00s | 5.30s | 44.1s | 0.01s |
Builds and rims are linear and betti near-linear in both backends; representatives is the superlinear one — comfortable to ~20k cells, expensive past 50k. Costs are sequential, so cycles from a 100k-cell cubical archive cost the build plus the extraction (~47s), while a 50-grid archive is ~2.5k cells, where it is hundredths of a second. Use witness to hold a large point cloud at a fixed landmark budget. torsion() runs an integer Smith normal form and is an offline diagnostic, not an online signal.
Env Step Profile
How fast the environment steps under random actions, and how that scales across gymnasium.vector.AsyncVectorEnv:
A 50×50 world with actions="egocentric". obs_mode is the observation — a separate axis from the action space:
obs_mode |
what the agent sees |
|---|---|
local (default) |
the occluded 7×7 egocentric patch of symbolic codes |
dict |
that patch, plus per-cell textures and absolute position |
vector |
position plus the texture block of the current cell |
global |
the whole grid, unoccluded |
Measured on an M5 Mac (18 cores) with 2–10% user/system usage.
python scripts/benchmarks/profiles/step_throughput.py
Documentation
docs/specs/topo_gym_overview.pdf— the detailed environment specification: world model, registry, generator schema, modes, reward semantics, complex backends, and the Texture/Top constructions. The authority on the benchmark.- docs/environments/ — per-environment pages (spaces, rewards, registered configurations).
- docs/reference.md — library internals: the cell
complex, the generator, TDA, the metrics interface, and
VisitedComplex— the incremental visited-state topology structure (cubical / Vietoris–Rips / witness backends, F_p or Z coefficients, representative cycles) for building custom topological agents. croissant.json+docs/manifest.csv— MLCommons Croissant metadata over the pinned registry (one record per environment id with its canonical config and certified topology), auto-generated byscripts/generate_croissant.py.
Contributing 🤝
- Discord: join us.
- Add an environment without writing code:
scripts/new_env.py— walkthrough in docs/contributing_environments.md. - Extend the framework (new families, shapes, mechanics): CONTRIBUTING.md. All new topology ships with certified tests — the homology engine is the referee.
Citation
If you use TopoGym in your research, please cite:
@software{carlson2026topogym,
author = {Carlson, Jason},
title = {TopoGym: Environments and Benchmarks for Topological
Exploration in Reinforcement Learning},
year = {2026},
url = {https://github.com/jcarlson212/TopoGym},
version = {0.4.1}
}
MIT. See also CITATION.cff.
Release files for topogym 0.4.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| topogym-0.4.1.tar.gz | 239.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| topogym-0.4.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 508.7 kB
Release files / topogym-0.4.1.tar.gz
| Download URL | topogym-0.4.1.tar.gz |
|---|---|
| Size | 239.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b242d8e46906d2f00795a92714fa9f8fd2aa7c5ce33272ab3f1d230ea85ef9ec
|
|
BLAKE2b-256 checksum How to use checksums |
5d01890175f3dd6ac80a253cb6b3147d9d2dda03f09f7667746a113317c9c2b1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency logRelease files / topogym-0.4.1-py3-none-any.whl
| Download URL | topogym-0.4.1-py3-none-any.whl |
|---|---|
| Size | 269.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
85c523b73b20a8bc8f83906ece4faf1ee19d0ccfac614b9102424d48db4fdfaf
|
|
BLAKE2b-256 checksum How to use checksums |
af8000286fba398f12823b6353e95e07893fb736333ebb9e08b0fb184926ace3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 19, 2026.
Transparency log