Skip to main content

Vamos

Python License

Vamos is a JAX-native Reinforcement Learning environment API designed for high-performance parallel execution with a Gymnasium-like interface rebuilt from the ground up to leverage JAX's functional programming paradigm and automatic vectorization.

Key Features

  • Stateless, Functional Design: Unlike Gymnasium where state is stored internally, Vamos passes state explicitly as function parameters. This enables seamless composition with JAX transformations (jit, vmap, grad).

  • Gymnasium-Familiar API: If you know Gymnasium, you'll feel at home. Vamos uses similar concepts (spaces, wrappers, step/reset) adapted for JAX's functional style. Builtin is many of the popular Gymnasium environments, wrappers, and make, which is highly extensible.

Installation

pip install vamos-rl

Quick Start

import jax
import vamos

env, params = vamos.make("CartPole-v1")

# Initialize
rng = jax.random.PRNGKey(0)
timestep, state = env.reset(params, rng)
# the timestep is a dataclass containing your step data (observation, reward, etc)

# Take a step
action = env.action_space.sample(rng)
timestep, state = env.step(state, action, params, rng)

print(f"Observation: {timestep.obs}")
print(f"Reward: {timestep.reward}")
print(f"Episode Over: {timestep.episode_over}")  # this is equal to computing `termination or truncation`

Vectorized Environments

Run multiple environments in parallel with VMapVectorEnv:

import jax
import vamos

vec_env, params = vamos.make_vec("CartPole-v1", num_envs=1024)

rng = jax.random.PRNGKey(0)
timestep, state = vec_env.reset(params, rng)  # Get the reset observation and state for all 1024 environments

# Step all 1024 environments simultaneously
actions = vec_env.action_space.sample(rng)  # Shape: (1024,)
timestep, state = vec_env.step(state, actions, params, rng)

Vamos offers three strategies to optimize automatically reset sub-environments when episodes end:

  • COMPLETE: Generate N reset states every step (maximum diversity)
  • OPTIMISTIC: Generate M << N states, reuse when needed (balanced)
  • PRECOMPUTED: Pre-generate a pool before training (zero overhead)

See the vector environment documentation for details on autoreset modes and strategies.

Gymnasium vs Vamos

Aspect Gymnasium Vamos
State management Internal (mutable) Explicit (functional)
Vectorization SyncVectorEnv (Python loops) vmap (hardware-accelerated)
JIT compilation Not supported Native support
Autodiff through env Not possible Supported via JAX
Parallelism Process-based Array-based (GPU/TPU)
Randomness Modifiable at Episode Resets Selectable at every timestep

Gymnasium style (stateful):

obs, info = env.reset()
obs, reward, term, trunc, info = env.step(action)

Vamos style (functional):

timestep, state = env.reset(params, rng)
timestep, state = env.step(state, action, params, rng)

Core Concepts

Timestep

All environment outputs are bundled in a Timestep dataclass:

@struct.dataclass
class Timestep:
    obs: ArrayTree          # Current observation
    reward: float           # Reward from last action
    termination: bool       # Episode ended (goal/failure)
    truncation: bool        # Episode cut off (time limit)
    info: dict              # Additional information

    @property
    def episode_over(self):
        return self.termination or self.truncation

Spaces

Define valid actions and observations: Vamos supports a significantly more limited set of spaces, just three Scalar for individual values like a Discrete set of actions, Array for a vector or matrix of data like an image and Dict for composing multiple spaces together.

from vamos.spaces import Scalar, Array, Dict

# Discrete action (0, 1, 2, 3, 4)
action_space = Scalar(5)

# Continuous bounded values
obs_space = Array(low=[-1.0, -1.0], high=[1.0, 1.0])

# Composite spaces
space = Dict({"position": Array(...), "velocity": Array(...)})

Wrappers

Compose environment modifications:

from vamos.wrappers.time_limit import TimeLimit

env, params = CartPoleEnv.new()
env, params = TimeLimit.wrap(env, params, max_episode_steps=500)

License

MIT License

Release files for vamos-rl 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vamos-rl 0.1.0
File Size Uploaded
vamos_rl-0.1.0.tar.gz 34.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vamos-rl 0.1.0
File Interpreter ABI Platform
vamos_rl-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 69.7 kB

Release files / vamos_rl-0.1.0.tar.gz

Download URL vamos_rl-0.1.0.tar.gz
Size 34.1 kB
Tags Source
SHA-256 checksum
How to use checksums
c41fdd779c0f711a837e871b05ea7ea91b617180c6947c3b93a59db64fec5408
BLAKE2b-256 checksum
How to use checksums
9930e98f417075a0d8a4b41f5d5bf745180f4898cd26c88f181e0799bf9a4eee
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.7

Release files / vamos_rl-0.1.0-py3-none-any.whl

Download URL vamos_rl-0.1.0-py3-none-any.whl
Size 35.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
49b2bcfe55cc9210a2bf609f1112b4570c80d4f4ea12022a6f3053b9f1911009
BLAKE2b-256 checksum
How to use checksums
1e38c7c25c44b1a83b903e14fc60c3773782d3b48fd2f3d5415f44ee14a3d1b6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.7

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page