Skip to main content

CartPole SwingUp environment for Gymnasium

Project description

Gymnasium CartPole SwingUp

PyPI version Python Versions License Tests GitHub release

A more challenging version of the classic CartPole environment for Gymnasium where the pole starts in a downward position.

Description

This package provides a port of the CartPole SwingUp environment to the modern Gymnasium API. It is based on:

The environment has been updated to work with the latest Gymnasium interface and includes enhanced rendering capabilities.

Installation

# Using pip
pip install gymnasium-cartpole-swingup

# Using uv
uv add gymnasium-cartpole-swingup

Usage

import gymnasium as gym
import gymnasium_cartpole_swingup  # This import is required to register the environment, even if unused

# Create the environment
env = gym.make("CartPoleSwingUp-v0", render_mode="human")
observation, info = env.reset(seed=42)

for _ in range(1000):
    action = env.action_space.sample()
    observation, reward, terminated, truncated, info = env.step(action)
    
    if terminated or truncated:
        observation, info = env.reset()

env.close()

Customizing Environment Parameters

You can customize the physics parameters of the environment by passing them to gym.make():

# Create an environment with custom parameters
env = gym.make(
    "CartPoleSwingUp-v0",
    render_mode="human",
    gravity=9.81,             # Gravitational acceleration (m/s²)
    cart_mass=1.0,            # Mass of the cart (kg)
    pole_mass=0.1,            # Mass of the pole (kg)
    pole_length=0.6,          # Length of the pole (m)
    force_mag=10.0,           # Force magnitude scale applied to cart
    friction=0.05,            # Friction coefficient
    x_threshold=2.5,          # Cart position limit (left/right boundary)
    cost_mode="default",      # Cost function mode ("default" or "pilco")
    sigma_c=0.25,             # Sigma parameter for PILCO cost function
)

Note: The import gymnasium_cartpole_swingup line is necessary to register the environment with Gymnasium, even though it may appear unused. If you're using auto-formatters or linters that remove unused imports, you can add a # noqa comment or disable that specific check:

import gymnasium_cartpole_swingup  # noqa: F401

Environment Details

  • State: Initially, the pole hangs downward ($\theta \approx \pi$)
  • Goal: Swing the pole upright and maintain balance
  • Action Space: Force applied to cart $[-1, 1]$ (scaled to $[-10, 10]$ N internally)
  • Observation Space: $[x, \dot{x}, \theta, \dot{\theta}]$
  • Reward: Higher when pole is upright and cart is centered

Observation Space Detail

The observation is a 4-dimensional vector:

Index Observation Description Min Max
0 $x$ Cart position along the track $-2.4$ $2.4$
1 $\dot{x}$ Cart velocity $-\infty$ $\infty$
2 $\theta$ Angle of the pole $-\pi$ $\pi$
3 $\dot{\theta}$ Angular velocity of the pole $-\infty$ $\infty$

Notes:

  • The angle $\theta$ is in radians and is kept within the range $[-\pi, \pi]$
  • When the pole is upright, $\theta = 0$
  • When the pole is hanging down, $\theta = \pi$ or $\theta = -\pi$

Action Space Detail

The action is a 1-dimensional continuous value:

Index Action Description Min Max
0 $F$ Horizontal force applied to the cart $-1.0$ $1.0$

Notes:

  • The force is scaled internally by a factor of $10.0$, resulting in an effective range of $[-10, 10]$ N
  • Positive values move the cart to the right
  • Negative values move the cart to the left

Reward Function

The environment supports two different reward (or cost) functions, which can be selected using the cost_mode parameter:

Default Mode (cost_mode="default")

The default reward function is a product of two components:

  • Pole angle component: $\cos(\theta)$

    • Maximum value of $1.0$ when the pole is upright ($\theta = 0$)
    • Minimum value of $-1.0$ when the pole is hanging down ($\theta = \pi$ or $\theta = -\pi$)
  • Cart position component: $\cos(x)$

    • Maximum value of $1.0$ when the cart is centered ($x = 0$)
    • Decreases as the cart moves away from center

Total reward = pole angle component $\times$ cart position component

PILCO Mode (cost_mode="pilco")

The PILCO (Probabilistic Inference for Learning COntrol) cost function is based on the squared distance between the pole tip position and the target position:

$reward = 1 - \exp(-\frac{d^2}{2\sigma_c^2})$

Where:

  • $d$ is the Euclidean distance between the current pole tip position and the target (upright) position
  • $\sigma_c$ is a parameter controlling the width of the cost function (default: 0.25)

This cost function is more focused on the pole tip position in Cartesian space rather than the angular position and cart position separately.

System Dynamics

The system dynamics follow the standard cart-pole physics model. The state update equations are:

$\ddot{x} = \frac{-2m_p l \dot{\theta}^2 \sin(\theta) + 3m_p g \sin(\theta)\cos(\theta) + 4F - 4b\dot{x}}{4(m_c + m_p) - 3m_p \cos^2(\theta)}$

$\ddot{\theta} = \frac{-3m_p l \dot{\theta}^2 \sin(\theta)\cos(\theta) + 6(m_c + m_p)g\sin(\theta) + 6(F - b\dot{x})\cos(\theta)}{4l(m_c + m_p) - 3m_p l \cos^2(\theta)}$

Where:

  • $m_c = 0.5$ (kg): Mass of the cart (default)
  • $m_p = 0.5$ (kg): Mass of the pole (default)
  • $l = 0.6$ (m): Length of the pole (default)
  • $g = 9.82$ (m/s²): Gravitational acceleration (default)
  • $b = 0.1$: Friction coefficient (default)
  • $F$: Applied force, scaled from action value to range $[-10, 10]$ N

All of these parameters can be customized when creating the environment as shown in the example above.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gymnasium_cartpole_swingup-0.1.3.tar.gz (10.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gymnasium_cartpole_swingup-0.1.3-py3-none-any.whl (9.2 kB view details)

Uploaded Python 3

File details

Details for the file gymnasium_cartpole_swingup-0.1.3.tar.gz.

File metadata

File hashes

Hashes for gymnasium_cartpole_swingup-0.1.3.tar.gz
Algorithm Hash digest
SHA256 4aecbce64c0ce8c107b916b4324a8cdacd67aca320811bab2e09f402c46f37b3
MD5 cc495e958042850de10ab80596d90f0c
BLAKE2b-256 f38a0ab30007fe13c5b2924f3938e03a552d598b14193481a4672aa97f4dc48c

See more details on using hashes here.

Provenance

The following attestation bundles were made for gymnasium_cartpole_swingup-0.1.3.tar.gz:

Publisher: publish-pypi.yml on nkiyohara/gymnasium-cartpole-swingup

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file gymnasium_cartpole_swingup-0.1.3-py3-none-any.whl.

File metadata

File hashes

Hashes for gymnasium_cartpole_swingup-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 3819f61b29ec526d7912f4896dea474e8a44f057b91a012dfbd5191d56ecf6f6
MD5 7205c76530a06fe296ac5e4952313176
BLAKE2b-256 72ee96019623a5ac391aa65383ef9eeea10ab832f8fb8a3e968e41eec5c3f214

See more details on using hashes here.

Provenance

The following attestation bundles were made for gymnasium_cartpole_swingup-0.1.3-py3-none-any.whl:

Publisher: publish-pypi.yml on nkiyohara/gymnasium-cartpole-swingup

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page