Matrix MDP
Easily generate an MDP from transition and reward matricies.
Want to learn more on the story behind this repo? Check the blog post here!
Installation
Assuming you are in the root directory of the project, run the following command:
pip install matrix-mdp-gym
Usage
import gymnasium as gym
import matrix_mdp
env = gym.make('matrix_mdp/MatrixMDP-v0')
Environment documentation
Description
A flexible environment to have a gym API for discrete MDPs with N_s states and N_a actions given:
- A vector of initial state distribution vector P_0(S)
- A transition probability matrix P(S' | S, A)
- A reward matrix R(S', S, A) of the reward for reaching S' after having taken action A in state S
Action Space
The action is a ndarray with shape (1,) representing the index of the action to execute.
Observation Space
The observation is a ndarray with shape (1,) representing index of the state the agent is in.
Rewards
The reward function is defined according to the reward matrix given at the creation of the environment.
Starting State
The starting state is a random state sampled from $P_0$.
Episode Truncation
The episode truncates when a terminal state is reached. Terminal states are inferred from the transition probability matrix as $\sum_{s' \in S} \sum_{s \in S} \sum_{a \in A} P(s' | s, a) = 0$
Arguments
p_à:ndarrayof shape(n_states, )representing the initial state probability distribution.p:ndarrayof shape(n_states, n_states, n_actions)representing the transition dynamics $P(S' | S, A)$.r:ndarrayof shape(n_states, n_states, n_actions)representing the reward matrix.
import gymnasium as gym
import matrix_mdp
gym.make('MatrixMDP-v0', p_0=p_0, p=p, r=r)
Version History
v0: Initial versions release
Acknowledgements
Thanks to Will Dudley for his help on learning how to put a Python package together/
Metadata
Release files for matrix-mdp-gym 1.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| matrix-mdp-gym-1.1.1.tar.gz | 5.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| matrix_mdp_gym-1.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 11.3 kB
Release files / matrix-mdp-gym-1.1.1.tar.gz
| Download URL | matrix-mdp-gym-1.1.1.tar.gz |
|---|---|
| Size | 5.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a1e3e2ec1805f5e12c5f2d774e95e7acc3e133a69f0c690bb659c6807fa8bd08
|
|
BLAKE2b-256 checksum How to use checksums |
da9f126c5807efcac8cd4b0b09ab907c5437ccc76eb01c749a0a3b1e0ed97fe5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.1 CPython/3.9.16
|
Release files / matrix_mdp_gym-1.1.1-py3-none-any.whl
| Download URL | matrix_mdp_gym-1.1.1-py3-none-any.whl |
|---|---|
| Size | 6.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a647ce63d11ad8bf9391a9816f1358fbb13100a1f733c6bd40b908062b581177
|
|
BLAKE2b-256 checksum How to use checksums |
d359b599f21b2db76259549fe3599327e2aa5aee6145a9be56e7e2d8aacaa40c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.1 CPython/3.9.16
|