Skip to main content

RewardGen

RewardGen is a python package that makes it easy to apply any reward model to your robot videos and plot the rewards as shown below. (All example videos at: https://philip-mit.github.io/rewardgen_view/)

Examples

https://github.com/user-attachments/assets/3c444096-d3dd-47c7-b09d-90b0756d0f72

Supported Models

ToDos

  • Enable fine-tuning of reward models on new datasets/demonstrations

File Structure

rewardgen/
├── rewardgen/         # Main package
│   ├── robometer/         # Robometer code
│   ├── sole.py            # SOLE-R1 code
│   ├── roboreward.py      # RoboReward code
│   ├── topreward.py       # TOPReward code
│   └── api_models.py      # OpenAI and Gemini APIs
├── test_videos/        # Example videos to test
├── model_outputs/      # Example videos showing model outputs
├── docs/   
│   ├── lerobot_dataset_reward_annotation.mdx  # Examples showing integration with lerobot datasets
└── pyproject.toml      # Dependencies (uv)

Install

Option 1: quick pip install

pip install -U rewardgen

Option 2: use uv for dependency management

# 1) Clone the repository
git clone https://github.com/Philip-MIT/rewardgen

# 2) Install `uv`
pip install uv

# 3) Sync environment
cd rewardgen
uv sync

# 4) Activate environment
source .venv/bin/activate

Optional: Pre-download model checkpoints

# SOLE-R1 (8B) 
python -c "from rewardgen.utils.model_utils import get_model_dir; get_model_dir('sole-r1')"

# Robometer (4B)
python -c "from rewardgen.utils.model_utils import get_model_dir; get_model_dir('robometer')"

# TOPReward (based on Qwen3-VL-8B)
python -c "from rewardgen.utils.model_utils import get_model_dir; get_model_dir('topreward')"

# RoboReward (8B)
python -c "from rewardgen.utils.model_utils import get_model_dir; get_model_dir('roboreward')"

> **Note:** Robometer is ~8GB. SOLE-R1, RoboReward, and TOPReward are ~17GB each.

Optional: Download all test videos and example model outputs

# 1) Install gcloud: https://cloud.google.com/sdk/docs/install

# 2) Go to target directory
# cd /path/to/rewardgen

# Optional: disable credentials so you don't have to authenticate
gcloud config set auth/disable_credentials True

# Download test videos
gcloud storage cp --recursive gs://roboreason-view-videos-philip/test_videos ./

# Download model outputs for all test videos
gcloud storage cp --recursive gs://roboreason-view-videos-philip/model_outputs ./

# Optional: re-enable credentials afterward if you disabled them above.
gcloud config set auth/disable_credentials False

Quick start: Example reward generation and plotting

# pip install -U rewardgen
from rewardgen import generate, video_plot

video_paths_external_view = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/view_external.mp4']
video_paths_wrist_view = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/view_wrist.mp4']
task_description = "Pick up the cube from the table."

# Robometer
response = generate(model="Robometer",  task_description=task_description, video_paths_external_view=video_paths_external_view,  verbose=False)
print(response.rewards)
print(response.success_probs)
output_robometer = {"model": "Robometer", "rewards": response.rewards[0]}

# SOLE-R1
response = generate(model="SOLE-R1",  task_description=task_description, video_paths_external_view=video_paths_external_view, video_paths_wrist_view=video_paths_wrist_view, verbose=False)
print(response.rewards)
print(response.output_text)
output_sole = {"model": "SOLE-R1", "rewards": response.rewards[0], "output_text": response.output_text[0]}

# Optional: Ground-truth rewards (available for test videos from sim environments)
import json
with open('test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/data.json', 'r') as f:
    data = json.load(f)

output_groundtruth = {"model": "Ground truth", "rewards": data['ground-truth rewards']}

# Plot
video_plot(outputs=[output_groundtruth, output_sole, output_robometer], plot_save_path='model_outputs/combined/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37.mp4', video_view_external_path=video_paths_external_view[0], video_view_wrist_path=video_paths_wrist_view[0], task_description=task_description)

Examples for generating across all models

Robometer

from rewardgen import generate

video_paths = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/view_external.mp4']
task_description = "Pick up the cube from the table."

response = generate(
    model="Robometer",  
    task_description=task_description, 
    video_paths=video_paths, 
    view_type="external",
    verbose=False
)
print(response.rewards)
print(response.success_probs)

SOLE-R1

from rewardgen import generate

video_paths = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37.mp4']
task_description = "Pick up the cube from the table."

response = generate(
    model="SOLE-R1",  
    task_description=task_description, 
    video_paths=video_paths, 
    view_type='external and wrist',
    verbose=False
)
print(response.rewards)
print(response.output_text)

output_sole = {"model": "SOLE-R1", "rewards": response.rewards[0], "output_text": response.output_text[0]}

# Plotting with show_output_text=True
video_plot(
    outputs=[output_sole], 
    plot_save_path='model_outputs/combined/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37.mp4', 
    video_path=video_paths[0],
    show_output_text=True,
    task_description=task_description,
    verbose=False
)

TOPReward

from rewardgen import generate

video_paths = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/view_external.mp4']
task_description = "Pick up the cube from the table."

response = generate(
    model="TOPReward",  
    task_description=task_description, 
    video_paths=video_paths, 
    view_type='external',
    verbose=False
)
print(response.rewards)

RoboReward

from rewardgen import generate

video_paths = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/view_external.mp4']
task_description = "Pick up the cube from the table."

response = generate(
    model="RoboReward",  
    task_description=task_description, 
    video_paths=video_paths, 
    view_type='external',
    verbose=False
)
print(response.rewards)

GPT-5 (and other OpenAI models)

from rewardgen import generate

video_paths = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/view_external.mp4']
task_description = "Pick up the cube from the table."

# requires OpenAI API key: https://developers.openai.com/api/docs/quickstart
API_KEY = "..."

response = generate(
    model="GPT-5",  
    task_description=task_description, 
    video_paths=video_paths, 
    view_type='external', 
    key=API_KEY, 
    verbose=False
)
print(response.rewards)
print(response.output_text)

Gemini-3-Pro (and other Google models)

from rewardgen import generate

video_paths = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/view_external.mp4']
task_description = "Pick up the cube from the table."

# requires Gemini API key: https://ai.google.dev/gemini-api/docs/api-key
API_KEY = "..."

response = generate(
    model="Gemini-3.1-Pro-Preview",  
    task_description=task_description, 
    video_paths=video_paths, 
    view_type='external', 
    key=API_KEY,
    verbose=False
)
print(response.rewards)
print(response.output_text)

Video plotting

from rewardgen import generate, video_plot

video_paths_external_view = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/view_external.mp4']
video_paths_wrist_view = ['test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/view_wrist.mp4']
task_description = "Pick up the cube from the table."

# Robometer
response = generate(model="Robometer",  task_description=task_description, video_paths_external_view=video_paths_external_view,  verbose=False)
output_robometer = {"model": "Robometer", "rewards": response.rewards[0]}

# SOLE-R1
response = generate(model="SOLE-R1",  task_description=task_description, video_paths_external_view=video_paths_external_view, video_paths_wrist_view=video_paths_wrist_view, verbose=False)
output_sole = {"model": "SOLE-R1", "rewards": response.rewards[0], "output_text": response.output_text[0]}

# Optional: Ground-truth rewards (available for test videos from sim environments)
import json
with open('test_videos/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37/data.json', 'r') as f:
    data = json.load(f)

output_groundtruth = {"model": "Ground truth", "rewards": data['ground-truth rewards']}

video_plot(
    outputs=[output_groundtruth, output_sole, output_robometer], 
    plot_save_path='model_outputs/combined/robosuite/lift/unsuccessful/robosuite_lift_episode_11_unsuccessful_max_reward_37.mp4', 
    video_view_external_path=video_paths_external_view[0], 
    video_view_wrist_path=video_paths_wrist_view[0],
    task_description=task_description,
    verbose=False
)

Reward generation and plotting across many videos

from rewardgen import generate
import glob
import json

video_paths = glob.glob('test_videos/robosuite/lift/unsuccessful/*.mp4')
task_description = "Pick up the cube from the table."

## REWARD GENERATION
# SOLE-R1 for all videos
response = generate(model="SOLE-R1",  task_description=task_description, video_paths=video_paths, view_type='external and wrist')

## PLOTTING
plot_save_dir = 'model_outputs/sole-r1/'
for video_idx in range(len(video_paths)):
    output_sole = {"model": "SOLE-R1", "rewards": response.rewards[video_idx]}
    # Optional: Ground-truth rewards (available for test videos from sim environments)
    with open(video_paths[video_idx].replace(".mp4", "/data.json"), 'r') as f:
        data = json.load(f)
    
    output_groundtruth = {"model": "Ground truth", "rewards": data['ground-truth rewards']}
    video_plot(
        outputs = [output_groundtruth, output_sole], 
        plot_save_path = plot_save_dir + video_paths[video_idx].split('test_videos/')[-1] , 
        video_path = video_paths[video_idx],
        task_description=task_description,
        verbose = False
    )


Acknowledgements

RewardGen builds upon the following repos:

Also thank you to Jack Vial for the SO-101 videos.

Release files for rewardgen 0.1.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rewardgen 0.1.1.3
File Size Uploaded
rewardgen-0.1.1.3.tar.gz 674.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rewardgen 0.1.1.3
File Interpreter ABI Platform
rewardgen-0.1.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 1.4 MB

Release files / rewardgen-0.1.1.3.tar.gz

Download URL rewardgen-0.1.1.3.tar.gz
Size 674.6 kB
Tags Source
SHA-256 checksum
How to use checksums
b2c6a8fcf7fa227ae186e7cee20ad824999f7ae106482b06ae61f77a21060335
BLAKE2b-256 checksum
How to use checksums
fe189c4d25db89bec9bcddb6346c26249e5e08f834b4decc29ca107be541edcf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.20

Release files / rewardgen-0.1.1.3-py3-none-any.whl

Download URL rewardgen-0.1.1.3-py3-none-any.whl
Size 754.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
843a92aea210f79b98033c426c1ac6d870c5a9d144ddb5ae3bf2a85f860556e5
BLAKE2b-256 checksum
How to use checksums
9aa86ace57b961d82f097ca0285fbddabe2885f62db96b0c59f44e61212ec79c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.10.20

Release history Release notifications | RSS feed

This release

0.1.1.3 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page