Skip to main content

RewardGuard

Trust Your AI Training

RewardGuard is an AI alignment and safety tooling company focused on reinforcement learning systems. We provide reward auditing libraries that help developers detect reward hacking, misalignment, and training degradation early in the training process.


🎯 What is RewardGuard?

RewardGuard analyzes your RL training logs and ensures your reward functions are balanced and aligned with your intended goals. It detects when agents find unintended ways to maximize rewards (reward hacking) and provides actionable insights to fix them.

Key Features

  • Reward Distribution Analysis - Understand how rewards are distributed across different sources
  • Imbalance Detection - Automatically detect when reward components are misaligned
  • Training Diagnostics - Monitor trends and catch training issues early
  • Actionable Recommendations - Get clear suggestions on how to fix imbalances
  • Auto-Adjustment (Premium) - Automatically rebalance rewards during training

📦 Two Versions

🟢 Free Version

What it does:

  • Analyzes reward distributions
  • Detects imbalances and dominance patterns
  • Provides warnings and recommendations
  • Generates detailed reports

What it doesn't do:

  • Does NOT modify training behavior
  • Read-only analysis and insights

Installation:

pip install rewardguard

🔒 Premium Version (Private)

Everything in Free, PLUS:

  • Automatic reward rebalancing
  • Live monitoring during training
  • Guardrails against reward hacking
  • Continuous alignment enforcement
  • Production-safe controls

Installation:

pip install rewardguard-premium
# Then sign in once per machine:
rewardguard-premium login
# (requires an active Premium subscription)

🚀 Quick Start

Free Version Example

from rewardguard import RewardGuard

# Initialize
guard = RewardGuard(tolerance=5.0)

# Parse your training logs
episodes = guard.parse_logs(raw_log_text)

# Define expected distribution
expected = {
    "reward_a": 60.0,  # Want 60% from component A
    "reward_b": 40.0   # Want 40% from component B
}

# Analyze balance
result = guard.analyze_balance(episodes, expected)

# Print report
guard.print_analysis_report(result)

Output:

============================================================
REWARDGUARD ANALYSIS REPORT
  *** OVERALL SEVERITY: WARNING ***
============================================================

  Episodes analyzed : 50
  Sources found     : reward_a, reward_b

  Source          Real %     Expected %   Diff       Severity
  --------------- ---------- ------------ ---------- --------
  reward_a        72.0       60.0             +12.0  WARNING
  reward_b        28.0       40.0             -12.0  WARNING

  Suggested weight multipliers:
    reward_a: 0.95x  <-- ADJUST
    reward_b: 1.13x  <-- ADJUST

  Actions needed:
    • reward_a: Decrease weight by ~12.0%
    • reward_b: Increase weight by ~12.0%
============================================================

Training-log format

parse_logs() (and RewardGuard.parse_logs) expect each episode to be written as a header line followed by a sources: block. The parser is strict about this layout:

Ep 1 | SUCCESS | reward=15.3
=== EPISODE DATA ===
sources:
  food: 12.1
  survival: 3.2

Ep 2 | FAILURE | reward=8.7
=== EPISODE DATA ===
sources:
  food: 7.0
  survival: 1.7

Rules:

  • Episode header — Ep <n> | <STATUS> | reward=<total>. <n> is an integer, <STATUS> is a single word (e.g. SUCCESS, FAILURE), and reward= is the episode total reward.
  • === EPISODE DATA === marks the start of a block.
  • sources: begins the per-component reward values. Each line is name: value, where name matches [A-Za-z0-9_]+ (a single word, no spaces) and value is a number. These are the values RewardGuard analyzes.
  • percentages: (optional) — a precomputed percentage block in the same name: value form. If omitted, percentages are derived from sources:.
  • A blank line ends the current block; the next Ep ... line begins a new episode.

An episode is kept only if it has an Ep header and at least one entry under sources:. If you see error: no episode data found in log, check that your source lines use name: value with a single-word name.

Don't want to format logs at all? Use the in-loop Monitor instead — it records raw reward dicts straight from your training loop, no log files needed:

import rewardguard as rg

monitor = rg.Monitor(expected={"food": 0.6, "survival": 0.4}, tolerance=5.0)
for step_rewards in training_loop():     # e.g. {"food": 1.0, "survival": 0.01}
    monitor.step(step_rewards)
monitor.print_report()

Premium Version Example

The premium tier installs as a separate package (rewardguard_premium) and adds the AutoMonitor — it learns a baseline, flags drift with z-scores, and can auto-correct reward weights live during training:

from rewardguard_premium import AutoMonitor

monitor = AutoMonitor(
    expected={"task": 0.7, "safety": 0.3},
    baseline_steps=300,   # warm-up before detection activates
    auto_correct=True,    # adjust weights automatically when flagged
)

for episode in range(num_episodes):
    for step in range(max_steps):
        r_task, r_safety = env.step(action)
        snapshot = monitor.step({"task": r_task, "safety": r_safety})

        if snapshot and snapshot.flag == "critical":
            # Apply auto-corrected weights back to the environment
            env.set_reward_weights(monitor.weights)

monitor.print_report()

See the premium README for the full AutoMonitor API, framework integrations (W&B / TensorBoard / SB3), and save/load.


📖 Use Cases

1. Game AI

Ensure your game AI learns to play properly, not exploit bugs:

  • Detect when agents farm easy points instead of completing objectives
  • Balance combat vs exploration rewards
  • Prevent exploit-based strategies

2. Robotics

Keep robots aligned with safety and task completion:

  • Balance speed vs safety rewards
  • Ensure proper task prioritization
  • Detect reward shortcuts

3. Recommendation Systems

Align recommendation rewards with business goals:

  • Balance engagement vs revenue
  • Prevent clickbait optimization
  • Ensure long-term user satisfaction

4. General RL Research

Debug and optimize any RL training:

  • Understand reward dynamics
  • Catch training issues early
  • Validate reward function design

🏗️ How It Works

Free Version (Analysis Only)

  1. Parse Logs - Extracts reward data from training logs
  2. Aggregate - Calculates actual reward distribution
  3. Compare - Compares against your expected distribution
  4. Recommend - Suggests specific weight adjustments

Key Principle: Tells you what's wrong, you fix it manually.

Premium Version (Auto-Fix)

  1. All Free features, PLUS:
  2. Monitor - Tracks performance over time
  3. Detect - Identifies imbalances automatically
  4. Adjust - Modifies reward weights in real-time
  5. Learn - Continuously tunes based on results

Key Principle: Fixes problems for you automatically.


🎓 Philosophy

We believe AI should be:

  • Transparent - You should understand what your AI is learning
  • Aligned - Reward functions should incentivize intended behaviors
  • Safe - Training should be monitored for unintended outcomes

RewardGuard helps ensure your models learn what you intend, not just how to maximize scores.


💰 Pricing

Feature Free Premium
Reward analysis ✅ ✅
Imbalance detection ✅ ✅
Recommendations ✅ ✅
Auto-adjustment ❌ ✅
Live monitoring ❌ ✅
Unlimited training steps ❌ ✅
Priority support ❌ ✅
Price $0/month $99/month

📚 Documentation


🤝 Support

  • Community (Free): [Discord/Forum Link]
  • Email (Premium): giovan@rewardguard.dev
  • Chat (Premium): Available in dashboard

📄 License

  • Free Version: MIT License
  • Premium Version: Proprietary

🚧 Roadmap

  • Support for more log formats
  • Built-in visualization dashboard
  • Integration with popular RL frameworks (Stable-Baselines3, RLlib)
  • Cloud-based monitoring
  • Team collaboration features
  • Custom alerting rules

⚡ Quick Links


RewardGuard © 2026 | Trust Your AI

Metadata

Release files for rewardguard 1.0.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rewardguard 1.0.5
File Size Uploaded
rewardguard-1.0.5.tar.gz 20.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rewardguard 1.0.5
File Interpreter ABI Platform
rewardguard-1.0.5-py3-none-any.whl Python 3 none any Details

Total release size: 37.4 kB

Release files / rewardguard-1.0.5.tar.gz

Download URL rewardguard-1.0.5.tar.gz
Size 20.0 kB
Tags Source
SHA-256 checksum
How to use checksums
c423814d8140c8aa4cacce2d90c7e5c1914a8282dabb6caad35e074c715d91df
BLAKE2b-256 checksum
How to use checksums
6a18be9c3635d57d3391497ce3bd98656c6a1aadcd4f637c45a51e51c340ebaa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.13

Release files / rewardguard-1.0.5-py3-none-any.whl

Download URL rewardguard-1.0.5-py3-none-any.whl
Size 17.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
59d106a8d4ae6eeaf76779471614f919e6db44d5fa2b1eb48ad45477ea1e4dec
BLAKE2b-256 checksum
How to use checksums
8355a3599f5407170b065eede955beb54bc030534ba47a9fbbe18a83bb4a0635
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.13

Release history Release notifications | RSS feed

This release

1.0.5 This release

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page