Skip to main content

A collection of judges for evaluating language model generations

Project description

JudgeZoo

This repo provides access to a set of commonly used LLM-based safety judges via a simple and consistent API. Our main focus is ease-of-use, correctness, and reproducibility.

How to use

You can create a judge model instance with a single line of code:

judge = Judge.from_name("strong_reject")

To get safety scores, just pass a list of conversations to score:

harmless_conversation = [
    {"role": "user", "content": "How do I make a birthday cake?"},
    {"role": "assistant", "content": "Step 1: Collect ingredients..."}
]

scores = judge([harmless_conversation])
print(scores)
>>> {"p_harmful": [0.02496337890625]}

All judgezoo return "p_harmful", which is a normalized score from 0 to 1. Depending on the original setup, the judge may also return discrete scores or harm categories (e.g. on a Likert scale). In these cases, the raw scores are also returned:

judge = Judge.from_name("adaptive_attacks")

scores = judge([harmful_conversation])
print(scores)
>>> {"p_harmful": 0.0, "rating": "1"}

Included judges

Name Argument Creator (Org/Researcher) Link to Paper Type Fine-tuned from
Adaptive Attacks adaptive_attacks Andriushchenko et al. (2024) arXiv:2404.02151 prompt-based
AdvPrefix advprefix Zhu et al. (2024) arXiv:2412.10321 prompt-based
AegisGuard* aegis_guard Ghosh et al. (2024) arXiv:2404.05993 fine-tuned LlamaGuard 7B
HarmBench harmbench Mazeika et al. (2024) arXiv:2402.04249 fine-tuned Gemma 2B
JailJudge jail_judge Liu et al. (2024) arXiv:2410.12855 fine-tuned Llama 2 7B
Llama Guard 3 llama_guard_3 Llama Team, AI @ Meta (2024) arXiv:2407.21783 fine-tuned Llama 3 8B
Llama Guard 4 llama_guard_4 Llama Team, AI @ Meta (2024) Meta blog fine-tuned Llama 4 12B
MD-Judge (v0.1 & v0.2) md_judge Li, Lijun et al. (2024) arXiv:2402.05044 fine-tuned Mistral-7B/LMintern2 7B
StrongREJECT strong_reject Souly et al. (2024) arXiv:2402.10260 fine-tuned Gemma 2b
StrongREJECT (rubric) strong_reject_rubric Souly et al. (2024) arXiv:2402.10260 prompt-based -
XSTestJudge xstest Röttge et al. (2023) arXiv:2308.01263 prompt-based

* there are two versions of this judge (permissive and defensive). You can switch between them using Judge.from_name("aegis_guard", defensive=[True/False])

Other

Prompt-based judges

While some judges (such as the HarmBench classifier) are finetuned local models, others rely on prompted foundation models. Currently, we support local foundation models and OpenAI models:

judge = Judge.from_name("adaptive_attacks", use_local_model=False, remote_foundation_model="gpt-4o")

scores = judge([harmless_conversation])
print(scores)
>>> {"p_harmful": 0.0, "rating": "1"}
judge = Judge.from_name("adaptive_attacks", use_local_model=True)

scores = judge([harmless_conversation])
print(scores)
>>> {"p_harmful": 0.0, "rating": "1"}

When not specified, the defaults in config.py are used.

Multi-turn interaction

Judges vary in how much of a conversation they can evaluate - many models only work for single-turn interactions. In these cases, we assume the first user message to be the prompt and the final assistant message to be the response to be judged. If you prefer a different setup, you can pass only single-turn conversations.

Reproducibility

Wherever possible, we use official code directly provided by the original authors to ensure correctness.

Finally, we warn if a user's setup diverges from the original implementation:

from judgezoo import Judge

judge = Judge.from_name("intention_analysis")
>>> WARNING:root:IntentionAnalysisJudge originally used gpt-3.5-turbo-0613, you are using gpt-4o. Results may differ from the original paper.

Installation (not on pypi yet)

pip install judgezoo

Tests

To run all tests, run

pytest tests/ --runslow

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

judgezoo-0.1.0.tar.gz (56.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

judgezoo-0.1.0-py3-none-any.whl (38.0 kB view details)

Uploaded Python 3

File details

Details for the file judgezoo-0.1.0.tar.gz.

File metadata

  • Download URL: judgezoo-0.1.0.tar.gz
  • Upload date:
  • Size: 56.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.2

File hashes

Hashes for judgezoo-0.1.0.tar.gz
Algorithm Hash digest
SHA256 4f4f23fbdd304fedce62203c280832c73f5865131d6c7a0706733fb2b39c2228
MD5 6933311e1bfdffbcb6398b26fb05eeb8
BLAKE2b-256 038ed66540aeda13091452fc3aafa712ee6cad9dd4c3e9fd2d564646d6e53c91

See more details on using hashes here.

File details

Details for the file judgezoo-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: judgezoo-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 38.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.2

File hashes

Hashes for judgezoo-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 965ffeeb34559010e8416b88eda376f367410842a55e6d80eb4970ede9a216ae
MD5 a09cd66135a14f7a87bc5c133629c14e
BLAKE2b-256 56229dc42d6b1be2eb9e38ca890aa4776669a3d7f652ec58e7db3913207b783d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page