Skip to main content

llm-behavior-eval · GitHub license pypi version

Deploy docs pyrefly ruff Unit tests Vulnerability scan

A Python 3.10+ toolkit for measuring undesirable LLM behaviors with Hugging Face models or local model paths.

It evaluates bias, hallucinations, prompt injection, refusal behavior, and Chinese censorship. All evaluations support Transformers instruct models. See Tested on for the models used to validate the pipeline.

What it evaluates

This toolkit evaluates five classes of behaviors:

  • Bias (BBQ, UNQOVER)

    • BBQ (Bias Benchmark for QA): hand‑crafted questions that probe stereotypes across protected dimensions. Supports paired splits:
      • bias (ambiguous) and unbias (disambiguated) for: gender, race, nationality, physical, age, religion.
      • Only BBQ provides both ambiguous and disambiguated versions.
    • UNQOVER: crowd‑sourced templates probing stereotypes; provides only the ambiguous/bias split for: religion, gender, race, nationality.
  • Hallucinations (HaluEval, Med‑Hallu)

    • HaluEval (halueval): general‑domain factuality/consistency checks.
    • Med‑Hallu (medhallu): medical‑domain hallucination benchmark.
  • Prompt Injection (Purple Llama)

    • Purple Llama Prompt Injection: measures susceptibility to instruction overriding and jailbreaks using curated prompt‑injection attacks. Reuses the hallucination judging pipeline with Yes/No grading.
  • Chinese censorship (CCPC-Bench)

    • CCPC-Bench: measures CCP-aligned political censorship using the pinned 500-row hirundo-io/ccpc-500 test split. This implements task version 6 for compatibility with the public benchmark. The frozen ccpc_binary_v4 judge emits bare True or False; raw metrics.csv reports fractional censorship_rate, stderr, and judged_fraction. Presentation summaries (summary_brief.csv and summary_full.csv) report Censorship rate (%) ⬇️ on the 0–100 scale, matching the Purple Llama percentage/direction convention. Compatible summary consumers should read that header without multiplying by 100; consumers requiring the previous fractional field should use metrics.csv. Version strings and revision hashes are unchanged.
    • CCPC upload through the bundled hirundo_notion_tools uploader is not yet supported.
  • Refusal (XSTest, OR-Bench)

    • XSTest: tests exaggerated safety behavior. It combines prompts that a model should answer with prompts that it should refuse, so it exposes both over-refusal on safe requests and under-refusal on unsafe requests.
    • OR-Bench: evaluates appropriate refusal behavior using the same safe/unsafe distinction. The evaluator counts both direct and indirect refusals as refusals.

Dataset identifiers:

  • BBQ: hirundo-io/bbq-<bias_type>-<bias|unbias>-free-text
  • UNQOVER: unqover/unqover-<bias_type>-bias-free-text
  • HaluEval: hirundo-io/halueval
  • Med‑Hallu: hirundo-io/medhallu
  • Prompt Injection (Purple Llama): hirundo-io/prompt-injection-purple-llama
  • Chinese censorship (CCPC-Bench): CLI preset chinese_censorship; Hugging Face repository hirundo-io/ccpc-500
  • XSTest: hirundo-io/XSTest
  • OR-Bench: hirundo-io/or-bench

Pass the behavior preset as the second positional CLI argument:

  • BBQ: bias:<bias_type> or unbias:<bias_type>
  • UNQOVER: unqover:bias:<bias_type>
  • Hallucinations:
    • HaluEval: hallu
    • Med‑Hallu: hallu-med
  • Prompt Injection:
    • Purple Llama: prompt-injection
  • Chinese censorship:
    • CCPC-Bench: chinese_censorship with the pinned --judge-model google/gemma-4-26B-A4B-it
  • Refusal:
    • XSTest: refusal:xstest
    • OR-Bench: refusal:orbench
    • Both: refusal:all

You can also run across all supported bias types using all:

  • BBQ (all ambiguous/bias splits): bias:all
  • BBQ (all unambiguous/unbias splits): unbias:all
  • UNQOVER (all bias splits): unqover:bias:all

Installation

Install Python 3.10.12 through 3.13, then create a virtual environment and install the package:

# Create and activate a virtual environment.
python -m venv .venv
source .venv/bin/activate

# Install the package.
pip install llm-behavior-eval

If you use uv, run uv venv .venv and uv pip install llm-behavior-eval instead. For a local checkout, install with pip install -e . or uv pip install -e ..

vLLM extra

The vllm extra is pinned to vllm>=0.23.0,<0.24 — the tested line for the text-only Gemma-4 judge (runner="generate", language_model_only=True) on torch==2.11. The upper bound is deliberate: newer vLLM releases require torch>=2.13 / cu13x wheels, so an open floor would silently pull an incompatible stack. The extra is optional — if the vLLM stack doesn't fit your environment, run the judge on the transformers backend (--judge-engine transformers), which needs no vLLM install.

The base transformers floor is >=5.10.4 for the same reason: that is the oldest release verified to load the gemma4_unified config. vLLM 0.23 itself allows transformers>=4.56.0 and its registry does contain Gemma4UnifiedForConditionalGeneration, so the architecture is supported — but an older transformers fails to recognise the config and the engine never starts.

Development Container

The repository ships a VS Code Dev Container definition (.devcontainer/). The setup script installs the base project dependencies to keep the image lean. If you need optional extras (for example MLflow or vLLM), set LLM_BEHAVIOR_EVAL_INSTALL_EXTRAS before the container runs:

# Example: install MLflow extra inside the devcontainer
export LLM_BEHAVIOR_EVAL_INSTALL_EXTRAS="mlflow"
bash .devcontainer/setup.sh

# Example: install both MLflow and vLLM (requires more disk space)
export LLM_BEHAVIOR_EVAL_INSTALL_EXTRAS="mlflow,vllm"
bash .devcontainer/setup.sh

If the requested extras exhaust the available disk, the script falls back to a base install so the container remains usable. Re-run the script with a smaller set of extras when needed.

Run the Evaluator

Use the CLI with the required model and behavior positional arguments. The behavior preset selects datasets for you.

llm-behavior-eval <model_repo_or_path> <behavior_preset>

You can pass comma-separated presets from one evaluator family, such as bias:gender,unbias:gender. The CLI runs one evaluator family per invocation and rejects mixed families because they use different evaluator and scoring paths. Run each family separately.

Examples

  • BBQ (bias) — evaluate a model on a biased split (free‑text):
llm-behavior-eval google/gemma-2b-it bias:gender
  • BBQ (unbias) — evaluate a model on an unambiguous split:
llm-behavior-eval meta-llama/Llama-3.1-8B-Instruct unbias:race
  • UNQOVER (bias) — use UNQOVER source datasets (UNQOVER does not support 'unbias'):
llm-behavior-eval google/gemma-2b-it unqover:bias:gender
  • BBQ (all bias types) — iterate all BBQ ambiguous splits:
llm-behavior-eval meta-llama/Llama-3.1-8B-Instruct bias:all
  • UNQOVER (all bias types) — iterate all UNQOVER bias splits:
llm-behavior-eval meta-llama/Llama-3.1-8B-Instruct unqover:bias:all
  • Hallucination (general) — HaluEval free‑text:
llm-behavior-eval google/gemma-2b-it hallu
  • Hallucination (medical) — Med-Hallu:
llm-behavior-eval meta-llama/Llama-3.1-8B-Instruct hallu-med
  • Prompt Injection — Purple Llama prompt injections:
llm-behavior-eval meta-llama/Llama-3.1-8B-Instruct prompt-injection
  • Chinese censorship. Run CCPC-Bench with an explicitly selected judge:
llm-behavior-eval google/gemma-3-12b-it chinese_censorship \
  --judge-model google/gemma-4-26B-A4B-it
  • XSTest (refusal) — measure over-refusal on safe prompts and refusal on unsafe prompts:
llm-behavior-eval meta-llama/Llama-3.1-8B-Instruct refusal:xstest
  • OR-Bench (refusal) — evaluate appropriate direct or indirect refusal behavior:
llm-behavior-eval meta-llama/Llama-3.1-8B-Instruct refusal:orbench
  • All refusal benchmarks — run XSTest and OR-Bench sequentially:
llm-behavior-eval meta-llama/Llama-3.1-8B-Instruct refusal:all

Interpreting refusal results

Both datasets label prompts as safe (the model should answer) or unsafe (the model should refuse). The evaluator uses a judge model to classify the generated answer. For XSTest, full and partial refusals count as refusals; for OR-Bench, both direct and indirect refusals count as refusals.

Use the two primary metrics together:

  • Safe refusal rate (%) ⬇️: the share of safe prompts that were refused. Lower is better because it measures over-refusal.
  • Unsafe refusal rate (%) ⬆️: the share of unsafe prompts that were refused. Higher is better because it measures appropriate refusal.
  • Heuristic safe refusal rate (%) ⬇️ and heuristic unsafe refusal rate (%) ⬆️: the same rates estimated from refusal phrases in each answer, without a judge. Use them as a quick diagnostic, not as the benchmark result.
  • Incomplete response rate (%) ⬇️: the share of configured samples where generation stopped because it reached the length limit. Truncated answers are still judged, so they can count toward the refusal rates.
  • Unknown finish reason rate (%) ⬇️: the share of configured samples with a finish reason other than a normal stop or a length limit. Those responses are not judged.
  • Judge unparseable rate (%) ⬇️: the share of configured samples where the judge did not produce a recognized refusal class. Those responses are excluded from the judge-based refusal rates.

The diagnostic rates use the configured sample count as their denominator. Unknown or unparseable responses reduce the judged sample set used for the primary rates.

CLI options

  • --max-samples <N> — cap how many rows to evaluate per dataset (defaults to 500). Use 0 or any negative value to run the entire split.
  • --use-4bit-judge/--no-use-4bit-judge — toggle 4-bit (bitsandbytes) loading for the judge model so you can keep the evaluator in full precision while fitting the judge onto smaller GPUs.
  • --model-token / --judge-token — supply Hugging Face credentials for the evaluated or judge models (the judge token defaults to the model token when omitted).
  • --judge-model — pick a different judge checkpoint; the default is google/gemma-3-12b-it.
  • --inference-engine vllm / --inference-engine transformers — switch between vLLM and transformers backends for the evaluated model. There are also --model-engine and --judge-engine flags for more explicit control.
  • --vllm-max-model-len / --vllm-gpu-memory-utilization — configure vLLM's maximum context length and GPU memory utilization. Leave the maximum length unset to use the model's native context; the GPU utilization default is 0.8. Override either only after confirming the target GPU's KV-cache capacity; increasing utilization increases that capacity, while lowering it decreases available KV-cache capacity.
  • --vllm-tokenizer-mode, --vllm-config-format, --vllm-load-format — forward advanced knobs directly to the underlying vLLM engine when you need to align tokenizer behavior, checkpoint formats, or tool-calling semantics with a particular deployment. Tokenizer mode accepts auto, slow, mistral, or custom.
  • --thinking-on/--thinking-off — enable thinking modes on tokenizers that support them. Unset uses the evaluator-family default: on for refusal, off for other behaviors. The judge always runs with thinking off.
  • --enable-thinking-arg-name — enable thinking argument name in tokenizer's apply_chat_template (e.g. 'enable_thinking').
  • --thinking-start-token / --thinking-end-token — Thinking start/end token to use for the model (e.g. ''/'').
  • --use-mlflow plus --mlflow-tracking-uri, --mlflow-experiment-name, and --mlflow-run-name — configure MLflow tracking for the run.

Need more control or wrappers around the library? Explore the scripts in examples/ to see how to call the evaluators from Python directly, customize additional knobs, or embed the run inside your own orchestration logic.

See examples/presets_customization.py for a minimal script-based workflow.

MLflow Integration (Optional)

Enable MLflow tracking with --use-mlflow to log simple parameters, metrics and artifacts.

Install: pip install llm-behavior-eval[mlflow] or pip install mlflow.

CLI example:

llm-behavior-eval google/gemma-2b-it bias:gender --use-mlflow

To find more documentation: see MLFLOW_INTEGRATION.md. Programmatic example: see examples/mlflow_example.py.

Output

Evaluation reports are saved as metrics CSV files and full response JSON files in the results directory. By default, the CLI writes to:

  • macOS: ~/Library/Application Support/llm-behavior-eval/results
  • Linux/Ubuntu: $XDG_DATA_HOME/llm-behavior-eval/results (or ~/.local/share/llm-behavior-eval/results if XDG_DATA_HOME is unset)
  • Windows: %LOCALAPPDATA%\llm-behavior-eval\results (fallback: %APPDATA%\llm-behavior-eval\results)

Override the default with --base-output-dir when you need a different path. You can also use --model-output-dir to explicitly override the name of the model under that base path; otherwise, the model path or repo ID will be used, with an added stub if using a LoRA adapter.

Outputs are organised as results/<model>/<dataset>_<dataset_type>_<text_format>/. Per‑model summaries are saved as results/<model>/summary_full.csv (full metrics) and results/<model>/summary_brief.csv.

summary_brief.csv contains the following columns: Dataset, Thinking, and one or more metric columns (Accuracy/Error/Attack success rate). Labels are inferred as follows:

  • BBQ: BBQ: <gender|race|nationality|physical|age|religion> <bias|unbias>
  • UNQOVER: UNQOVER: <religion|gender|race|nationality> <bias>
  • Hallucination: halueval or medhallu
  • Prompt Injection: prompt-injection-purple-llama
  • Chinese censorship: chinese_censorship
  • Refusal: XSTest or or-bench

Tested on

Validated the pipeline on the following models:

  • "google/gemma-3-12b-it"

  • "meta-llama/Meta-Llama-3.1-8B-Instruct"

  • "meta-llama/Llama-3.2-3B-Instruct"

  • "google/gemma-7b-it"

  • "google/gemma-2b-it"

  • "google/gemma-3-4b-it"

Using the next models as judges:

  • "google/gemma-3-12b-it"

  • "meta-llama/Llama-3.3-70B-Instruct"

License

This project is licensed under the MIT License. See the LICENSE file for more information.

Metadata

Release files for llm-behavior-eval 0.1.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-behavior-eval 0.1.9
File Size Uploaded
llm_behavior_eval-0.1.9.tar.gz 111.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-behavior-eval 0.1.9
File Interpreter ABI Platform
llm_behavior_eval-0.1.9-py3-none-any.whl Python 3 none any Details

Total release size: 196.0 kB

Release files / llm_behavior_eval-0.1.9.tar.gz

Download URL llm_behavior_eval-0.1.9.tar.gz
Size 111.8 kB
Tags Source
SHA-256 checksum
How to use checksums
27354247d0d656ec6498409874e2b7250966f143980659e583b4d40eabd69dab
BLAKE2b-256 checksum
How to use checksums
531668b2bb1a378229781140aa4296a6bc5b3ac02fba945f42b0b031762642e2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release files / llm_behavior_eval-0.1.9-py3-none-any.whl

Download URL llm_behavior_eval-0.1.9-py3-none-any.whl
Size 84.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8d10050d7acb767b9081b0053b6c636c0804cdeda3e85fef6490107f60bac022
BLAKE2b-256 checksum
How to use checksums
fd180d979d60f37a7b54ff373c101f7ac71a73793cf225af4fd96d075b8f2d50
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.9 This release

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.3

2 release files

0.1.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page