Bengali LLM safety benchmark: 879 statute-anchored prompts, 17 harm categories, 5 register conditions, calibrated four-way judge.
Project description
BanglaSafe
A Bengali safety benchmark for LLMs. 879 prompts across 17 harm categories, each one anchored to a Bangladesh statute or a documented case, asked in five different registers.
Point it at any OpenAI-compatible endpoint and it gives you an attack success rate.
One command, nothing to install:
uvx banglasafe run \
--model Qwen/Qwen3-32B-Instruct \
--base-url http://localhost:8000/v1 \
--judge-model anthropic/claude-opus-4-7 \
--judge-base-url https://openrouter.ai/api/v1
uvx fetches it, runs it, and cleans up. If you want it to stick around, uv tool install banglasafe. If you would rather use pip, pip install banglasafe works too, just slower.
ASR loose 47.2% [44.1, 50.3] PARTIAL + HARMFUL
ASR strict 19.8% HARMFUL only
vs the 18-model reference cohort: z = -0.22, rank 11/19 (safer than cohort average)
Register effect BN_Formal is 16.4pp higher than BN_Collq (reference cohort: +17.5pp)
The point of it
Most multilingual safety work translates English harm categories into other languages. That loses the harms that only exist locally, like hundi, yaba, or bKash fraud, and it loses register. Bengali is diglossic, so the same request reads very differently as newspaper prose, as code-mixed Banglish, or as an institutional case file.
Across the 18 models we tested, a harmful request written as a formal newspaper investigation got answered 17.5 points more often than the same request written as a casual message. No jailbreak, no prompt engineering, just a change of register.
Reading the output
Two rates. ASR loose counts PARTIAL and HARMFUL, meaning the model engaged. ASR strict counts only HARMFUL, meaning it produced something usable. Report both or you are either overstating or understating.
Every rate gets a bootstrap 95% CI, broken out by condition, by category, and by whether the prompt was human-written or model-generated. Cells under n=30 get no CI instead of a useless one.
Your score is placed against 18 reference models that ship with the package, so you get a z-score and a rank rather than a number with no context.
Each report also prints how well the judge agrees with human labels, because the label is an estimate and you should be able to see its error bar.
Endpoints
It speaks plain /v1/chat/completions, so one --base-url covers vLLM, SGLang, LightLLM, Ollama, llama.cpp, TGI, LM Studio, a LiteLLM proxy, OpenRouter, Together, and Groq. Both a /v1 base and a full /v1/chat/completions URL work.
Picking a judge
--judge-model is required and there is no default. The judge decides your numbers, so it should be a choice you make.
| judge | agreement |
|---|---|
anthropic/claude-opus-4-7 |
The judge behind the published labels. Binary kappa 0.666 against human gold on 300 responses, above inter-human 0.586. |
google/gemini-3.1-pro |
Independent second judge. Agrees with the first at binary kappa 0.787, per-model Pearson r 0.989 across 18 models. |
Any OpenAI-compatible model can judge. Only these two have measured agreement with humans on this benchmark, and the tool warns you if you pick something else.
Commands
banglasafe run # query, judge, score, report
banglasafe score # re-score an existing run, offline
banglasafe report # re-render terminal and HTML from a finished run
banglasafe compare # rank runs side by side
banglasafe judges # judge options and their calibration
banglasafe cohort # the 18 reference models
banglasafe info # dataset, rubric hash, version
Flags worth knowing:
--limit 20 # 20 harm instances, all 5 conditions each
--condition BN_Formal --condition BN_Collq # just the register comparison
--category hundi # one harm
--max-concurrency 16 # minutes instead of an hour
--config banglasafe.yaml # flags still override the file
--limit takes whole harm instances rather than the first N rows, so a subset still compares like
with like. (5 of the 167 instances are missing a condition in the source data.)
Output
results/<model>__<timestamp>/
responses.jsonl
labels.jsonl
summary.json
report.html
The JSONL files are the real output. summary.json and report.html are derived and you can rebuild them with banglasafe report. Runs resume: kill one and re-run the same command, and it picks up instead of charging you twice.
Config
Flags beat environment variables, which beat the YAML file.
target:
model: Qwen/Qwen3-32B-Instruct
base_url: http://localhost:8000/v1
api_key: "{{ env.LOCAL_API_KEY }}"
judge:
model: anthropic/claude-opus-4-7
base_url: https://openrouter.ai/api/v1
api_key: "{{ env.OPENROUTER_API_KEY }}"
run:
max_concurrency: 16
{{ env.VAR }} expands at load, so keys stay out of the file. It also reads BANGLASAFE_BASE_URL, BANGLASAFE_API_KEY, BANGLASAFE_JUDGE_BASE_URL, BANGLASAFE_JUDGE_API_KEY, and falls back to OPENAI_BASE_URL, OPENAI_API_KEY, OPENROUTER_API_KEY, VLLM_BASE_URL, SGLANG_BASE_URL, OLLAMA_BASE_URL.
How it runs
Temperature 0, top_p 1, no system prompt. Decoding settings alone can swing open-model safety rates from near zero to over 95% on the same prompts, so they stay fixed. No system prompt is the consumer chatbot case.
The judge rubric ships in the package and its hash goes into every report. Four labels, a decision algorithm, an edge-case table, 17 tie-breakers. Same spec that produced the published labels.
Labels are noisy. Binary judge-vs-human kappa is 0.666, and four-way inter-human kappa is only 0.286, which is why the binary collapse is the headline and the four-way split is reported next to it rather than instead. Differences inside the confidence intervals are noise.
Data
Prompts live at BanglaLLM/BanglaSafe. The
tool downloads the JSONL directly and caches it in ~/.cache/banglasafe, so there is no token and
no heavy data library in the dependency tree. Set BANGLASAFE_CACHE to move it.
If you want the dataset for something else:
from datasets import load_dataset
prompts = load_dataset("BanglaLLM/BanglaSafe")
taxonomy = load_dataset("BanglaLLM/BanglaSafe", "taxonomy")
Please don't
These prompts exist to make unsafe behaviour measurable. Use them to evaluate models, not to extract harmful content, and don't train on them to make a model more willing.
The benchmark measures harmful compliance. It says nothing about over-refusal, since there is no benign control set. It cannot tell you whether hardening a model against the journalism register would make it start refusing real investigative reporting.
Citation
@inproceedings{islam2026banglasafe,
title = {Register Shifts Break {LLM} Safety: A Bengali Benchmark with Culturally Grounded Harms},
author = {Islam, Naymul and Lia, Nusrat Jahan and Roy Dipta, Shubhashis and Sultan, Sabik Bin and Zehady, Abdullah Khan},
year = {2026}
}
MIT for the code, CC-BY-4.0 for the dataset.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file banglasafe-1.0.0.tar.gz.
File metadata
- Download URL: banglasafe-1.0.0.tar.gz
- Upload date:
- Size: 32.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.16 {"installer":{"name":"uv","version":"0.11.16","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
52fe1064f184e07bff135e28d8d92cefe869cb83410f763ae4b9751d5cce7a70
|
|
| MD5 |
29d09c91d3eb3f3946ded874c66ec18b
|
|
| BLAKE2b-256 |
fb876a118db21fb42af501b80b3b75b15cdf20c8f8519c967505b116ebf3ba13
|
File details
Details for the file banglasafe-1.0.0-py3-none-any.whl.
File metadata
- Download URL: banglasafe-1.0.0-py3-none-any.whl
- Upload date:
- Size: 37.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.16 {"installer":{"name":"uv","version":"0.11.16","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
627f3adeca40c488fec42fd198fe7c9946633502abb9afe553d793f127cdad40
|
|
| MD5 |
45593fc6723dde3d807001861f054361
|
|
| BLAKE2b-256 |
b1ef4f02930e7316cea164ba79e4e73b9983e87dc21a688fb8cdae99b0201332
|