FairLMs
Fairness definitions and bias metrics for large language models — a companion library for studying and evaluating bias in LMs.
The long-term goal is a stable, sklearn-style API: import a metric, call compute(...), and get a result — without caring where the implementation lives.
Status: All 33 metrics share one contract — configuration in the constructor, data as a validated container passed to
compute(model, data). Leafmain.pyfiles andexamples/are short public-API demos. Implementation math still lives underfairlms/definition/.
Install
pip install fairlms
That's all most users need — fairlms is on PyPI,
so installing does not require access to this repository.
The project is FairLMs; both the PyPI distribution and the import name are
fairlms, all-lowercase per PEP 8.
Optional extras
pip install "fairlms[openai]" # API-backed decoder metrics (CR, CTF, BA)
pip install "fairlms[dev]" # pytest
pip install "fairlms[all]" # openai + common extras
From source
To track unreleased work, or to develop the library:
git clone https://github.com/michaellarionov/FairLMs.git && cd FairLMs && pip install -e ".[dev]"
An editable install means your edits take effect immediately, with no reinstall. You can also install a specific commit or tag directly:
pip install "git+https://github.com/michaellarionov/FairLMs.git@v0.3.1"
Prefer a tag over @main if you do this: main tracks unreleased work, so any
push can change behaviour under you. Tags don't move.
The import name is fairlms regardless of the repository name.
Requirements
Python ≥ 3.10 — the floor for torch, transformers and datasets.
Installed automatically: torch, transformers, datasets, numpy, pandas,
scipy, scikit-learn, wordfreq, nltk. The last three are needed at import
time by counterfactual_auc / normalized_position_distance,
lexical_frequency_proportion, and morphological_choice_divergence
respectively.
Versioning
Semantic versioning, currently pre-1.0 — the public API may still change between
minor versions. The version is single-sourced in
fairlms/_version.py; pyproject.toml reads it via
[tool.setuptools.dynamic], so bump that one value.
To cut a release, bump _version.py, then tag. Pushing a v* tag triggers
.github/workflows/publish.yml, which builds,
validates with twine check --strict, checks the tag matches the version, and
uploads to PyPI via Trusted Publishing (OIDC — no API token is stored anywhere):
git tag -a "v$(python scripts/package_version.py)" -m "Release" && git push origin --tags
Users then upgrade with pip install --upgrade fairlms.
Breaking changes
| Version | Change | Migration |
|---|---|---|
| 0.3.0 | Project renamed fairllms → fairlms |
pip install fairlms, import fairlms |
| 0.2.0 | Package renamed fairLLMs → fairllms (PEP 8) |
import fairllms |
| 0.2.0 | Metric config is keyword-only | WEAT(pooling="cls"), not WEAT("cls") |
Metric calling conventions were not broken in 0.2.0: the previous keyword
style (compute(model=m, T1_terms=[...])) still works and emits a
DeprecationWarning naming its replacement. Those shims are scheduled for
removal in a later release, so migrate when convenient.
Quick start
from fairlms.metrics import CrowSPairsScore, list_metrics
from fairlms.datasets import CrowSPairs
from fairlms.models import HuggingFaceModel
model = HuggingFaceModel("bert-base-uncased", task="mlm")
result = CrowSPairsScore().compute(model, CrowSPairs(n_max=50))
print(result.score, result.by_category)
# Discover metrics
print(list_metrics())
The contract
Every metric follows scikit-learn's estimator conventions:
Metric(**config).compute(model, data) -> MetricResult
- Configuration goes in the constructor, keyword-only, and is introspectable
via
get_params()/set_params()— so metrics can be cloned or swept. - Data is the second positional argument: a
FairnessDataset, a plain sequence, or a typed container fromfairlms.metrics.datafor metrics that need several labelled sets. - Unknown keywords raise
TypeErrorinstead of silently using a default.
from fairlms.metrics import WEAT, SEAT, WordSets
words = WordSets(target_1=t1, target_2=t2, attribute_1=a1, attribute_2=a2)
WEAT(n_samples=10_000).compute(model, words)
SEAT(pooling="cls").compute(model, words) # same data, different metric
WEAT(pooling="cls").get_params() # {'n_samples': 10000, 'pooling': 'cls'}
The published association tests ship pre-wrapped in fairlms.data, so a
standard run needs no term lists at all — and swapping the checkpoint does not
change the call:
from fairlms.data import weat_c1, list_word_sets
from fairlms.metrics import SEAT
from fairlms.models import HuggingFaceModel
metric = SEAT(n_samples=1_000)
for ckpt in ("bert-base-uncased", "roberta-base"):
result = metric.compute(HuggingFaceModel(ckpt, task="encoder"), weat_c1)
print(ckpt, result.score, result.details["p_value"])
list_word_sets() # ['weat_c1', …, 'seat_c1', …]
weat_c1 is a WordSets, so user-supplied evidence goes through the exact
same call with no adapter code:
mine = WordSets(target_1=names_a, target_2=names_b,
attribute_1=pleasant, attribute_2=unpleasant)
metric.compute(model, mine)
Containers validate at construction, so mistakes fail immediately:
>>> ContextSets(["a sentence"], ...)
TypeError: target_1 must be a mapping of term -> list of context sentences, got
list. (A flat list of sentences is not accepted; CEAT samples contexts per term,
so terms must be keyed.)
Metrics that need no model
Five metrics score predictions you already have. They also exist as plain
functions, mirroring sklearn.metrics:
from fairlms.metrics import equal_opportunity_gap, accuracy_disparity
equal_opportunity_gap(y_true, y_pred, groups, g1="A", g2="B") # -> float
accuracy_disparity(scores_stereotype, scores_counter) # -> float
Also available: inference_bias_score, fair_inference_score,
context_based_disparity. Use the class form
(EqualOpportunityGap, …) when you want the full MetricResult with
diagnostics.
Dataset diagnostics
Dataset diagnostics are a separate, dataset-first API. They consume explicit
evidence and audit intent rather than a model, and return a structured report
whose components can be ready, blocked, or not_applicable. Missing
evidence is never represented as a score of zero.
The first diagnostic is axis-level representativeness (b_rep): smoothed
KL(observed || reference) in nats. It requires an explicit reference and its
provenance; the package does not infer a population prior from a dataset name.
from fairlms.diagnostics import (
DatasetAuditSpec,
ReferenceDistribution,
RepresentationEvidence,
audit_representativeness,
)
evidence = RepresentationEvidence(
axis="community",
counts={"amber": 3, "teal": 1},
source="My benchmark, evaluation split",
)
reference = ReferenceDistribution(
axis="community",
probabilities={"amber": 0.5, "teal": 0.5},
source="Benchmark design specification v1",
purpose="design_target",
population="Intended benchmark composition",
)
spec = DatasetAuditSpec(
target_name="my-unregistered-benchmark",
target_kind="benchmark_dataset",
task_family="free_text",
design_stance="stress_test",
references={"community": reference},
)
report = audit_representativeness(evidence, spec)
result = report.components["b_rep"]
print(result.status.value, result.value)
print(report.to_json())
population_proxy and stress_test use the same mathematics but not the same
interpretation. A stress test may deliberately over-sample a category; the
report therefore warns that divergence is descriptive evidence rather than an
automatic fairness failure. Use RepresentationEvidence.from_records(...) or
.from_dataframe(...) with explicit field names, support, and optional value
mapping for unfamiliar schemas. Adapters preserve the complete JSON-safe value
mapping in provenance so a recoding can be reproduced. Reference probabilities
outside an absolute 1e-9 sum tolerance are rejected; values inside that
tolerance are canonicalized onto the probability simplex and the report records
both the input sum and whether canonicalization occurred.
Start with Preparing audit evidence for the
evidence-layer boundary and input checklist. The detailed
representativeness guide
covers supported input paths, raw text, coverage, references, and a complete
b_rep example.
Scoring instrument audits
Scorer diagnostics audit scores that already exist at row level. They do not
infer groups from raw text or run a model/scorer to generate scores.
score_mean_gap is the descriptive maximum absolute group-mean difference in
the scorer's native score units. score_rate_gap first applies one explicit,
serialized score-to-event rule to every group, then reports the maximum absolute
group event-rate difference as a proportion. score_wasserstein_1_gap compares
the complete one-dimensional empirical score distributions and reports their
maximum pairwise Wasserstein-1 distance in the scorer's native score units.
score_counterfactual_sensitivity separately averages absolute score changes
inside complete, explicitly declared two-condition pairs.
from fairlms.diagnostics import (
DatasetAuditSpec,
ScoreRateTransform,
ScoredGroups,
ScorerMeanGap,
ScorerRateGap,
ScorerWasserstein1Gap,
audit_scores,
)
evidence = ScoredGroups(
axis="cohort",
groups=("amber", "amber", "teal", "teal"),
scores=(0.1, 0.3, 0.8, 1.0),
score_name="example_safety_score",
source="Existing row-level score export v1",
score_range=(0.0, 1.0),
)
rate_transform = ScoreRateTransform(
event_name="score_at_or_above_policy_threshold",
threshold=0.5,
direction="higher",
inclusive=True,
provenance={"rule_source": "Example policy v1"},
)
spec = DatasetAuditSpec(
target_name="example-score-table",
target_kind="score_table",
task_family="scored_rows",
design_stance="stress_test",
references={},
requested_components=(
"score_mean_gap",
"score_rate_gap",
"score_wasserstein_1_gap",
),
)
report = audit_scores(
evidence,
spec,
diagnostics=(
ScorerMeanGap(),
ScorerRateGap(transform=rate_transform),
ScorerWasserstein1Gap(),
),
)
mean_result = report.components["score_mean_gap"]
rate_result = report.components["score_rate_gap"]
w1_result = report.components["score_wasserstein_1_gap"]
print(mean_result.value, mean_result.details["unit"]) # 0.7 score_units
print(rate_result.value, rate_result.details["unit"]) # 1.0 proportion
print(w1_result.value, w1_result.details["unit"]) # 0.7 score_units
Paired sensitivity uses independent evidence because group marginals do not preserve which rows are counterparts:
from fairlms.diagnostics import (
PairedScores,
ScorerCounterfactualSensitivity,
)
paired = PairedScores(
axis="declared_identity_intervention",
pair_ids=("p1", "p1", "p2", "p2"),
conditions=("baseline", "swap", "baseline", "swap"),
scores=(0.1, 0.4, 0.8, 0.3),
condition_roles=("baseline", "swap"),
score_name="example_safety_score",
source="Existing paired score export v1",
pairing_basis="Reviewed minimal identity-token substitutions",
score_range=(0.0, 1.0),
)
paired_spec = DatasetAuditSpec(
target_name="example-paired-score-table",
target_kind="score_table",
task_family="paired_sentences",
design_stance="stress_test",
references={},
requested_components=("score_counterfactual_sensitivity",),
)
paired_report = audit_scores(
paired,
paired_spec,
diagnostic=ScorerCounterfactualSensitivity(),
)
print(paired_report.components["score_counterfactual_sensitivity"].value) # 0.4
There is no implicit threshold, direction, or boundary rule. score_range
validates scores; it is neither a threshold nor a Wasserstein normalization. A
missing rate transform is blocked, not a zero gap. Mean, rate, and W1 gaps
answer different questions, and none is a causal claim, fairness pass/fail rule,
or error-rate metric. Paired sensitivity supports identity-isolated causal
language only if the pairs really differ solely in the declared intervention;
the package validates pair completeness but cannot verify that semantic claim
from scores. W1 and paired sensitivity are symmetric primary values and carry
no higher/lower direction.
higher plus inclusive=True is the paper-exact score >= threshold
definition; the report marks lower-tail and exclusive-boundary variants
separately. Treat each value as a per-dataset, per-scorer diagnostic, and rate
gaps additionally as per-rule, rather than using them to rank datasets or
unrelated scorer scales. See
Preparing audit evidence and the runnable
score_rate_gap and
score_wasserstein_1_gap
and
score_counterfactual_sensitivity
examples.
Shared loaders:
from fairlms.datasets import CrowSPairs, BBQ, StereoSet
from fairlms.models import load_masked_lm
crows = CrowSPairs().load()
bbq = BBQ(categories=["Age"]).load()
loaded = load_masked_lm("bert-base-uncased")
Leaf runners under definition/ are short demos of the same public API:
python -m fairlms.definition.encoder_only.intrinsic_bias.probability_based.pseudo_log_likelihood_metrics.cps.main
See also examples/ at the repository root.
Package layout
fairlms/
├── metrics/ # Public API: CrowSPairsScore, WEAT, … (all expose compute)
│ ├── data.py # Validated input containers (WordSets, ProbeSet, …)
│ └── functional.py # sklearn.metrics-style functions (model-free metrics)
├── diagnostics/ # Dataset/result-table evidence, applicability, and reports
├── datasets/ # CrowSPairs, StereoSet, BBQ, BiasInBios, WinoBias, …
├── models/ # HuggingFaceModel, OpenAIModel, load_* helpers
├── utils/ # PLL / masking / association / path helpers
├── data/ # Bundled CrowS-Pairs + BBQ files; exports WEAT/SEAT word sets
├── artifacts/ # Preferred output dir for metric CSVs
└── definition/ # Internal implementations + short public-API demos (main.py)
Repo-root examples/ has additional runnable snippets.
Every metric exposes the same method: compute(...).
Datasets
| Class | Source | Notes |
|---|---|---|
CrowSPairs |
Bundled CSV under fairlms/data/crows_pairs/ |
Stereotype / anti pairs |
StereoSet |
Hugging Face (stereoset / McGill-NLP/stereoset) |
Pairs or triples |
BBQ |
Bundled jsonl under fairlms/data/bbq/ |
Optional context_condition filter |
BiasInBios |
Hugging Face LabHC/bias_in_bios |
Profession / gender helpers |
WinoBias |
Hugging Face wino_bias |
Occupation direction helpers |
XNLIReligionPairs |
Hugging Face XNLI + templates | Religion swap pairs |
Loaders prefer canonical files in fairlms/data/, then fall back to legacy copies under definition/ so existing scripts keep working.
Bundled word sets
Association tests need four labelled term lists rather than a corpus, so they
are importable constants instead of loader classes. Each one is an
already-validated WordSets, ready to pass straight to compute:
| Name | Test | Terms |
|---|---|---|
weat_c1 … weat_c4 |
Caliskan et al. (2017) C1–C4 | Race, gender, disease, age |
seat_c1 … seat_c4 |
May et al. (2019), expanded name lists | Same four axes |
from fairlms.data import weat_c2, get_word_set, WORD_SET_LABELS
get_word_set("seat_c1") # same objects, by name
WORD_SET_LABELS["weat_c2"] # 'C2 – Gender (Male/Female names × Career/Family)'
Both families work with WEAT and SEAT; the seat_* lists are the larger
name sets the sentence templates were sized for. All eight have balanced target
lists, which SEAT requires. CEAT is not included — it consumes ContextSets
(terms keyed to context sentences), not four flat lists.
Models
from fairlms.models import (
HuggingFaceModel,
load_masked_lm, # task="mlm"
load_encoder, # task="encoder"
load_sequence_classifier,
load_seq2seq,
load_causal_lm,
OpenAIModel,
)
HuggingFaceModel("roberta-base", task="sequence_classification").load()
OpenAIModel("davinci-002").load() # needs OPENAI_API_KEY; pip install fairlms[openai]
Set HF_TOKEN (or HUGGING_FACE_HUB_TOKEN) for gated models such as Llama-2.
Design principles
- One verb for metrics —
compute(sklearn’sfit/predictanalogue). - Separate metrics from datasets — reuse the same metric on CrowS-Pairs, StereoSet, or custom data.
- Stable public surface — internals under
definition/can change without breaking user code. - Book-aligned taxonomy —
definition/{encoder_only,encoder_decoder,decoder_only}/{intrinsic_bias,extrinsic_bias}/…mirrors the conceptual organization of the accompanying textbook. - Applicability before computation — dataset diagnostics report missing or incompatible evidence instead of manufacturing a numeric result.
Development
pip install -e ".[dev]"
pytest # contract suite over every metric in the registry
# Run a leaf metric script:
python -m fairlms.definition.encoder_only.intrinsic_bias.similarity_based.weat.main
tests/test_common.py is the analogue of scikit-learn's check_estimator: it
runs the parameter/repr/keyword contract across METRIC_REGISTRY, so a new
metric that breaks the shape fails there rather than surprising a user. It needs
no network or model weights.
Tests that do need model weights skip cleanly when the weights aren't cached, so the suite is hermetic:
HF_HUB_OFFLINE=1 pytest -q # what CI runs; a few seconds, no downloads
CI (.github/workflows/test.yml) runs this on
Python 3.10 and 3.13 for every push and pull request. A second clean-install
job builds the wheel, installs it into a fresh environment, and imports it from a
directory with no source checkout on sys.path.
That second job exists because an editable install cannot catch a whole class of
packaging bug: fairlms.metrics eagerly imports every metric family, so any
third-party module imported at module scope under fairlms/definition/ is a
hard requirement of import fairlms. If such a dependency is only listed in an
extra, pip install fairlms produces a package that cannot be imported — while
every local test still passes, because the developer's environment already has
it. tests/test_packaging.py also guards this
statically, walking the AST and naming the offending file.
Publishing an edit
Edits reach users through a tagged release, not through main:
- Edit locally — your editable install picks changes up immediately.
pytest— the contract suite catches API breakage before it ships.- Commit and push. CI verifies the matrix and the clean install.
- Bump
fairlms/_version.py, then tag and push the tag. - The publish workflow uploads to PyPI; users get it with
pip install --upgrade fairlms.
Metric result CSVs should go under fairlms/artifacts/ (via fairlms.utils.results_to_csv); leaf-local *_results.csv files are gitignored.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file fairlms-0.3.1.tar.gz.
File metadata
- Download URL: fairlms-0.3.1.tar.gz
- Upload date:
- Size: 1.7 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
795c31f3d2ccf2fd1a6c2fc80fbb52b7b1a1f257d4609538aa0ef72f53bd3ac4
|
|
| MD5 |
e464f8415d7228a7d8fedb7bdfca1a94
|
|
| BLAKE2b-256 |
91bb18f3c53241c336fa7a99f11a490d4014af5d82b53339e8f39370292a11c9
|
Provenance
The following attestation bundles were made for fairlms-0.3.1.tar.gz:
Publisher:
publish.yml on michaellarionov/FairLMs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
fairlms-0.3.1.tar.gz -
Subject digest:
795c31f3d2ccf2fd1a6c2fc80fbb52b7b1a1f257d4609538aa0ef72f53bd3ac4 - Sigstore transparency entry: 2542474205
- Sigstore integration time:
-
Permalink:
michaellarionov/FairLMs@b199ed8683655a79f91397ff0c22202cc9ccca4b -
Branch / Tag:
refs/tags/v0.3.1 - Owner: https://github.com/michaellarionov
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b199ed8683655a79f91397ff0c22202cc9ccca4b -
Trigger Event:
push
-
Statement type:
File details
Details for the file fairlms-0.3.1-py3-none-any.whl.
File metadata
- Download URL: fairlms-0.3.1-py3-none-any.whl
- Upload date:
- Size: 1.7 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8f26459c0fd2422961b69bb06a39db448f8f93de41f14a66e13e9df31c37a3df
|
|
| MD5 |
5387e7e0850fe515ea78f3567d2ac6dc
|
|
| BLAKE2b-256 |
14df724617ccd3730781088a593e316d066a7e7c0b91845865135e7ed8181b78
|
Provenance
The following attestation bundles were made for fairlms-0.3.1-py3-none-any.whl:
Publisher:
publish.yml on michaellarionov/FairLMs
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
fairlms-0.3.1-py3-none-any.whl -
Subject digest:
8f26459c0fd2422961b69bb06a39db448f8f93de41f14a66e13e9df31c37a3df - Sigstore transparency entry: 2542474572
- Sigstore integration time:
-
Permalink:
michaellarionov/FairLMs@b199ed8683655a79f91397ff0c22202cc9ccca4b -
Branch / Tag:
refs/tags/v0.3.1 - Owner: https://github.com/michaellarionov
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b199ed8683655a79f91397ff0c22202cc9ccca4b -
Trigger Event:
push
-
Statement type: