Skip to main content

Dropwise-Metrics

Dropwise-Metrics is a lightweight TorchMetrics-compatible toolkit for performing Monte Carlo Dropout–based uncertainty estimation in Transformers. It enables confidence-aware decision making by revealing how certain a model is about its predictions — packaged as plug-and-play PyTorch Metric classes.


Features

  • Enable dropout during inference for Bayesian-like uncertainty estimation
  • Compute predictive entropy, confidence, and per-class standard deviation
  • Modular support for classification, QA, token tagging, and regression
  • Works seamlessly with Hugging Face Transformers and PyTorch
  • TorchMetrics-compatible: .update() + .compute()
  • Supports batch inference, CPU/GPU, and customizable num_passes
  • Cleanly packaged and extensible for research or production

Supported Tasks

  • sequence-classification — e.g. distilbert-base-uncased-finetuned-sst-2-english
  • token-classification — e.g. dslim/bert-base-NER
  • question-answering — e.g. deepset/bert-base-cased-squad2
  • regression — e.g. roberta-base with a custom head

Note: Your model must contain dropout layers for MC sampling to work (most Hugging Face models do).


Installation

pip install dropwise-metrics

Or install from source:

git clone https://github.com/aryanator/dropwise-metrics.git
cd dropwise-metrics
pip install -e .

Example Usage (Metric Style)

Sequence Classification

from transformers import AutoModelForSequenceClassification, AutoTokenizer
from dropwise_metrics.metrics.entropy import PredictiveEntropyMetric

model = AutoModelForSequenceClassification.from_pretrained("distilbert-base-uncased-finetuned-sst-2-english")
tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased-finetuned-sst-2-english")

metric = PredictiveEntropyMetric(model, tokenizer, task_type="sequence-classification", num_passes=20)
metric.update(["The movie was fantastic!", "Awful experience."])
results = metric.compute()

print(results[0])

Token Classification (NER)

from transformers import AutoModelForTokenClassification, AutoTokenizer
from dropwise_metrics.metrics.entropy import PredictiveEntropyMetric

model = AutoModelForTokenClassification.from_pretrained("dslim/bert-base-NER")
tokenizer = AutoTokenizer.from_pretrained("dslim/bert-base-NER")

metric = PredictiveEntropyMetric(model, tokenizer, task_type="token-classification", num_passes=15)
metric.update(["Hugging Face is based in New York City."])
results = metric.compute()

print(results[0]['token_predictions'])

Question Answering

from transformers import AutoModelForQuestionAnswering, AutoTokenizer
from dropwise_metrics.metrics.entropy import PredictiveEntropyMetric

model = AutoModelForQuestionAnswering.from_pretrained("deepset/bert-base-cased-squad2")
tokenizer = AutoTokenizer.from_pretrained("deepset/bert-base-cased-squad2")

question = "Where is Hugging Face based?"
context = "Hugging Face Inc. is a company based in New York City."
qa_input = f"{question} [SEP] {context}"

metric = PredictiveEntropyMetric(model, tokenizer, task_type="question-answering", num_passes=10)
metric.update([qa_input])
results = metric.compute()

print(results[0]['answer'])

Regression

from transformers import AutoModelForSequenceClassification, AutoTokenizer
from dropwise_metrics.metrics.entropy import PredictiveEntropyMetric

model = AutoModelForSequenceClassification.from_pretrained("roberta-base", num_labels=1)
tokenizer = AutoTokenizer.from_pretrained("roberta-base")

metric = PredictiveEntropyMetric(model, tokenizer, task_type="regression", num_passes=20)
metric.update(["The child is very young."])
results = metric.compute()

print(results[0]['predicted_score'], "+/-", results[0]['uncertainty'])

Output Dictionary (per sample)

Common fields returned:

  • predicted_class: Most probable class (classification)
  • predicted_score: Scalar prediction (regression)
  • confidence: Highest softmax probability
  • entropy: Predictive entropy (lower = more confident)
  • std_dev: Per-class standard deviation
  • probs: Raw softmax probabilities
  • margin: Confidence gap between top-2 classes
  • answer: Predicted span (question answering only)
  • token_predictions: Per-token predictions (NER only)

Run Tests

python test_entropy.py

Folder Structure

dropwise_metrics/
├── base.py
├── metrics/
│   └── entropy.py
├── tasks/
│   ├── __init__.py
│   ├── sequence_classification.py
│   ├── token_classification.py
│   ├── question_answering.py
│   └── regression.py

License

MIT License

Built with ❤️ for robust, explainable, uncertainty-aware AI systems.

Metadata

Release files for dropwise-metrics 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dropwise-metrics 0.1.1
File Size Uploaded
dropwise_metrics-0.1.1.tar.gz 7.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dropwise-metrics 0.1.1
File Interpreter ABI Platform
dropwise_metrics-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 14.8 kB

Release files / dropwise_metrics-0.1.1.tar.gz

Download URL dropwise_metrics-0.1.1.tar.gz
Size 7.2 kB
Tags Source
SHA-256 checksum
How to use checksums
abceabb26d56199ac37f2b2baae27325f5410bd4d34e0ec20bbde2c88f9bfe1e
BLAKE2b-256 checksum
How to use checksums
3fbecded66513bda26fa8c31f5ac469680d0c93e61409e46155af7cf1b0ae120
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.12.4

Release files / dropwise_metrics-0.1.1-py3-none-any.whl

Download URL dropwise_metrics-0.1.1-py3-none-any.whl
Size 7.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b4862ff10ba15bbac39984b80d5ec5fafca9f213902527be138461e1908bd54a
BLAKE2b-256 checksum
How to use checksums
09dcbfd300f96017214cfea22aa5ce601c46298b1724a9f6c78c84dbea32f255
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.12.4

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page