Skip to main content

Measuring Epistemic Humility in Multimodal Large Language Models

License PyPI HuggingFace

📦 Installation

Install the latest release from PyPI:

pip install HumbleBench

🚀 Quickstart (Python API)

The following snippet demonstrates a minimal example to evaluate your model on HumbleBench.

from HumbleBench import download_dataset, evaluate
from HumbleBench.utils.entity import DataLoader

# Download the HumbleBench dataset
dataset = download_dataset()

# Prepare data loader (batch_size=16, no-noise images)
data = DataLoader(dataset=dataset,
                    batch_size=16, 
                    use_noise_image=False,  # For HumbleBench-GN, set this to True
                    nota_only=False)        # For HumbleBench-E, set this to True

# Run inference
results = []
for batch in data:
    # Replace the next line with your model's inference method
    predictions = your_model.infer(batch)
    # Expect predictions to be a list of dicts matching batch keys, plus 'prediction'
    # Example: 
    results.extend(predictions)

# Compute evaluation metrics
metrics = evaluate(
    input_data=results,
    model_name_or_path='YourModel',
    use_noise_image=False,  # For HumbleBench-GN, set this to True
    nota_only=False         # For HumbleBench-E, set this to True
)
print(metrics)

If you prefer to reproduce the published results, load one of our provided JSONL files (at results/common, results/noise_image, or results/nota_only):

from HumbleBench.utils.io import load_jsonl
from HumbleBench import evaluate

path = 'results/common/Model_Name/Model_Name.jsonl'
data = load_jsonl(path)
metrics = evaluate(
    input_data=data,
    model_name_or_path='Model_Name',
    use_noise_image=False,  # For HumbleBench-GN, set this to True
    nota_only=False,        # For HumbleBench-E, set this to True
)
print(metrics)

🧩 Advanced Usage: Command-Line Interface

⚠️WARNING⚠️: If you wanna use our implemented models, please make sure you install all the requirements of respective model by yourself. And we use Conda to manage the python environment, so maybe you need to modify the env_name to your env's name.

HumbleBench provides a unified CLI for seamless integration with any implementation of our model interface.

1. Clone the Repository

git clone git@github.com:maifoundations/HumbleBench.git
cd HumbleBench

2. Implement the Model Interface

Create a subclass of MultiModalModelInterface and define the infer method:

# my_model.py
from HumbleBench.models.base import register_model, MultiModalModelInterface

@register_model("YourModel")
class YourModel(MultiModalModelInterface):
    def __init__(self, model_name_or_path, **kwargs):
        super().__init__(model_name_or_path, **kwargs)
        # Load your model and processor here
        # Example:
        # self.model = ...
        # self.processor = ...

    def infer(self, batch: List[Dict]) -> List[Dict]:
        """
        Args:
            batch: List of dicts with keys:
                - label: one of 'A', 'B', 'C', 'D', 'E'
                - question: str
                - type: 'Object'/'Attribute'/'Relation'/...
                - file_name: path to image file
                - question_id: unique identifier
        Returns:
            List of dicts with an added 'prediction' key (str).
        """
        # Your inference code here
        return predictions

3. Configure Your Model

Edit configs/models.yaml to register your model and specify its weights:

models:
  YourModel:
    params:
      model_name_or_path: "/path/to/your/checkpoint"

4. Run Evaluation from the Shell

#!/bin/bash
export CUDA_VISIBLE_DEVICES=0,1,2,3

python main.py \
    --model "YourModel" \
    --config configs/models.yaml \
    --batch_size 16 \
    --log_dir results/common \
    [--use-noise] \
    [--nota-only]
  • --model: Name registered via @register_model
  • --config: Path to your models.yaml
  • --batch_size: Inference batch size
  • --log_dir: Directory to save logs and results
  • --use-noise: Optional flag to assess HumbleBench-GN
  • --nota-only: Optional flag to assess HumbleBench-E

5. Contribute to HumbleBench!

🙇🏾🙇🏾🙇🏾

We have implemented many popular models in the models directory, along with corresponding shell scripts (including support for noise-image experiments) in the shell directory. If you’d like to add your own model to HumbleBench, feel free to open a Pull Request — we’ll review and merge it as soon as possible.

📮 Contact

For bug reports or feature requests, please open an issue or email us at bingkuitong@gmail.com.

Release files for HumbleBench 1.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for HumbleBench 1.0.2
File Size Uploaded
humblebench-1.0.2.tar.gz 10.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for HumbleBench 1.0.2
File Interpreter ABI Platform
humblebench-1.0.2-py3-none-any.whl Python 3 none any Details

Total release size: 23.4 kB

Release files / humblebench-1.0.2.tar.gz

Download URL humblebench-1.0.2.tar.gz
Size 10.5 kB
Tags Source
SHA-256 checksum
How to use checksums
5a04e5c5cfd18b8759449223a03801a32f26617b34c1dbed9812a4de09d03e6a
BLAKE2b-256 checksum
How to use checksums
c23e56bf0c7841e1300afec1d261f93bbc6fe0fa703b4adb631b7aafbd91e305
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.14

Release files / humblebench-1.0.2-py3-none-any.whl

Download URL humblebench-1.0.2-py3-none-any.whl
Size 12.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
142b7bbaa55015008b86bb13240daef3045039e1657a37aa49036b7057f2d4a4
BLAKE2b-256 checksum
How to use checksums
995a7efcc03509091b14c7e0db376acbf8859f0ff1ed6814def4e30b4b340032
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.10.14

Release history Release notifications | RSS feed

This release

1.0.2 This release

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page