Skip to main content

TruthScore

truthscore is a fast, modular reimplementation of RAGAS's FactualCorrectness metric, supporting both open-weight and hosted LLMs. It evaluates factual consistency between a user response and a reference passage by breaking down answers into claims and verifying them using Natural Language Inference (NLI).

It is a metric component of the TruthBench framework and is intended for scalable, cost-efficient factuality evaluation. TruthBench is the meta-evaluation framework: it applies controlled, graded perturbations to ground-truth answers so that factuality metrics can be scored on how well their judgements track the injected error severity. truthscore is an LLM+NLI factual-correctness metric that can be evaluated with it, and is shipped as its own installable package. Both are described in our EvalLLM 2025 paper.


🔍 What it does

  1. Claim Decomposition: The LLM-generated response is split into atomic factual claims using a lightweight LLM.
  2. Entailment Scoring: Each claim is passed to an NLI model with the reference passage as context.
  3. Final Score: The score reflects how many claims are entailed by the context, in the range [0.0, 1.0].

For more details, see FactualCorrectness.


✨ Key Features

  • 🔁 RAGAS-compatible: Faithfully reimplements the FactualCorrectness metric logic from RAGAS
  • ✅ Open-weight LLM support: Works with open-weight models (e.g., Gemma, LLaMA, Mistral via Ollama)
  • 🧠 Plug-and-play: Swap in custom NLI models
  • ⚙️ GPU-accelerated: Recommended for claim decomposition + NLI
  • 🧪 Evaluated: Competitive benchmark results (see TruthBench)

📦 Installation

For full open-weight support (LLM hosted with Ollama + CrossEncoders NLI):

pip install truthscore[open]

Otherwise, install the lightweight version and pick the dependencies that best suit your setup:

pip install truthscore

Regarding ollama installation, please check Ollama.

🚀 Quick Start

💡 Open-weight (fully local)

from langchain_ollama import OllamaLLM
from ragas import SingleTurnSample
from ragas.llms import LangchainLLMWrapper

from truthscore import OpenFactualCorrectness

test_data = {
    "user_input": "What happened in Q3 2024?",
    "reference": "The company saw an 8% rise in Q3 2024, driven by strong marketing and product efforts.",
    "response": "The company experienced an 8% increase in Q3 2024 due to effective marketing strategies and product efforts."
}
sample = SingleTurnSample(**test_data)

evaluator_llm = LangchainLLMWrapper(OllamaLLM(model="gemma3:27b", base_url="http://localhost:11434"))
metric = OpenFactualCorrectness(llm=evaluator_llm)
score = metric.single_turn_score(sample)

print(score)  # e.g. 1.0

☁️ Hosted LLM (e.g., OpenAI)

from openai import OpenAI
from ragas import SingleTurnSample
from ragas.llms import LangchainLLMWrapper

from truthscore import OpenFactualCorrectness

evaluator_llm = LangchainLLMWrapper(OpenAI())
metric = OpenFactualCorrectness(llm=evaluator_llm)

# test_data same as above
score = metric.single_turn_score(SingleTurnSample(**test_data))

⚙️ Custom NLI Models

import torch
from langchain_ollama import OllamaLLM
from ragas import SingleTurnSample
from ragas.llms import LangchainLLMWrapper
from sentence_transformers import CrossEncoder

from truthscore import OpenFactualCorrectness

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
nli_model = CrossEncoder("cross-encoder/nli-deberta-v3-large")
nli_model.model.to(device)

evaluator_llm = LangchainLLMWrapper(OllamaLLM(model="gemma3:27b", base_url="http://localhost:11434"))
metric = OpenFactualCorrectness(llm=evaluator_llm, nli_model=nli_model)

# test_data same as above
score = metric.single_turn_score(SingleTurnSample(**test_data))

📊 Background

This metric was evaluated across a 500-example benchmark using perturbation levels A0–A4 on top of the Google Natural Questions dataset using truthbench.

See full results in the project overview.

Citation

If you use TruthScore in your research, please cite our EvalLLM 2025 paper:

@inproceedings{gharsallah-etal-2025-peut,
    title = "Peut-on faire confiance aux juges ? Validation de m{\'e}thodes d'{\'e}valuation de la factualit{\'e} par perturbation des r{\'e}ponses",
    author = {Gharsallah, Sarra  and
      Robaldo, Ad{\`e}le  and
      Tokareva, Mariia  and
      Gatti Pinheiro, Giovanni  and
      Guendouz, Ilyana  and
      Troncy, Rapha{\"e}l  and
      Papotti, Paolo  and
      Michiardi, Pietro},
    editor = "Bechet, Fr{\'e}d{\'e}ric  and
      Chifu, Adrian-Gabriel  and
      Pinel-sauvagnat, Karen  and
      Favre, Benoit  and
      Maes, Eliot  and
      Nurbakova, Diana",
    booktitle = "Actes de l'atelier {\'E}valuation des mod{\`e}les g{\'e}n{\'e}ratifs (LLM) et challenge 2025 (EvalLLM)",
    month = "6",
    year = "2025",
    address = "Marseille, France",
    publisher = "ATALA {\&} ARIA",
    url = "https://aclanthology.org/2025.jeptalnrecital-evalllm.19/",
    pages = "228--252",
    language = "fra"
}

Metadata

Release files for truthscore 0.4.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for truthscore 0.4.2
File Size Uploaded
truthscore-0.4.2.tar.gz 6.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for truthscore 0.4.2
File Interpreter ABI Platform
truthscore-0.4.2-py3-none-any.whl Python 3 none any Details

Total release size: 13.4 kB

Release files / truthscore-0.4.2.tar.gz

Download URL truthscore-0.4.2.tar.gz
Size 6.5 kB
Tags Source
SHA-256 checksum
How to use checksums
9a2c1e6bfac2b8027f4750ea792895ea64451db796ed3e7cae1adf9951fa9bde
BLAKE2b-256 checksum
How to use checksums
ad16631d30f4ecc81963d54e62c873a2813f5d4b5e5242b1f123fbe111fa5355
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.2.1 CPython/3.11.6 Linux/5.15.167.4-microsoft-standard-WSL2

Release files / truthscore-0.4.2-py3-none-any.whl

Download URL truthscore-0.4.2-py3-none-any.whl
Size 7.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c69be58ca8a3f44df76949820739bd536c02da4a3e8cd17e26a33e8db5570680
BLAKE2b-256 checksum
How to use checksums
26d3a188f167aa4a75b464a74c734fd8d4ea372b18529bc5f0f498914b9d6ac2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.2.1 CPython/3.11.6 Linux/5.15.167.4-microsoft-standard-WSL2

Release history Release notifications | RSS feed

This release

0.4.2 This release

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page