Skip to main content

Gaussia

PyPI version PyPI - Python Version PyPI - Downloads PyPI - License

AI evaluation framework for measuring fairness, quality, and safety of AI models and assistants.

Installation

pip install gaussia

With specific metric dependencies:

pip install gaussia[toxicity]              # Toxicity analysis
pip install gaussia[bias]                  # Bias detection
pip install gaussia[privacy-presidio]      # Privacy with the Presidio detector
pip install gaussia[privacy-huggingface]   # Privacy with a HuggingFace NER detector
pip install gaussia[evalhub]               # EvalHub provider adapter
pip install gaussia[metrics]               # All metrics
pip install gaussia[all]                   # Everything

Roast Me ships its own extras, deliberately outside metrics and all, so whoever only profiles an assistant pays for neither retrieval nor training:

pip install gaussia[roastme]      # The three corpus-reading probe engines
pip install gaussia[roastme-rl]   # The above, plus the reinforcement-learning search

Most metrics judge with a LangChain-compatible chat model, which you install separately: langchain-openai, langchain-anthropic, langchain-google-genai, langchain-groq, langchain-ollama.

Quick Start

from gaussia import Retriever, Dataset, Batch
from gaussia.metrics.context import Context
from langchain_openai import ChatOpenAI

# 1. Define your data source
class MyRetriever(Retriever):
    def load_dataset(self) -> list[Dataset]:
        return [
            Dataset(
                session_id="session-1",
                assistant_id="assistant-1",
                language="en",
                context="France is a country in Western Europe.",
                conversation=[
                    Batch(
                        qa_id="q1",
                        query="Where is France?",
                        assistant="France is located in Western Europe.",
                        ground_truth_assistant="France is a country in Western Europe.",
                    )
                ],
            )
        ]

# 2. Run a metric. `run` takes the retriever *class* and instantiates it for you.
metrics = Context.run(MyRetriever, model=ChatOpenAI(model="gpt-4o-mini", temperature=0.0))

Metrics are imported from their own modules — gaussia.metrics declares the surface but imports nothing, so no metric drags in another's optional dependencies.

Metrics

Metric Description Install extra
Context Evaluates response alignment with provided context
Conversational Dialogue quality via Grice's maxims (memory, language, quality, quantity, relation, manner)
BestOf King-of-the-hill tournament comparison of multiple assistants
Agentic Agent evaluation with pass@K and tool correctness
RoleAdherence Whether the assistant stays inside its defined role, scored from judge first-token logprobs
Toxicity Cluster-based toxicity profiling with demographic and sentiment analysis [toxicity]
Bias Bias detection across protected attributes using guardians [bias]
Privacy Domain-adjusted detection score rating one PII/PHI detector's fitness for a regulated domain [privacy-presidio] / [privacy-huggingface]
PrivacyRanker The same score across several detectors, ranked, to decide which one to ship [privacy-presidio] / [privacy-huggingface]
Humanity Emotion, empathy, and human-like quality analysis [humanity]
Regulatory Compliance evaluation against regulatory documents [regulatory]
VisionSimilarity VLM description comparison via semantic similarity [vision]
VisionHallucination Hallucination detection in VLM outputs [vision]

Features

Guardians

Pluggable bias detection backends. The guardian is passed as a class, like the retriever:

from gaussia.guardians import IBMGranite, LLamaGuard
from gaussia.metrics.bias import Bias

metrics = Bias.run(MyRetriever, guardian=IBMGranite)

PII Detectors

Pluggable detection backends behind the PIIDetector contract, the same way guardians sit behind Guardian. A detector carries the expert [0, 1] scalars that place it in your domain, so it is passed as an instance:

from gaussia.detectors.presidio import PresidioDetector
from gaussia.metrics.privacy import Privacy
from gaussia.schemas.privacy import PrivacyDomainConfig

domain = PrivacyDomainConfig(
    classes=frozenset({"email_address", "phone_number"}),
    criticality_weights={"email_address": 0.6, "phone_number": 0.4},
    fn_severity_weights={"email_address": 0.7, "phone_number": 0.3},
    regulatory_framework="GDPR",
)
detector = PresidioDetector(name="presidio", domain_fit=0.9, regulatory_fit=0.8)

metrics = Privacy.run(MyRetriever, detector=detector, domain_config=domain)

PrivacyRanker takes detectors=[...] instead and ranks them under the same domain config.

Role Adherence

Whether the assistant stayed in the role it was given, scored per turn and aggregated per session. The judge derives a calibrated [0, 1] score from first-token logprobs, so it needs a provider that exposes them — StructuredOutputJudgeStrategy is the fallback for providers that do not:

from gaussia.metrics.role_adherence import RoleAdherence, LLMJudgeStrategy
from langchain_openai import ChatOpenAI

strategy = LLMJudgeStrategy(model=ChatOpenAI(model="gpt-4o-mini"))
metrics = RoleAdherence.run(MyRetriever, scoring_strategy=strategy)

Statistical Modes

Choose between frequentist and Bayesian aggregation:

from gaussia import FrequentistMode, BayesianMode
from gaussia.metrics.context import Context

metrics = Context.run(MyRetriever, model=judge, statistical_mode=FrequentistMode())
metrics = Context.run(MyRetriever, model=judge, statistical_mode=BayesianMode())

Synthetic Data Generation

Generate evaluation datasets from documents:

from gaussia.generators import BaseGenerator, create_markdown_loader
from langchain_openai import ChatOpenAI

generator = BaseGenerator(model=ChatOpenAI(model="gpt-4o-mini"))
loader = create_markdown_loader()

datasets = await generator.generate_dataset(
    context_loader=loader,
    source="./docs/knowledge_base.md",
    assistant_id="my-assistant",
)

Roast Me

Adversarial evaluation as a search problem: profile an assistant's weaknesses from tagged probes, then look for the categories of realistic question that break it reproducibly. A generator subsystem, not a metric — nothing subclasses Gaussia. What enters the metric pipeline is the Roast Dataset it emits, which existing metrics consume unchanged.

from gaussia.generators.roastme import ProbeLibrary, Profiler, to_dataset

probes = ProbeLibrary(engines).generate(documents, catalogue)
result = Profiler(contract=contract, target=your_adapter).profile(probes)

dataset = to_dataset(
    probes,
    result.outcomes,
    session_id="roast-run-1",
    assistant_id="support-assistant",
    context="Roast Me run over the policy knowledge base",
)

Every number Roast Me produces is a judge-only measurement: one language model's estimate of whether another one misbehaved. No grader here has been calibrated against human labels, so read a violation rate as evidence to look at, never as a measured error rate.

Explainability

Token-level attribution analysis. The method is a class, not a string:

from transformers import AutoModelForCausalLM, AutoTokenizer
from gaussia.explainability import AttributionExplainer, Lime

model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-0.6B")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-0.6B")

explainer = AttributionExplainer(model, tokenizer)
result = explainer.explain(
    prompt=tokenizer.apply_chat_template([{"role": "user", "content": "What is gravity?"}], tokenize=False),
    target="Gravity is the force of attraction between objects.",
    method=Lime,
)
print(result.get_top_k(5))

Prompt Optimization

Optimize prompts using evolutionary and multi-objective strategies:

from gaussia.prompt_optimizer import GEPAOptimizer, MIPROv2Optimizer

EvalHub Provider

Run Gaussia as an EvalHub BYOF provider:

python -m gaussia.integrations.evalhub.adapter

Documentation

Full documentation available at docs.gaussia.ai.

Requirements

  • Python >= 3.11

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

gaussia-1.2.0.tar.gz (899.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

gaussia-1.2.0-py3-none-any.whl (966.9 kB view details)

Uploaded Python 3

File details

Details for the file gaussia-1.2.0.tar.gz.

File metadata

  • Download URL: gaussia-1.2.0.tar.gz
  • Upload date:
  • Size: 899.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for gaussia-1.2.0.tar.gz
Algorithm Hash digest
SHA256 4b0c2b68362a96bf59d16c0da6801a833c2680e6902275ecfc7e1148273b5956
MD5 758a3dc2ad73ab3f5b9ce292c788c136
BLAKE2b-256 2a93ec90ce612d7bd93afcbfc3de3dabfc8217d38b0a24783577267b02c2937c

See more details on using hashes here.

File details

Details for the file gaussia-1.2.0-py3-none-any.whl.

File metadata

  • Download URL: gaussia-1.2.0-py3-none-any.whl
  • Upload date:
  • Size: 966.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for gaussia-1.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 f8b3236d7832ac47202ff711f99187fc6afdf6f2c47085b4686b30cff63cee9f
MD5 edf3fba6571eb61e9338f4ca4183e473
BLAKE2b-256 76c3d77239f93d9a86813da4510f2a6c27d4f197d4a4e9e9dd9b2805bcf6f0ae

See more details on using hashes here.

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page