Skip to main content

A library for evaluating LLM-based applications.

Project description

GLLM Evaluator SDK

A comprehensive evaluation framework for Generative AI applications including LLM outputs, AI Agent responses, and RAG (Retrieval-Augmented Generation) systems.

Overview

The GLLM Evaluator SDK provides a robust, extensible framework designed to make AI evaluation as simple and seamless as possible across the GDP Labs ecosystem. Built with integration-first philosophy, it enables teams to easily assess the quality of generated content from any AI system while seamlessly connecting with experiment tracking and observability platforms.

Philosophy

Easy Evaluation Everywhere: Standardize evaluation practices across all GDP Labs AI applications with minimal setup and maximum flexibility.

Integration-First Design: Built to work seamlessly with your existing experiment tracking, observability, and MLOps infrastructure.

Extensible by Design: Add new evaluators, metrics, and integrations without breaking existing workflows.

Key Features

  • 🌐 GDP Labs Ecosystem Ready: Standardized evaluation framework across all internal AI applications
  • 🔌 Seamless Integration: Easy integration with experiment tracking and observability platforms
  • 🚀 Async-First Design: High-performance async evaluation with parallel processing
  • 🔧 Extensible Architecture: Easy to add new evaluators and metrics for any use case
  • 🤖 LLM as a Judge: Advanced language models for nuanced, contextual evaluation
  • 📐 Traditional Metrics: Support for classical evaluation metrics and custom scoring functions
  • 🔗 Popular Evaluator Integration: Integration with popular evaluators such as RAGAS, DeepEval, and LangChain
  • Zero-Config Start: Get started with sensible defaults, customize as needed

Installation

Prerequisites

Mandatory:

  1. Python 3.11+ — Install here
  2. pip — Install here
  3. uv — Install here
  4. gcloud CLI (for authentication) — Install here, then log in using:
    gcloud auth login
    

Install from Artifact

Because gllm-evals is a private library hosted in a secure Google Cloud repository, you must provide an access token to install it. The command below handles this authorization inline by using an access token from the gcloud CLI.

uv pip install --extra-index-url https://oauth2accesstoken:$(gcloud auth print-access-token)@glsdk.gdplabs.id/gen-ai-internal/simple/ gllm-evals

Local Development Setup

Prerequisites

  1. Python 3.11+ — Install here

  2. pip — Install here

  3. uv — Install here

  4. gcloud CLI — Install here, then log in using:

    gcloud auth login
    
  5. Git — Install here

  6. Access to the GDP Labs SDK GitHub repository


1. Clone Repository

git clone git@github.com:GDP-ADMIN/gl-sdk.git
cd gl-sdk/libs/gllm-evals

2. Setup Authentication

Because gllm-evals is a private library, you first need to configure uv to authenticate with our secure Google Cloud repositories. Set the following environment variables to authenticate with internal package indexes:

export UV_INDEX_GEN_AI_INTERNAL_USERNAME=oauth2accesstoken
export UV_INDEX_GEN_AI_INTERNAL_PASSWORD="$(gcloud auth print-access-token)"

3. Quick Setup

Run:

make setup

4. Activate Virtual Environment

source .venv/bin/activate

Local Development Utilities

The following Makefile commands are available for quick operations:

Install uv

make install-uv

Install Pre-Commit

make install-pre-commit

Install Dependencies

make install

Update Dependencies

make update

Run Tests

make test

Adding the Package

Once authorization is configured, you can add gllm-evals to your project:

uv add gllm-evals

Dependencies

The SDK requires:

  • gllm-core and gllm-inference for LLM interactions
  • pydantic for data validation

Quick Start

Basic Usage

import asyncio
import os
from gllm_evals.evaluator.geval_generation_evaluator import GEvalGenerationEvaluator

async def main():
    # Initialize the evaluator
    evaluator = GEvalGenerationEvaluator(
        model_credentials=os.getenv("GOOGLE_API_KEY")
    )

    # Prepare evaluation data
    data = {
        "query": "What is the capital of France?",
        "expected_response": "Paris is the capital of France.",
        "generated_response": "The capital of France is Paris.",
        "retrieved_context": "Paris is the capital and largest city of France."
    }

    # Evaluate
    result = await evaluator.evaluate(data)
    print(result)

if __name__ == "__main__":
    asyncio.run(main())

Multimodal Evaluation

Use [ATTACHMENT:<id-or-uri>] placeholders in evaluation fields when the judge needs image context. Attachments may be declared by ID through LLMTestCase.attachments, or referenced inline with https://, http://, file://, or data: URIs.

import asyncio

from gllm_evals import AttachmentRef, LLMTestCase, evaluate
from gllm_evals.evaluator.geval_generation_evaluator import GEvalGenerationEvaluator


async def main():
    data = [
        LLMTestCase(
            input="Describe this image: [ATTACHMENT:sample_image]",
            actual_output="The image shows a mountain landscape.",
            expected_output="A mountainous outdoor landscape is visible.",
            retrieved_context=["Reference image: [ATTACHMENT:sample_image]"],
            attachments={
                "sample_image": AttachmentRef(
                    uri="https://picsum.photos/id/29/640/480.jpg",
                    mime_type="image/jpeg",
                )
            },
        )
    ]

    results = await evaluate(data=data, evaluators=[GEvalGenerationEvaluator()])
    print(results)


if __name__ == "__main__":
    asyncio.run(main())

You can place image placeholders in input, actual_output, expected_output, retrieved_context, and expected_context. The evaluator resolves placeholders before invoking the judge; unresolved placeholders fail validation instead of being sent as raw text.

For CSV and spreadsheet datasets, put row-level attachments in an attachments JSON cell:

input,actual_output,expected_output,attachments
"Describe [ATTACHMENT:image_1]","A mountain landscape.","A mountain landscape.","{""image_1"":{""uri"":""https://picsum.photos/id/29/640/480.jpg"",""mime_type"":""image/jpeg""}}"

See examples/evaluate/example_multimodal_evaluate_from_csv.py for a runnable CSV example that references bundled local image files, including a chart image used for a chart-comprehension question.

Inline URI and data URL placeholders are also supported without a named attachment entry:

LLMTestCase(
    input="Describe [ATTACHMENT:https://example.com/image.jpg]",
    actual_output="A product photo.",
)

Pass multimodal=False to suppress multimodal judging rules on the metrics that expose the option — LMBasedMetric and the GEval-family metrics (DeepEvalGEvalMetric and its subclasses, e.g. groundedness, completeness, and redundancy). Placeholder resolution still occurs, so raw [ATTACHMENT:...] text is not sent to the judge.

Batch Evaluation

import asyncio
import os
from gllm_evals.dataset.dict_dataset import DictDataset
from gllm_evals.evaluator.geval_generation_evaluator import GEvalGenerationEvaluator
from gllm_evals.runner import Runner
from gllm_evals.experiment_tracker.csv_experiment_tracker import CSVExperimentTracker

async def batch_evaluation():
    # Initialize evaluator
    evaluator = GEvalGenerationEvaluator(
        model_credentials=os.getenv("GOOGLE_API_KEY"),
        run_parallel=True  # Enable parallel processing
    )

    # Create dataset
    dataset = DictDataset([
        {
            "query": "What is the capital of France?",
            "expected_response": "Paris",
            "generated_response": "Paris is the capital of France.",
            "retrieved_context": "Paris is the capital of France."
        },
        {
            "query": "What is 1 + 1?",
            "expected_response": "2",
            "generated_response": "The answer is 2.",
            "retrieved_context": "1 + 1 equals 2."
        }
    ])

    # Run evaluation
    runner = Runner(evaluator, batch_size=10)
    results = await runner.evaluate(dataset)

    # Track results
    tracker = CSVExperimentTracker(score_key="generation/score")
    tracker.log_batch(results)

    print(f"Evaluation Results: {tracker.get_results()}")

if __name__ == "__main__":
    asyncio.run(batch_evaluation())

Custom Metrics

Create domain-specific metrics easily:

from gllm_evals.metrics.metric import BaseMetric
from gllm_evals.types import LLMTestCase, MetricOutput

class DomainSpecificMetric(BaseMetric):
    """Custom metric for domain-specific evaluation."""

    name = "domain_accuracy"

    async def _evaluate(self, data: LLMTestCase) -> MetricOutput:
        # Your domain-specific evaluation logic
        score = self.calculate_domain_score(data)
        return {"score": score, "explanation": "Domain-specific reasoning"}

Architecture

Core Components

1. Evaluators

  • BaseEvaluator: Abstract base class for all evaluators - extend for any evaluation scenario
  • GEvalGenerationEvaluator: Production-ready GEval-backed evaluator for text generation quality with rule-based scoring

2. Metrics

  • BaseMetric: Abstract base class for metrics - create custom metrics for any domain
  • LMBasedMetric: Generic LM-powered metric evaluation with customizable prompts

3. Datasets

  • BaseDataset: Abstract base class for datasets - support any data format
  • DictDataset: Simple dictionary-based dataset implementation

4. Runner

  • Runner: Runner class for batch evaluation

Metrics

Below is a list of metrics that are currently supported by the SDK.

Metric Description Type Score Range
LMBasedMetric An all purpose metric that can be used to evaluate any metric that can be expressed as a LM prompt LM-based -
DeepEvalGEvalMetric A versatile evaluation metric framework that can be used to create custom evaluation metrics with configurable criteria, evaluation steps, and rubrics LM-based -
GEvalCompletenessMetric A metric that can be used to evaluate the completeness of the generated output DeepEval GEval 1-3
GEvalRedundancyMetric A metric that can be used to evaluate the redundancy of the generated output DeepEval GEval 1-3
GEvalGroundednessMetric A metric that can be used to evaluate the groundedness of the generated output DeepEval GEval 1-3
GEvalLanguageConsistencyMetric A metric that can be used to evaluate language consistency between query and generated response DeepEval GEval 0-1
GEvalRefusalMetric A metric that can be used to evaluate refusal behavior from query and expected response DeepEval GEval 0-1
GEvalRefusalAlignmentMetric A metric that can be used to evaluate refusal alignment between expected and generated responses DeepEval GEval 0-1

Evaluators

Below is a list of evaluators that are currently supported by the SDK.

Evaluator Description Type
GEvalGenerationEvaluator An evaluator that can be used to evaluate the quality of the generated output LLM-based

Datasets

Below is a list of datasets that are currently supported by the SDK.

Dataset Description
DictDataset A dataset that loads data from a dictionary
HuggingFaceDataset A dataset that loads data from a HuggingFace dataset

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

gllm_evals_binary-0.1.23-cp313-cp313-win_amd64.whl (2.5 MB view details)

Uploaded CPython 3.13Windows x86-64

gllm_evals_binary-0.1.23-cp313-cp313-manylinux_2_31_x86_64.whl (3.6 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.31+ x86-64

gllm_evals_binary-0.1.23-cp313-cp313-macosx_13_0_arm64.whl (2.9 MB view details)

Uploaded CPython 3.13macOS 13.0+ ARM64

gllm_evals_binary-0.1.23-cp312-cp312-win_amd64.whl (2.5 MB view details)

Uploaded CPython 3.12Windows x86-64

gllm_evals_binary-0.1.23-cp312-cp312-manylinux_2_31_x86_64.whl (3.6 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.31+ x86-64

gllm_evals_binary-0.1.23-cp312-cp312-macosx_13_0_arm64.whl (2.8 MB view details)

Uploaded CPython 3.12macOS 13.0+ ARM64

gllm_evals_binary-0.1.23-cp311-cp311-win_amd64.whl (2.5 MB view details)

Uploaded CPython 3.11Windows x86-64

gllm_evals_binary-0.1.23-cp311-cp311-manylinux_2_31_x86_64.whl (3.3 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.31+ x86-64

gllm_evals_binary-0.1.23-cp311-cp311-macosx_13_0_arm64.whl (2.8 MB view details)

Uploaded CPython 3.11macOS 13.0+ ARM64

File details

Details for the file gllm_evals_binary-0.1.23-cp313-cp313-win_amd64.whl.

File metadata

File hashes

Hashes for gllm_evals_binary-0.1.23-cp313-cp313-win_amd64.whl
Algorithm Hash digest
SHA256 01b2b3927c384c2648aa75e3b9da2e43eab3399a595d9ca958d0450e1931f108
MD5 ab716c8505c46521ab0ac5372e55884d
BLAKE2b-256 0bd432d019626537a7f3b97cc0ddf16aa04c507d00a573f4ce1bc0eabcf91a9b

See more details on using hashes here.

Provenance

The following attestation bundles were made for gllm_evals_binary-0.1.23-cp313-cp313-win_amd64.whl:

Publisher: build-binary.yml on GDP-ADMIN/gl-sdk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file gllm_evals_binary-0.1.23-cp313-cp313-manylinux_2_31_x86_64.whl.

File metadata

File hashes

Hashes for gllm_evals_binary-0.1.23-cp313-cp313-manylinux_2_31_x86_64.whl
Algorithm Hash digest
SHA256 640c3ed2e2e695432d3138fb78bbac1c14c23bb05bf445bc103c65970a8d0d4a
MD5 ae98e77b254b8ab9d5e9b6824269b9ce
BLAKE2b-256 ffb33ee85bcdcf1c6167c01678a1b6ef4a7d641632803015d83eee44de60f9ef

See more details on using hashes here.

File details

Details for the file gllm_evals_binary-0.1.23-cp313-cp313-macosx_13_0_arm64.whl.

File metadata

File hashes

Hashes for gllm_evals_binary-0.1.23-cp313-cp313-macosx_13_0_arm64.whl
Algorithm Hash digest
SHA256 b1c9b4a9e42dd7cc5eeb115f8fc396848f3b44db3c0f8b9b73dcb6f91ab79b36
MD5 c207ed02f308eea76c336f4093147ae0
BLAKE2b-256 c2dae1ef70b22e987c05632a8d13b66a0f5a28b9282c1e24c182cd2e474ad124

See more details on using hashes here.

Provenance

The following attestation bundles were made for gllm_evals_binary-0.1.23-cp313-cp313-macosx_13_0_arm64.whl:

Publisher: build-binary.yml on GDP-ADMIN/gl-sdk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file gllm_evals_binary-0.1.23-cp312-cp312-win_amd64.whl.

File metadata

File hashes

Hashes for gllm_evals_binary-0.1.23-cp312-cp312-win_amd64.whl
Algorithm Hash digest
SHA256 eb7788278e96d67e29f75cc43ac29439c265fbfaf4ede3ddb798d63cb316907f
MD5 11e4673b28ea8ad199509bf31c91dd7f
BLAKE2b-256 d6d61016583c5dafa72940487323e863695d1be36938b8b2314a59f1c8e57bb1

See more details on using hashes here.

Provenance

The following attestation bundles were made for gllm_evals_binary-0.1.23-cp312-cp312-win_amd64.whl:

Publisher: build-binary.yml on GDP-ADMIN/gl-sdk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file gllm_evals_binary-0.1.23-cp312-cp312-manylinux_2_31_x86_64.whl.

File metadata

File hashes

Hashes for gllm_evals_binary-0.1.23-cp312-cp312-manylinux_2_31_x86_64.whl
Algorithm Hash digest
SHA256 d7bbaf23403c036acd796ce526854768a0ca3bc81ea238495eeaec7f569ea629
MD5 5d7ed25a5b2a11ff97066e476fbb84ca
BLAKE2b-256 bc0cbd2bc431f88e384b2b054f892942959d1cec9acde758b83750e7e3a68869

See more details on using hashes here.

File details

Details for the file gllm_evals_binary-0.1.23-cp312-cp312-macosx_13_0_arm64.whl.

File metadata

File hashes

Hashes for gllm_evals_binary-0.1.23-cp312-cp312-macosx_13_0_arm64.whl
Algorithm Hash digest
SHA256 8035fff9c4d996250ab05f4ff28a1337c7fa9b0a7c52048cd725e8c3c61ca7a7
MD5 6ede7ece38bc13084d2e9844683b3e66
BLAKE2b-256 287f38da1f26a22921a1add67bd7eac38eb4617dcabe60fac7bf9a96e862cdad

See more details on using hashes here.

Provenance

The following attestation bundles were made for gllm_evals_binary-0.1.23-cp312-cp312-macosx_13_0_arm64.whl:

Publisher: build-binary.yml on GDP-ADMIN/gl-sdk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file gllm_evals_binary-0.1.23-cp311-cp311-win_amd64.whl.

File metadata

File hashes

Hashes for gllm_evals_binary-0.1.23-cp311-cp311-win_amd64.whl
Algorithm Hash digest
SHA256 4e9f20451eb7a91cd313ff0180af75189c80fd2c78435578f46c03a365285936
MD5 de1b8c10715b2f6dcf852c48dc1c7ed6
BLAKE2b-256 2375719e3c7ee22852062bf38ba077bfe6417496fe27ea008c52c486012e2940

See more details on using hashes here.

Provenance

The following attestation bundles were made for gllm_evals_binary-0.1.23-cp311-cp311-win_amd64.whl:

Publisher: build-binary.yml on GDP-ADMIN/gl-sdk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file gllm_evals_binary-0.1.23-cp311-cp311-manylinux_2_31_x86_64.whl.

File metadata

File hashes

Hashes for gllm_evals_binary-0.1.23-cp311-cp311-manylinux_2_31_x86_64.whl
Algorithm Hash digest
SHA256 3b379c9f7c2ba38b1d9d69a27ec0457495235a0dee8e507c23dacf9e281b0f67
MD5 8166ae9554db076b3c86ed25fee01b15
BLAKE2b-256 ff6fe9122ce724acf4678e14cf4c9a583d70300ff57727f3d6c099e1e0012f06

See more details on using hashes here.

File details

Details for the file gllm_evals_binary-0.1.23-cp311-cp311-macosx_13_0_arm64.whl.

File metadata

File hashes

Hashes for gllm_evals_binary-0.1.23-cp311-cp311-macosx_13_0_arm64.whl
Algorithm Hash digest
SHA256 738cdafacd29e954f94a94c3bba259ebd683c3af8638dc6cd0495be585b99c30
MD5 ee57c5090f338c4cf85fa8364bb0b1d4
BLAKE2b-256 098ac63ee9c68915db78791ca98d2d8614353c9ab839ddddf8ad9539ecd2abde

See more details on using hashes here.

Provenance

The following attestation bundles were made for gllm_evals_binary-0.1.23-cp311-cp311-macosx_13_0_arm64.whl:

Publisher: build-binary.yml on GDP-ADMIN/gl-sdk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page