Skip to main content

Semantic Testing for LLMs - Test your AI outputs with semantic similarity validation

Project description

🧪 PromptEval

Semantic Testing for LLMs - Test your AI outputs with semantic similarity validation.

Version Python License

🚀 Quick Start

Installation

pip install prompteval-core

CLI Usage

# Run tests
prompteval run adapter.yml --tests tests.yml --api-key $PROMPTEVAL_API_KEY

# Validate YAML files
prompteval validate tests.yml

# Generate HTML report
prompteval report results.json --output report.html

# Check license/quota
prompteval licenses --api-key $PROMPTEVAL_API_KEY

Python SDK

from prompteval import PromptEval

# Initialize client
client = PromptEval(api_key="pe_xxxxx")

# Run tests from files
result = client.run_from_files(
    adapter_path="adapter.yml",
    tests_path="tests.yml"
)

# Check results
print(f"Success rate: {result.success_rate}%")
print(f"Passed: {result.total_passed}/{result.total_tests}")

if not result.success:
    for test in result.failed_tests:
        print(f"❌ {test.id_test}: {test.similarity:.1%} similarity")

📋 Configuration Files

Adapter YAML

Define your LLM endpoint configuration:

name: fitness-llm-api
description: Fitness LLM Testing

# Endpoint
endpoint:
  url: https://fitness-llm.com/v1/chat
  method: POST
  timeout: 10

# Request template
request:
  headers:
    Content-Type: application/json
  
  template:
    prompt: "{{PROMPT}}"
    type: "{{TYPE}}"
    max_tokens: 150

# Response extraction
response:
  type: json
  path: choices.0.message.content

# Validation
validation:
  ml_threshold: 0.75
  use_semantic: true

execution:
  parallel_limit: 10
  batch_delay: 0.3
  output_dir: ./output

Test Cases YAML

Define your test cases:

tests:
  - name: duration_basic
    id: FIT-001
    description: Pregunta sobre tiempo de entrenamiento
    prompt: "Ask about training duration"
    context:
      PROMPT: "Ask about training duration"
      TYPE: "duration"
    expected: "Desde cuando estas entrenando este ejercicio o rutina"
    variants:
      - "Hace cuanto tiempo empezaste con este entrenamiento"
      - "Cuanto tiempo llevas entrenando este ejercicio"
      - "Por favor, dime cuándo empezaste con este entrenamiento."
      - "Cuéntame, ¿desde cuándo entrenas así?"
    threshold: 0.70
    tags:
      - duration

🔧 SDK Reference

PromptEval Client

from prompteval import PromptEval

client = PromptEval(
    api_key="pe_xxxxx",           # Required
    base_url="https://...",       # Optional (default: production)
    timeout=300                   # Optional (default: 300s)
)

Running Tests

# From files
result = client.run_from_files("adapter.yml", "tests.yml")

# From dictionaries
result = client.run(
    adapter={"name": "test", "endpoint": {...}},
    tests=[{"id": "T1", "expected": "..."}]
)

# From raw YAML
result = client.run_from_yaml(yaml_string)

EvalResult Object

result.success          # bool - True if all tests passed
result.success_rate     # float - Percentage of passed tests
result.total_tests      # int - Total number of tests
result.total_passed     # int - Number of passed tests
result.total_failed     # int - Number of failed tests
result.total_errors     # int - Number of errors
result.duration_ms      # float - Total duration in milliseconds
result.test_results     # List[TestResult] - All test results
result.failed_tests     # List[TestResult] - Only failed tests
result.passed_tests     # List[TestResult] - Only passed tests

TestResult Object

test.id_test            # str - Test identifier
test.description        # str - Test description
test.passed             # bool - Whether test passed
test.expected           # str - Expected result
test.actual             # str - Actual result
test.similarity         # float - Semantic similarity (0-1)
test.threshold          # float - Required threshold
test.duration_ms        # float - Test duration
test.error              # str - Error message if any

Account Methods

# Get licenses
licenses = client.get_licenses()
for lic in licenses:
    print(f"{lic.plan}: {lic.tests_remaining} tests remaining")

# Get usage
usage = client.get_usage(license_id)
print(f"Tests this month: {usage['tests_this_month']}")

# Get API keys
keys = client.get_api_keys(license_id)

🔑 Environment Variables

# Set API key (recommended)
export PROMPTEVAL_API_KEY=pe_xxxxx

# Then use without --api-key flag
prompteval run adapter.yml --tests tests.yml

🎯 CI/CD Integration

GitHub Actions

name: LLM Tests

on: [push, pull_request]

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      
      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: '3.11'
      
      - name: Install PromptEval
        run: pip install prompteval
      
      - name: Run Tests
        env:
          PROMPTEVAL_API_KEY: ${{ secrets.PROMPTEVAL_API_KEY }}
        run: |
          prompteval run adapter.yml --tests tests.yml --output results.json
          prompteval report results.json --output report.html
      
      - name: Upload Report
        uses: actions/upload-artifact@v4
        with:
          name: test-report
          path: report.html

📊 Semantic Validation

PromptEval uses sentence transformers to compute semantic similarity between expected and actual outputs. This allows flexible matching that understands meaning, not just exact text.

Example:

  • Expected: "The capital of France is Paris"
  • Actual: "Paris is the capital city of France"
  • Similarity: 94%

Threshold configuration:

  • 0.90+ - Very strict (nearly exact match)
  • 0.75-0.89 - Strict (same meaning, different words)
  • 0.60-0.74 - Moderate (similar concept)
  • <0.60 - Loose (related topic)

🔗 Links

📄 License

Copyright (c) 2026 PromptEval. All Rights Reserved.

This source code is proprietary and confidential. Unauthorized copying, distribution, or use is strictly prohibited.


Made with ❤️ by the PromptEval Team

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

prompteval_core-0.1.5.tar.gz (3.5 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

prompteval_core-0.1.5-py3-none-any.whl (4.2 MB view details)

Uploaded Python 3

File details

Details for the file prompteval_core-0.1.5.tar.gz.

File metadata

  • Download URL: prompteval_core-0.1.5.tar.gz
  • Upload date:
  • Size: 3.5 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.10

File hashes

Hashes for prompteval_core-0.1.5.tar.gz
Algorithm Hash digest
SHA256 bb90c0d5fd394eb1af023d863333892e304ec3b4e13e3746f7a403a2b223486f
MD5 23213dc0c0fd45c8cc6cc2a3e854b201
BLAKE2b-256 350417f5df08aa0251a32f8175dd39738060153224ccab2bc12e6b1fdb973ee8

See more details on using hashes here.

File details

Details for the file prompteval_core-0.1.5-py3-none-any.whl.

File metadata

File hashes

Hashes for prompteval_core-0.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 86d537c802967f7b55d664c319d3ae1f30895fc4a2db4fab291f17002f507207
MD5 c9311f9ec0b42431e5e18f4a57af192a
BLAKE2b-256 4310ed99a28de91a07acbf0a3d28e45743139755873c7a9511309770089591ce

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page