Skip to main content

giskardlogo giskardlogo

Evals, Red Teaming and Test Generation for Agentic Systems

Modular, Lightweight, Dynamic and Async-first

GitHub release License Downloads CI Giskard on Discord

Docs • Website • Community


Install

pip install giskard           # checks (+ agents, llm, core)
pip install "giskard[scan]"   # + vulnerability / quality scan
pip install "giskard[openai]" # provider SDK for LLM judges / generators

Requires Python 3.12+.

Extra Adds
(none) giskard-checks and dependencies
scan giskard-scan
openai / anthropic / … provider SDKs (see pyproject.toml optional deps)

Telemetry: optional aggregated analytics via giskard-core. No prompts or outputs are sent. Opt out with export DO_NOT_TRACK=1 or export GISKARD_TELEMETRY_DISABLED=1 (or the same keys in a .env file in the working directory). Set them before import to skip creating ~/.giskard/id; setting them later still stops further sends. Details: giskard-core README.


Giskard is an open-source Python library for testing and evaluating agentic systems. The v3 architecture is a modular set of focused packages — each carrying only the dependencies it needs — built from scratch to wrap anything: an LLM, a black-box agent, or a multi-step pipeline.

Status Package Description
✅ Stable giskard-checks Testing & evaluation — scenario API, built-in checks, LLM-as-judge
✅ Stable giskard-scan Agent vulnerability scanner + RAG/quality evaluation — red teaming, prompt injection, jailbreaks & harmful content (vulnerability_scan, successor of v2 Scan), plus knowledge-base quality eval (quality_scan, successor of v2 RAGET)

These build on three foundational libraries — giskard-core (shared utilities & telemetry), giskard-llm (provider-agnostic LLM routing), and giskard-agents (agent & workflow orchestration) — which are pulled in automatically and rarely used directly.

Giskard Checks — create and apply evals for testing agents

pip install giskard-checks

Giskard Checks is a lightweight library for creating evaluations (evals) that test LLM-based systems — from simple assertions to LLM-as-judge assessments. Unlike traditional unit tests, evals are designed for non-deterministic outputs where the same input can produce different valid responses.

Use Giskard Checks to:

  • Catch regressions — verify your system still behaves correctly after changes
  • Validate RAG quality — check if answers are grounded in retrieved context
  • Enforce safety rules — ensure outputs conform to your content policies
  • Evaluate multi-turn agents — test full conversations, not just single exchanges

Built-in evals include string matching, comparisons, regex, semantic similarity, and LLM-as-judge checks (Groundedness, Conformity, LLMJudge).

Concepts

  • Target — your system under test: any sync/async callable (inputs) -> outputs (optionally with trace)
  • Scenario — one eval: interactions + checks
  • Check — assertion or LLM judge over the trace
  • Suite — many scenarios run together

giskard.agents.Generator is an LLM client for workflows/judges — not the same as giskard.checks input generators (LLMGenerator) that synthesize user messages.

Quickstart

import asyncio
from giskard.checks import Scenario, Groundedness


def get_answer(inputs: str) -> str:
    return "Paris"  # replace with your model / agent


async def main() -> None:
    scenario = (
        Scenario("test_france_capital")
        .interact(inputs="What is the capital of France?", outputs=get_answer)
        .check(
            Groundedness(
                name="answer is grounded",
                context="France is in Western Europe. Its capital is Paris.",
            )
        )
    )
    result = await scenario.run()
    result.print_report()


asyncio.run(main())

Groundedness is an LLM judge — install a provider extra (e.g. pip install "giskard[openai]") and set the matching API key. Default model: openai/gpt-4o-mini.

See the full docs for Suites, LLMJudge, multi-turn scenarios, and more.


Giskard Scan — vulnerability scanner for AI agents

pip install "giskard[scan]"   # or: pip install giskard-scan

Giskard Scan is the red-teaming and vulnerability scanning layer for agentic systems. It generates adversarial test suites automatically from a plain-language description of your agent, covering prompt injection, harmful content, stereotypes, misinformation, and more.

Use Giskard Scan to:

  • Red-team your agent — automatically generate adversarial inputs across OWASP LLM Top-10 threat categories
  • Run prompt-injection probes — built-in dataset of injection payloads ready to use
  • Extend with custom generators — pass your own ScenarioGenerator instances to generate_suite, or register them on vulnerability_suite_generator_registry

Quickstart

import asyncio
from giskard.scan import vulnerability_scan


async def my_agent(inputs: str) -> str:
    # Replace with your agent / model call
    return f"Echo: {inputs}"


async def main() -> None:
    await vulnerability_scan(
        target=my_agent,
        description="A customer support chatbot for an e-commerce platform.",
        languages=["en"],
    )


asyncio.run(main())

Scan generators also need an LLM provider extra and API key (same as Checks judges above).

Looking for Giskard v2?

Giskard v2 included Scan (automatic vulnerability detection) and RAGET (RAG evaluation test set generation).

For LLM agents, both are superseded in v3 by giskard-scan: use vulnerability_scan in place of the v2 LLM scan, and quality_scan (with KnowledgeBase) in place of RAGET.

v3 works with ML models too — wrap one as a target and evaluate it with giskard-checks or giskard-scan. What the examples below cover is the v2-only automatic tabular scan — the detector suite that introspects a giskard.Model + giskard.Dataset to auto-detect performance, bias, and robustness issues — along with the giskard.testing ML test suite and the Giskard Hub. These are not planned for v3.

pip install "giskard[llm]>2,<3"

Scan — automatically detect performance, bias & security issues

Wrap your model and run the scan:

import giskard
import pandas as pd


# Replace my_llm_chain with your actual LLM chain or model inference logic
def model_predict(df: pd.DataFrame):
    """The function takes a DataFrame and must return a list of outputs (one per row)."""
    return [my_llm_chain.run({"query": question}) for question in df["question"]]


giskard_model = giskard.Model(
    model=model_predict,
    model_type="text_generation",
    name="My LLM Application",
    description="A question answering assistant",
    feature_names=["question"],
)

scan_results = giskard.scan(giskard_model)
display(scan_results)

Scan Example

RAGET — generate evaluation datasets for RAG applications

Automatically generate questions, reference answers, and context from your knowledge base:

import pandas as pd
from giskard.rag import generate_testset, KnowledgeBase

# Load your knowledge base documents
df = pd.read_csv("path/to/your/knowledge_base.csv")
knowledge_base = KnowledgeBase.from_pandas(df, columns=["column_1", "column_2"])

testset = generate_testset(
    knowledge_base,
    num_questions=60,
    language="en",
    agent_description="A customer support chatbot for company X",
)

RAGET Example

Full v2 docs

👋 Community

We welcome contributions from the AI community! Read this guide to get started, and join our thriving community on Discord.

Follow the progress and share feedback: v3 Announcement · Roadmap

🌟 Leave us a star, it helps the project to get discovered by others and keeps us motivated to build awesome open-source tools! 🌟

❤️ If you find our work useful, please consider sponsoring us on GitHub. With a monthly sponsoring, you can get a sponsor badge, display your company in this readme, and get your bug reports prioritized. We also offer one-time sponsoring if you want us to get involved in a consulting project, run a workshop, or give a talk at your company.

Metadata

Release files for giskard 3.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for giskard 3.0.1
File Size Uploaded
giskard-3.0.1.tar.gz 12.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for giskard 3.0.1
File Interpreter ABI Platform
giskard-3.0.1-py3-none-any.whl Python 3 none any Details

Total release size: 22.7 kB

Release files / giskard-3.0.1.tar.gz

Download URL giskard-3.0.1.tar.gz
Size 12.2 kB
Tags Source
SHA-256 checksum
How to use checksums
5f2d5901eb4a59e394ab93c9604acdf2a36a3e17e29ebc659f2e96a4b41dbab1
BLAKE2b-256 checksum
How to use checksums
b7c093b1dcb7b47c142d9406234dd53718516887827f6a6c9b9e2aa9a8a1e4c3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / giskard-3.0.1-py3-none-any.whl

Download URL giskard-3.0.1-py3-none-any.whl
Size 10.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
43f0367d60938797186faea20d02d0c76e3fa32388774c7380b4c5ea5168321c
BLAKE2b-256 checksum
How to use checksums
9c1628a1609370c99e908e7d77481c66d22c6533dee98bd1b9a6b28d418d8c33
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.22 {"installer":{"name":"uv","version":"0.12.22","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

3.0.1 This release

2 release files

3.0.0

2 release files

2.19.1

2 release files

2.19.0

2 release files

2.18.0

2 release files

2.17.0

2 release files

2.16.2

2 release files

2.16.1

2 release files

2.16.0

2 release files

2.15.5

2 release files

2.15.4

2 release files

2.15.3

2 release files

2.15.2

2 release files

2.15.1

2 release files

2.14.6

2 release files

2.14.5

2 release files

2.14.4

2 release files

2.14.3

2 release files

2.14.2

2 release files

2.13.0

2 release files

2.11.0

2 release files

2.10.0

2 release files

2.9.1

2 release files

2.9.0

2 release files

2.8.0

2 release files

2.7.7

2 release files

2.7.6

2 release files

2.7.5

2 release files

2.7.4

2 release files

2.7.3

2 release files

2.7.2

2 release files

2.7.1

2 release files

2.7.0

2 release files

2.6.0

2 release files

2.5.3

2 release files

2.5.2

2 release files

2.5.1

2 release files

2.5.0

2 release files

2.4.0

2 release files

2.3.2

2 release files

2.3.1

2 release files

2.2.0

2 release files

2.1.3

2 release files

2.1.2

2 release files

2.1.1

2 release files

2.0.7

2 release files

2.0.6

2 release files

2.0.5

2 release files

2.0.4

2 release files

2.0.3

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.9.4

2 release files

1.9.3

2 release files

1.9.1

2 release files

1.9.0

2 release files

1.8.0

2 release files

1.7.3

2 release files

1.7.2

2 release files

1.7.1

2 release files

1.7.0

2 release files

1.6.0

2 release files

1.5.0

2 release files

1.4.0

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

1 release file

1.0.0

2 release files

0.1.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page