Inspect Flow is a workflow stack built on Inspect AI that enables research organizations to run AI evaluations at scale

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

Project description

Inspect Flow

Workflow orchestration for Inspect AI that enables you to run evaluations at scale with repeatability and maintainability.

Why Inspect Flow?

As evaluation workflows grow in complexity—running multiple tasks across different models with varying parameters—managing these experiments becomes challenging. Inspect Flow addresses this by providing:

Declarative Configuration: Define complex evaluations with tasks, models, and parameters in type-safe schemas
Repeatable & Shareable: Encapsulated definitions of tasks, models, configurations, and Python dependencies ensure experiments can be reliably repeated and shared
Powerful Defaults: Define defaults once and reuse them everywhere with automatic inheritance
Parameter Sweeping: Matrix patterns for systematic exploration across tasks, models, and hyperparameters

Inspect Flow is designed for researchers and engineers running systematic AI evaluations who need to scale beyond ad-hoc scripts.

Getting Started

Prerequisites

Before using Inspect Flow, you should:

Have familiarity with Inspect AI
Have an existing Inspect evaluation or use one from inspect-evals

Installation

pip install inspect-flow

Optional: VS Code extension

Optionally install the Inspect AI VS Code Extension which includes features for viewing evaluation log files.

Basic Example

FlowSpec is the main entrypoint for defining evaluation runs. At its core, it takes a list of tasks to run. Here's a simple example that runs two evaluations:

from inspect_flow import FlowSpec, FlowTask

FlowSpec(
    log_dir="logs",
    tasks=[
        FlowTask(
            name="inspect_evals/gpqa_diamond",
            model="openai/gpt-4o",
        ),
        FlowTask(
            name="inspect_evals/mmlu_0_shot",
            model="openai/gpt-4o",
        ),
    ],
)

To run the evaluations, run the following command in your shell. This will create a virtual environment for this spec run and install the dependencies. Note that task and model dependencies (like the inspect-evals and openai Python packages) are inferred and installed automatically.

flow run config.py

This will run both tasks and display progress in your terminal.

Progress bar in terminal

Python API

You can run evaluations from Python instead of the command line.

from inspect_flow import FlowSpec, FlowTask
from inspect_flow.api import run

spec = FlowSpec(
    log_dir="logs",
    tasks=[
        FlowTask(
            name="inspect_evals/gpqa_diamond",
            model="openai/gpt-4o",
        ),
        FlowTask(
            name="inspect_evals/mmlu_0_shot",
            model="openai/gpt-4o",
        ),
    ],
)
run(spec=spec)

Matrix Functions

Often you'll want to evaluate multiple tasks across multiple models. Rather than manually defining every combination, use tasks_matrix to generate all task-model pairs:

from inspect_flow import FlowSpec, tasks_matrix

FlowSpec(
    log_dir="logs",
    tasks=tasks_matrix(
        task=[
            "inspect_evals/gpqa_diamond",
            "inspect_evals/mmlu_0_shot",
        ],
        model=[
            "openai/gpt-5",
            "openai/gpt-5-mini",
        ],
    ),
)

To preview the expanded config before running it, you can run the following command in your shell to ensure the generated config is the one that you intend to run.

flow config matrix.py

This command outputs the expanded configuration showing all 4 task-model combinations (2 tasks × 2 models).

log_dir: logs
dependencies:
- inspect-evals
tasks:
- name: inspect_evals/gpqa_diamond
  model:
    name: openai/gpt-5
- name: inspect_evals/gpqa_diamond
  model:
    name: openai/gpt-5-mini
- name: inspect_evals/mmlu_0_shot
  model:
    name: openai/gpt-5
- name: inspect_evals/mmlu_0_shot
  model:
    name: openai/gpt-5-mini

tasks_matrix and models_matrix are powerful functions that can operate on multiple levels of nested matrixes which enable sophisticated parameter sweeping. Let's say you want to explore different reasoning efforts across models—you can achieve this with the models_matrix function.

from inspect_ai.model import GenerateConfig
from inspect_flow import FlowSpec, models_matrix, tasks_matrix

FlowSpec(
    log_dir="logs",
    tasks=tasks_matrix(
        task=[
            "inspect_evals/gpqa_diamond",
            "inspect_evals/mmmu_0_shot",
        ],
        model=models_matrix(
            model=[
                "openai/gpt-5",
                "openai/gpt-5-mini",
            ],
            config=[
                GenerateConfig(reasoning_effort="minimal"),
                GenerateConfig(reasoning_effort="low"),
                GenerateConfig(reasoning_effort="medium"),
                GenerateConfig(reasoning_effort="high"),
            ],
        ),
    ),
)

For even more concise parameter sweeping, use configs_matrix to generate configuration variants. This produces the same 16 evaluations (2 tasks × 2 models × 4 reasoning levels) as above, but with less boilerplate:

from inspect_flow import FlowSpec, configs_matrix, models_matrix, tasks_matrix

FlowSpec(
    log_dir="logs",
    tasks=tasks_matrix(
        task=[
            "inspect_evals/gpqa_diamond",
            "inspect_evals/mmmu_0_shot",
        ],
        model=models_matrix(
            model=[
                "openai/gpt-5",
                "openai/gpt-5-mini",
            ],
            config=configs_matrix(
                reasoning_effort=["minimal", "low", "medium", "high"],
            ),
        ),
    ),
)

Run evaluations

Before running evaluations, preview the resolved configuration with --dry-run:

flow run matrix.py --dry-run

This creates the virtual environment, installs all dependencies, imports tasks from the registry, applies all defaults, and expands all matrix functions—everything except actually running the evaluations. It's invaluable for verifying that dependencies can be installed, tasks are properly configured, and the exact settings are what you expect. Unlike flow config which just parses the config file, --dry-run performs the full setup process.

To run the config:

flow run matrix.py

This will run all 16 evaluations (2 tasks × 2 models × 4 reasoning levels). When complete, you'll find a link to the logs at the bottom of the task results summary.

Log path printed in terminal

To view logs interactively, run:

inspect view --log-dir logs

Eval logs rendered by Inspect View

Learning More

See the following articles to learn more about using Flow:

Flow Concepts: Flow type system, config structure and basics.
Defaults: Define defaults once and reuse them everywhere with automatic inheritance.
Matrixing: Systematic parameter exploration with matrix and with functions.
Reference: Detailed documentation on the Flow Python API and CLI commands.

Development

To work on development of Inspect Flow, clone the repository and install with the -e flag and [dev, doc] optional dependencies:

git clone https://github.com/meridianlabs-ai/inspect_flow
cd inspect_flow
uv sync
source .venv/bin/activate

Optionally install pre-commit hooks via

make hooks

Run linting, formatting, and tests via

make check
make test

Project details

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

jjallaire

Release history Release notifications | RSS feed

0.8.0

Apr 21, 2026

0.7.0

Mar 24, 2026

0.6.0

Mar 16, 2026

0.5.0

Mar 6, 2026

0.4.1

Feb 20, 2026

0.4.0

Feb 19, 2026

0.3.0

Feb 3, 2026

0.2.2

Jan 27, 2026

This version

0.2.1

Jan 23, 2026

0.2.0

Jan 22, 2026

0.1.4

Jan 15, 2026

0.1.3

Jan 6, 2026

0.1.2

Dec 15, 2025

0.1.1

Dec 11, 2025

0.1.0

Dec 5, 2025

0.0.1

Nov 18, 2025

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

inspect_flow-0.2.1.tar.gz (44.6 kB view details)

Uploaded Jan 23, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

inspect_flow-0.2.1-py3-none-any.whl (58.6 kB view details)

Uploaded Jan 23, 2026 Python 3

File details

Details for the file inspect_flow-0.2.1.tar.gz.

File metadata

Download URL: inspect_flow-0.2.1.tar.gz
Upload date: Jan 23, 2026
Size: 44.6 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for inspect_flow-0.2.1.tar.gz
Algorithm	Hash digest
SHA256	`7c10d291ae0f34c0f353cd3095599db94e7d45da6ebfbb7c53c8ba51a2c498b6`
MD5	`2b36b5daff3009565f8d999fbb2b785b`
BLAKE2b-256	`cb5a7a7c1a3c06a0a90f1a203d78ef43710bc71fa1f66816da145c0f7268fcc5`

See more details on using hashes here.

Provenance

The following attestation bundles were made for inspect_flow-0.2.1.tar.gz:

Publisher: release.yaml on meridianlabs-ai/inspect_flow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: inspect_flow-0.2.1.tar.gz
- Subject digest: 7c10d291ae0f34c0f353cd3095599db94e7d45da6ebfbb7c53c8ba51a2c498b6
- Sigstore transparency entry: 846346434
- Sigstore integration time: Jan 23, 2026
Source repository:
- Permalink: meridianlabs-ai/inspect_flow@5858db3a179baed4dadf24913e5e2a8679e3a649
- Branch / Tag: refs/heads/main
- Owner: https://github.com/meridianlabs-ai
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: release.yaml@5858db3a179baed4dadf24913e5e2a8679e3a649
- Trigger Event: push

File details

Details for the file inspect_flow-0.2.1-py3-none-any.whl.

File metadata

Download URL: inspect_flow-0.2.1-py3-none-any.whl
Upload date: Jan 23, 2026
Size: 58.6 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.7

File hashes

Hashes for inspect_flow-0.2.1-py3-none-any.whl
Algorithm	Hash digest
SHA256	`96156db2e61e058618cfc98637253ba0b26768613d43877068135d5f800c69e9`
MD5	`0e86815476189aeedb17b3a5ab17562b`
BLAKE2b-256	`779be58b67ec04dd26595d9328adb1946bef72760015896c3a8e1e2782950c3b`

See more details on using hashes here.

Provenance

The following attestation bundles were made for inspect_flow-0.2.1-py3-none-any.whl:

Publisher: release.yaml on meridianlabs-ai/inspect_flow

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: inspect_flow-0.2.1-py3-none-any.whl
- Subject digest: 96156db2e61e058618cfc98637253ba0b26768613d43877068135d5f800c69e9
- Sigstore transparency entry: 846346441
- Sigstore integration time: Jan 23, 2026
Source repository:
- Permalink: meridianlabs-ai/inspect_flow@5858db3a179baed4dadf24913e5e2a8679e3a649
- Branch / Tag: refs/heads/main
- Owner: https://github.com/meridianlabs-ai
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: release.yaml@5858db3a179baed4dadf24913e5e2a8679e3a649
- Trigger Event: push

inspect-flow 0.2.1

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Project description

Inspect Flow

Why Inspect Flow?

Getting Started

Prerequisites

Installation

Optional: VS Code extension

Basic Example

Python API

Matrix Functions

Run evaluations

Learning More

Development

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance