Skip to main content

AutoGEPA

Thin orchestration package around DSPy's GEPA optimizer.
Explore the Documentation »
Report Bug Request Feature

Table of Contents
  1. About
  2. Quick Start
  3. Usage
  4. AutoData
  5. API
  6. Contributing
  7. License

About

AutoGEPA automates the tedious parts of setting up a DSPy optimization pipeline: converting raw data into dspy.Examples, generating a metric file with an LLM, splitting datasets, running baselines, and training with GEPA.

  • Automatic field inference — Maps row columns to DSPy Signature fields automatically
  • LLM-generated metrics — Drafts evaluation metrics automatically, saving them as reproducible .py files for human review
  • End-to-end pipeline — Datasets → Baseline → GEPA optimization → Compare & promote
  • Zero-config defaults — Sensible defaults for all hyperparameters and LMs

Requires DSPy 3.1+.

(back to top)

Quick Start

Install

Install AutoGEPA with uv (recommended):

uv add dspy-auto-gepa

Or with pip:

pip install dspy-auto-gepa

Basic Usage

import dspy
from dspy_auto_gepa import AutoGEPA

# Configure models
lm = dspy.LM("openrouter/openai/gpt-oss-120b")
large_lm = dspy.LM("openrouter/moonshotai/kimi-k2.5")
dspy.configure(lm=lm)

class TicketSignature(dspy.Signature):
    """Classify support tickets."""
    message: str = dspy.InputField()
    urgency: str = dspy.OutputField()
    sentiment: str = dspy.OutputField()

program = dspy.ChainOfThought(TicketSignature)

rows = [
    {"message": "The server room AC is out and equipment is overheating.", "urgency": "high", "sentiment": "negative"},
    {"message": "Can someone clean conference room B next week?", "urgency": "low", "sentiment": "neutral"},
]

# Fields are automatically inferred from the module's signature
auto = AutoGEPA(
    name="TicketSignature-v1_0_0",
    rows=rows,
    module=program,
    metric_lm=large_lm,
    reflection_lm=large_lm,
)

results = auto.run(force=False)  # Set True to re-run even if a saved model exists

# Check if a cached model was loaded
if results.loaded_from:
    print(f"Loaded existing model from {results.loaded_from}")
else:
    print(f"Baseline score: {results.baseline:.4f}")
    print(f"Optimized score: {results.optimized:.4f}")
    print(f"Improvement: {results.improvement:.4f}")
    print(f"Saved optimized program to {results.saved_to}")

(back to top)

Usage

When your row columns match your module's signature fields, you don't need to specify any field mappings. AutoGEPA infers input and output fields directly from the dspy.Signature attached to your module.

If your row columns don't match the module's signature fields, AutoGEPA will raise a clear error telling you which fields are missing and suggesting you use dict mappings:

ValueError: Row columns do not match module signature fields. Missing from rows: ['message', 'sentiment', 'urgency']. Extra in rows: ['msg_text', 'sent', 'urg']. Pass input_fields/output_fields to map row columns to signature fields, or ensure row columns match exactly.

Loading datasets from files

Pass a file path directly as rows — format is auto-detected from the extension:

auto = AutoGEPA(
    rows="data/train.jsonl",   # .jsonl, .json, .csv, .parquet all work
    module=program,
    name="TicketSignature",
    metric_lm=large_lm,
    reflection_lm=large_lm,
)

With dict field mappings

When your row columns have different names than your module's signature fields:

# Row columns: msg_text, urg, sent
# Signature fields: message, urgency, sentiment

auto = AutoGEPA(
    rows=rows,
    module=program,
    name="TicketSignature",
    input_fields={"msg_text": "message"},    # row_col → sig_field
    output_fields={"urg": "urgency", "sent": "sentiment"},
    metric_lm=large_lm,
    reflection_lm=large_lm,
)

results = auto.run()

Advanced: step-by-step control

If you prefer fine-grained control over each stage, you can call the individual methods that run() orchestrates under the hood:

# Optional: generate the metric file first for human inspection
metric_file = auto.build_metric()
print(f"Metric written to {metric_file}")
# After reviewing, proceed:

ds = auto.datasets()

baseline = auto.run_baseline(datasets=ds)

optimized = auto.train(datasets=ds)

final = auto.run_baseline(module=optimized, datasets=ds)

# Or compare and promote
comparison = auto.compare(
    optimized_module=optimized,
    datasets=ds,
)
auto.promote(
    optimized_module=optimized,
    destination=auto._run_dir / "optimized_ticket_classifier.json",
)

(back to top)

AutoData

AutoData is a built-in synthetic data generator that creates training datasets using LLMs. It reads your module's dspy.Signature to understand the task, then generates realistic input/output rows — either from scratch or seeded with a few examples.

Two generation modes cover different use cases:

  • "split" (default) — generates inputs first, then outputs separately. Works well for classification-style tasks where outputs are constrained (e.g., enum labels).
  • "signature" — generates complete rows in one shot using the full signature. Better for complex, tightly-coupled outputs like ReAct reasoning traces.

Quick example

import dspy
from dspy_auto_gepa import AutoData, AutoDataConfig

lm = dspy.LM("openrouter/openai/gpt-oss-120b")
dspy.configure(lm=lm)

class TicketSignature(dspy.Signature):
    """Classify support tickets by urgency and sentiment."""
    message: str = dspy.InputField()
    urgency: str = dspy.OutputField()
    sentiment: str = dspy.OutputField()

program = dspy.ChainOfThought(TicketSignature)

seed_rows = [
    {"message": "The server room AC is out and equipment is overheating.", "urgency": "high", "sentiment": "negative"},
    {"message": "Can someone clean conference room B next week?", "urgency": "low", "sentiment": "neutral"},
    {"message": "Thanks for fixing the VPN, works perfectly now!", "urgency": "medium", "sentiment": "positive"},
]

config = AutoDataConfig(
    n=100,
    generation_mode="split",
    seed_examples=seed_rows,
    output_path=".auto_gepa/TicketSignature/generated/rows.jsonl",
)

gen = AutoData(module=program, data_lm=lm, config=config, name="TicketSignature")
result = gen.generate()

print(f"Generated {result.n_produced} of {result.n_requested} rows")
print(f"Time: {result.generation_time_seconds:.1f}s")
for row in result.rows[:3]:
    print(row)

Feeding generated data into AutoGEPA

The typical workflow is: generate data with AutoData, then pass the rows straight into AutoGEPA for optimization:

from dspy_auto_gepa import AutoData, AutoDataConfig, AutoGEPA

# Step 1: Generate synthetic data
gen = AutoData(module=program, data_lm=lm, config=config, name="TicketSignature")
result = gen.generate()
rows = result.rows

# Step 2: Optimize with AutoGEPA
auto = AutoGEPA(
    name="TicketSignature",
    rows=rows,
    module=program,
    metric_lm=lm,
    reflection_lm=lm,
)
results = auto.run()

Loading seeds from files

Use from_csv or from_json to bootstrap generation from existing data:

gen = AutoData.from_csv("seeds.csv", module=program, data_lm=lm)
result = gen.generate(n=200)

AutoDataConfig reference

Parameter Type Default Description
n int 100 Number of rows to generate
generation_mode Literal["split", "signature"] "split" How to generate rows
seed_examples list[dict] | None None Optional seed rows to guide generation
output_path str | Path | None None Where to save (.jsonl, .csv, .parquet)
diversity_categories str "" Comma-separated topics for diversity (signature mode)
data_lm dspy.LM | None None LM for generation. Falls back to dspy.settings.lm
judge_lm dspy.LM | None None LM for quality scoring. Falls back to data_lm
judge_enabled bool True Score generated rows with an LLM judge
balance_outputs bool True Balance categorical output distribution
num_threads int 16 Parallel generation threads
seed int 42 Random seed
force bool False Overwrite existing output file

GenerationResult

gen.generate() returns a GenerationResult with:

Field Type Description
rows list[dict] Generated rows
n_requested int How many were requested
n_produced int How many were actually generated
n_failed int n_requested - n_produced
generation_time_seconds float Wall-clock time
quality_scores list[float] | None Per-row judge scores (if judge_enabled)

(back to top)

API

Constructor

AutoGEPA(...) accepts all configuration fields directly:

Parameter Type Description
rows list[dict] | DataFrame | str | Path | None Training data. Accepts list[dict], pandas DataFrame, polars DataFrame/LazyFrame, file path (str/Path to .jsonl, .json, .csv, .parquet), or any object with .to_dicts() or .to_pandas()
module dspy.Module | None The DSPy module to optimize
name str | None Task name for artifact subdirectory
input_fields list[str] | dict[str, str] | None Input field names. List for exact match, dict for {row_column: signature_field} mapping. Inferred from signature if omitted
output_fields list[str] | dict[str, str] | None Output field names. Same format as input_fields. Inferred from signature if omitted
metric Path | str | None Path to a custom metric .py file (skips generation)
split tuple[float, ...] Train/val/test split ratios. Default (0.7, 0.2, 0.1)
seed int Random seed for reproducibility. Default 42
artifact_dir Path | str Root directory for artifacts. Default ".auto_gepa"
metric_lm dspy.LM | None LM used for metric generation. Defaults to dspy.LM("openrouter/openai/gpt-oss-120b")
reflection_lm dspy.LM | None LM used for GEPA reflection. Defaults to dspy.LM("openrouter/moonshotai/kimi-k2.5")
gepa_auto Literal["light", "medium", "heavy"] GEPA optimization intensity. Default "light"
num_threads int Parallel threads for evaluation. Default 16
metric_generator_signature Type[dspy.Signature] | None Custom signature class for metric generation
metric_generator_module Type[dspy.Module] | None Custom module class for metric generation

Methods

Method Signature Description
build_metric (rows, module, name, metric, out_path, force=False) → Path Generates the metric .py file explicitly. Skips if a custom metric path is provided. out_path overrides the default save location. Use force=True to overwrite an existing generated metric
run (rows, module, name, metric, force=False) → RunResult Orchestrates the full pipeline: datasets → baseline → train → compare → promote. If force=False and a saved model exists at .auto_gepa/<name>/optimized_<name>.json, loads it and skips training. Returns a RunResult with baseline, optimized, improvement, saved_to (or loaded_from if cached)
datasets (rows, module, name, metric, force=False) → Datasets Converts rows to dspy.Examples and splits into train/val/test. Uses constructor defaults if args omitted. name sets the artifact subdirectory. force=True overwrites an existing metric file
run_baseline (module=None, datasets) → float Evaluates the unoptimized module. Uses module from constructor if not overridden
train (module=None, datasets) → dspy.Module Runs GEPA optimization. Uses module from constructor if not overridden
compare (optimized_module, datasets, baseline_module=None) → dict Side-by-side score comparison. Uses constructor module as baseline_module if not overridden
promote (optimized_module, destination) → Path Saves the optimized program to the given destination
load_metric () → callable Lazily loads the generated metric function

Result Types

  • RunResult — returned by run():
    • baseline: float — score before optimization
    • optimized: float — score after optimization
    • improvement: float — absolute difference
    • saved_to: Path \| None — where the optimized program was saved
    • loaded_from: Path \| None — if a cached model was loaded instead of training
  • Datasets — returned by datasets():
    • train: list[dspy.Example]
    • val: list[dspy.Example]
    • test: list[dspy.Example]

Field Resolution Behaviour

AutoGEPA resolves fields in this order:

  1. Both provided explicitly (list[str] | dict[str, str]) — uses exactly what you gave it. Lists mean exact column names. Dicts mean {row_column: signature_field} mapping.
  2. Neither provided — infers both from the module's DSPy Signature. Raises a clear error if row columns don't match signature fields, listing what's missing and what's extra.
  3. Only one provided — if the other can be inferred from module signature or remaining row keys, great. If not, raises an error.

(back to top)

Contributing

Quick workflow:

  1. Fork and branch: git checkout -b feature/name
  2. Make changes
  3. Commit and push
  4. Open a Pull Request

(back to top)

License

MIT (as declared in pyproject.toml).


Built by thememium

Metadata

Release files for dspy-auto-gepa 0.1.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dspy-auto-gepa 0.1.9
File Size Uploaded
dspy_auto_gepa-0.1.9.tar.gz 34.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dspy-auto-gepa 0.1.9
File Interpreter ABI Platform
dspy_auto_gepa-0.1.9-py3-none-any.whl Python 3 none any Details

Total release size: 72.6 kB

Release files / dspy_auto_gepa-0.1.9.tar.gz

Download URL dspy_auto_gepa-0.1.9.tar.gz
Size 34.6 kB
Tags Source
SHA-256 checksum
How to use checksums
8639d1cc05ddf14313363808192dfe9461a48d54ad2eb49c140b0ad0425cd54c
BLAKE2b-256 checksum
How to use checksums
8a5fc8999908a2daffbea63dd29636265364091001b3be38d8182370d5abc4b3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / dspy_auto_gepa-0.1.9-py3-none-any.whl

Download URL dspy_auto_gepa-0.1.9-py3-none-any.whl
Size 38.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
adedb9c2ca25620b1dcbf3c70f1a4251de5d5cf66e4240b4177e9a027f171492
BLAKE2b-256 checksum
How to use checksums
bd89b1a94dcacee72e7a46be527bc8390c2e747d5ad395dbc6b2d7185d28fd5f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.1.9 This release

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page