Skip to main content

Format Guard

Make LLM output safe to consume.

Validate structured AI output, automatically repair schema violations, retry with the model's own validation error, and return a typed object or a safe flagged failure.

Format Guard is a Python library for the failure mode that appears everywhere when LLMs are connected to real software:

LLM response
     ↓
"Almost valid" JSON
     ↓
database / API / spreadsheet
     ↓
💥 validation error, bad data, or downstream failure

Format Guard puts a validation and repair layer between the model and the application:

Raw LLM text
     ↓
JSON cleanup
     ↓
Pydantic schema validation
     ↓
 ┌───────────────┐
 │ Valid output? │
 └───────┬───────┘
       yes │ no
           │
           ▼
      Repair prompt
           │
           ▼
        LLM retry
           │
           ▼
     Validate again
           │
     ┌─────┴─────┐
     │           │
   valid       failed
     │           │
     ▼           ▼
Typed object   Flag / fallback

Why Format Guard?

LLMs are excellent at producing useful content, but production software usually needs something stricter:

  • an integer must actually be an integer
  • required fields cannot disappear
  • extra fields may need to be rejected
  • JSON must be parseable
  • downstream systems should not receive model prose around JSON
  • a failed response should be repairable without rebuilding the entire workflow
  • repeated failures should be observable
  • applications need a predictable success/failure contract

Format Guard treats formatting and schema correctness as an engineering boundary rather than assuming the model will always obey instructions.

Features

Strict schema validation

Use any Pydantic BaseModel as the contract between your LLM and application.

Automatic repair

When validation fails, Format Guard builds a repair prompt containing the original request, invalid output, validation error, required JSON schema, and explicit repair instructions.

JSON cleanup

Common formatting problems are normalized before validation:

  • Markdown JSON fences
  • generic Markdown code fences
  • trailing commas

Bounded retries

Configure exactly how many repair attempts are allowed.

guard(
    schema=Customer,
    llm_fn=llm,
    prompt="Extract the customer details.",
    max_retries=3,
)

Safe fallback

If all retries fail, optionally return a predefined Pydantic object rather than leaving the application without a usable value.

Failure flagging

Every exhausted validation flow is explicitly marked as:

result.flagged is True

Observability

Results expose attempts, repairs, validation failures, repair rate, fallback usage, raw output, final output, and validation errors.

Provider-agnostic design

The core guard accepts a simple callable:

llm_fn(prompt: str) -> str

The validation layer therefore does not need to know whether the model is Groq, OpenAI, Gemini, a local model, or your own service.

Installation

From PyPI

pip install format-guard

Then:

from format_guard import guard

Development installation

git clone https://github.com/Mainak156/FormatGuard.git
cd FormatGuard

python -m venv .fguard
.\.fguard\Scripts\Activate.ps1

python -m pip install -e ".[dev]"

Quick Start

from pydantic import BaseModel
from format_guard import guard


class Customer(BaseModel):
    name: str
    age: int
    email: str


def llm(prompt: str) -> str:
    return """
    {
        "name": "Mainak",
        "age": 21,
        "email": "mainak@example.com"
    }
    """


result = guard(
    schema=Customer,
    llm_fn=llm,
    prompt="Extract the customer's name, age and email.",
    max_retries=3,
)

if result.success:
    customer = result.value
    print(customer.name)
else:
    print("Validation failed:", result.error)

Repair Flow

Suppose the first response is:

{
  "name": "Mainak",
  "age": "twenty one",
  "email": "mainak@example.com"
}

The schema requires:

age: int

Format Guard rejects the response and constructs a repair prompt containing the validation failure.

The next model response can then be:

{
  "name": "Mainak",
  "age": 21,
  "email": "mainak@example.com"
}

The result records that the output was repaired:

result.success
# True

result.attempts
# 2

result.repaired
# True

result.flagged
# False

Fallbacks

from format_guard import guard

fallback = Customer(
    name="Unknown",
    age=0,
    email="unknown@example.com",
)

result = guard(
    schema=Customer,
    llm_fn=llm,
    prompt="Extract customer information.",
    max_retries=2,
    fallback=fallback,
)

if result.fallback_used:
    print("Using safe fallback:", result.value)

Metrics

result.metrics.attempts
result.metrics.repairs
result.metrics.validation_failures
result.metrics.repair_rate
result.metrics.successful

Example:

Attempts: 2
Repairs: 1
Validation failures: 1
Repair rate: 50.00%
Successful: True

JSON Cleanup

The deterministic cleanup layer handles:

  • ```json ... ```
  • ``` ... ```
  • trailing commas before } or ]

It does not attempt to silently rewrite arbitrary natural-language output into data. Invalid semantic output should reach the repair loop instead.

Groq Integration

Install the optional Groq dependency:

python -m pip install "format-guard[groq]"

Set:

GROQ_API_KEY=your_key_here

Then:

from format_guard import guard
from format_guard.providers import GroqProvider
from pydantic import BaseModel


class Customer(BaseModel):
    name: str
    age: int
    email: str


provider = GroqProvider(
    model="openai/gpt-oss-120b",
)

result = guard(
    schema=Customer,
    llm_fn=provider,
    prompt="Extract the customer name, age and email as JSON.",
    max_retries=3,
)

print(result.value if result.success else result.error)

LangChain Structured Output

Install:

python -m pip install "format-guard[langchain]"

Then:

from format_guard.providers import LangChainGroqProvider

provider = LangChainGroqProvider(
    model="openai/gpt-oss-120b",
)

customer = provider.generate_structured(
    prompt="Extract customer information.",
    schema=Customer,
)

print(customer)

This adapter is intentionally separate from the raw benchmark path. The benchmark measures the incremental effect of Format Guard rather than provider-native structured-output enforcement.

REST API

Install:

python -m pip install "format-guard[api]"

Start:

uvicorn format_guard.api:app --reload

Health:

GET /health

Validation:

POST /validate

Example request:

{
  "prompt": "Extract customer information.",
  "schema": {
    "name": "string",
    "age": "integer",
    "email": "string"
  },
  "max_retries": 3,
  "model": "openai/gpt-oss-120b"
}

Architecture

format_guard/
├── __init__.py
├── core.py
├── models.py
├── cleaner.py
├── repair.py
├── exceptions.py
├── api.py
├── providers/
│   ├── base.py
│   ├── mock.py
│   ├── groq.py
│   └── langchain_groq.py
└── benchmark/
    ├── models.py
    ├── providers.py
    └── runner.py

The architecture separates core validation, provider adapters, the API layer, and benchmark/evaluation code.

Benchmark

The benchmark compares:

Raw LLM output
       │
       ├── BEFORE
       │    └── schema validation
       │
       └── AFTER
            └── Format Guard
                 ├── validation
                 ├── repair
                 └── retry

Dataset

  • 200 synthetic prompts
  • 5 categories
  • 40 prompts per category
  • CRM
  • Support
  • Sales
  • Meeting
  • Onboarding

Live benchmark models

  • GPT-OSS 120B
  • GPT-OSS 20B
  • Qwen 3.8 27B

The benchmark architecture is provider-agnostic and can be extended with additional adapters.

Benchmark Results

The completed live benchmark evaluated 200 prompts per model, for 600 model/prompt evaluations.

Model Before After Improvement Avg. Retries Repair Rate Flagged Extra Cost
GPT-OSS 120B 5% 100% +95 pp 0.95 95% 0 $0.028606
GPT-OSS 20B 9% 100% +91 pp 0.91 91% 0 $0.013078
Qwen 3.8 27B 0% 100% +100 pp 1.00 100% 0 $0.115355

Interpretation

On this benchmark, Format Guard increased schema-valid output from 0–9% to 100% for all three tested models.

The improvement figures are percentage-point changes, not universal claims about model accuracy.

Repairs introduce additional model calls and therefore additional token cost. Qwen 3.8 27B had the largest measured repair overhead in this run, while GPT-OSS 20B had the lowest.

Benchmark chart

Format Guard — Valid Output Rate

Limitation

These results are benchmark-specific. They should not be interpreted as a universal statement about the tested models on arbitrary production prompts. The benchmark intentionally evaluates raw responses before provider-native structured-output enforcement.

Running the Benchmark

Generate the dataset:

python benchmarks\generate_dataset.py

Smoke test:

python benchmarks\smoke_test.py

Full checkpointed benchmark:

python benchmarks\run_full_benchmark.py

Generate report and chart:

python benchmarks\generate_report.py

Results:

benchmarks/
└── results/
    ├── groq_full_benchmark.json
    └── format_guard_before_after.png

Testing

Current test status:

94 passed in 2.33s

Run:

python -m pytest -v

Coverage includes:

  • public API
  • validation
  • retries
  • repair prompts
  • fallback behavior
  • JSON cleanup
  • provider adapters
  • Groq integration
  • LangChain integration
  • benchmark engine
  • benchmark dataset
  • provider cost tracking
  • observability metrics
  • REST API

Public API

from format_guard import (
    FormatGuard,
    guard,
    GuardMetrics,
    ValidationResult,
)

Important ValidationResult fields:

result.success
result.value
result.attempts
result.repaired
result.error
result.raw_output
result.final_output
result.flagged
result.fallback_used
result.metrics

Use Cases

  • Database ingestion
  • CRM updates
  • Spreadsheet automation
  • REST/API payload generation
  • Document field extraction
  • Agent tool-call validation
  • Structured classification pipelines

Typical boundary:

LLM → Pydantic validation → application

Design Principles

  1. Validation belongs outside the model. Prompting alone should not be the application's only safety mechanism.
  2. Fail explicitly. Failed output should be represented as failure rather than silently converted into questionable data.
  3. Retry with context. Repair attempts receive the original request, invalid output, validation error, and schema.
  4. Keep providers replaceable. The core guard should not depend on one LLM vendor.
  5. Measure reliability. Repair frequency, attempts, and cost should be observable.
  6. Prefer deterministic preprocessing. Simple JSON cleanup happens before another model call.

Project Structure

FormatGuard/
├── benchmarks/
│   ├── data/
│   │   └── customer_prompts.json
│   ├── results/
│   │   ├── groq_full_benchmark.json
│   │   └── format_guard_before_after.png
│   ├── generate_dataset.py
│   ├── generate_report.py
│   ├── run_full_benchmark.py
│   └── smoke_test.py
├── examples/
│   ├── benchmark_demo.py
│   ├── groq_example.py
│   └── langchain_structured_example.py
├── src/
│   └── format_guard/
│       ├── __init__.py
│       ├── api.py
│       ├── cleaner.py
│       ├── core.py
│       ├── exceptions.py
│       ├── models.py
│       ├── repair.py
│       ├── benchmark/
│       │   ├── __init__.py
│       │   ├── models.py
│       │   ├── providers.py
│       │   └── runner.py
│       └── providers/
│           ├── __init__.py
│           ├── base.py
│           ├── groq.py
│           ├── langchain_groq.py
│           └── mock.py
├── tests/
├── .env.example
├── .gitignore
├── pyproject.toml
├── README.md
└── LICENSE

Environment

.env.example:

GROQ_API_KEY=

Never commit .env or API keys.

PyPI

The distribution name is:

format-guard

The Python import name is:

format_guard

After configuring the package metadata:

python -m pip install --upgrade build twine
python -m pytest -v
python -m build
python -m twine check dist/*

Then publish:

python -m twine upload dist/*

For GitHub-hosted production releases, PyPI Trusted Publishing through GitHub Actions is preferred over storing a long-lived API token locally.

Roadmap

Completed

  • Pydantic schema validation
  • JSON cleanup
  • Automatic repair loop
  • Bounded retries
  • Safe fallback
  • Failure flagging
  • Metrics
  • Provider abstraction
  • Groq provider
  • LangChain structured-output adapter
  • FastAPI /validate
  • 200-prompt synthetic benchmark
  • Checkpointed benchmark runner
  • Token/cost tracking
  • 3-model live benchmark
  • Before/after visualization
  • 94 automated tests

Next

  • Add more independent LLM providers
  • Reach the roadmap target of 6+ LLMs
  • Add richer failure categories
  • Add latency measurements
  • Add confidence intervals
  • Add GitHub Actions CI
  • Publish stable releases to PyPI
  • Integrate Format Guard into larger agent workflows

Contributing

git clone https://github.com/Mainak156/FormatGuard.git
cd FormatGuard

python -m venv .fguard
.\.fguard\Scripts\Activate.ps1

python -m pip install -e ".[dev]"
python -m pytest -v

Please add tests for behavioral changes.

License

MIT License. See LICENSE.

Author

Mainak Sen

AI/ML Developer focused on LLM applications, agent reliability, evaluation, and production-oriented AI systems.

GitHub: https://github.com/Mainak156

LinkedIn: https://www.linkedin.com/in/techmainak001


Format Guard turns unreliable LLM formatting into a validated application contract.

Metadata

Release files for format-guard 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for format-guard 0.1.0
File Size Uploaded
format_guard-0.1.0.tar.gz 30.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for format-guard 0.1.0
File Interpreter ABI Platform
format_guard-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 51.2 kB

Release files / format_guard-0.1.0.tar.gz

Download URL format_guard-0.1.0.tar.gz
Size 30.3 kB
Tags Source
SHA-256 checksum
How to use checksums
60ab2f4945b1311d2de31b9a9be08bbfa0d7cc188d6fdab950f90ceead06a8f5
BLAKE2b-256 checksum
How to use checksums
55f0c76d4d6dabebc4f9c89da899fe742beeea48fc173a45f59a067dfd2cd73e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.8

Release files / format_guard-0.1.0-py3-none-any.whl

Download URL format_guard-0.1.0-py3-none-any.whl
Size 20.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3abc399f4e59d8d745fc0391bd59b7a226b46b7b80ced56cb0df1b609d501d41
BLAKE2b-256 checksum
How to use checksums
55a2498697afc1cf74d950864defba0ab9e27dce29e96d47b9ff8043f640c63d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.8

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page