Format Guard
Make LLM output safe to consume.
Validate structured AI output, automatically repair schema violations, retry with the model's own validation error, and return a typed object or a safe flagged failure.
Format Guard is a Python library for the failure mode that appears everywhere when LLMs are connected to real software:
LLM response
↓
"Almost valid" JSON
↓
database / API / spreadsheet
↓
💥 validation error, bad data, or downstream failure
Format Guard puts a validation and repair layer between the model and the application:
Raw LLM text
↓
JSON cleanup
↓
Pydantic schema validation
↓
┌───────────────┐
│ Valid output? │
└───────┬───────┘
yes │ no
│
▼
Repair prompt
│
▼
LLM retry
│
▼
Validate again
│
┌─────┴─────┐
│ │
valid failed
│ │
▼ ▼
Typed object Flag / fallback
Why Format Guard?
LLMs are excellent at producing useful content, but production software usually needs something stricter:
- an integer must actually be an integer
- required fields cannot disappear
- extra fields may need to be rejected
- JSON must be parseable
- downstream systems should not receive model prose around JSON
- a failed response should be repairable without rebuilding the entire workflow
- repeated failures should be observable
- applications need a predictable success/failure contract
Format Guard treats formatting and schema correctness as an engineering boundary rather than assuming the model will always obey instructions.
Features
Strict schema validation
Use any Pydantic BaseModel as the contract between your LLM and application.
Automatic repair
When validation fails, Format Guard builds a repair prompt containing the original request, invalid output, validation error, required JSON schema, and explicit repair instructions.
JSON cleanup
Common formatting problems are normalized before validation:
- Markdown JSON fences
- generic Markdown code fences
- trailing commas
Bounded retries
Configure exactly how many repair attempts are allowed.
guard(
schema=Customer,
llm_fn=llm,
prompt="Extract the customer details.",
max_retries=3,
)
Safe fallback
If all retries fail, optionally return a predefined Pydantic object rather than leaving the application without a usable value.
Failure flagging
Every exhausted validation flow is explicitly marked as:
result.flagged is True
Observability
Results expose attempts, repairs, validation failures, repair rate, fallback usage, raw output, final output, and validation errors.
Provider-agnostic design
The core guard accepts a simple callable:
llm_fn(prompt: str) -> str
The validation layer therefore does not need to know whether the model is Groq, OpenAI, Gemini, a local model, or your own service.
Installation
From PyPI
pip install format-guard
Then:
from format_guard import guard
Development installation
git clone https://github.com/Mainak156/FormatGuard.git
cd FormatGuard
python -m venv .fguard
.\.fguard\Scripts\Activate.ps1
python -m pip install -e ".[dev]"
Quick Start
from pydantic import BaseModel
from format_guard import guard
class Customer(BaseModel):
name: str
age: int
email: str
def llm(prompt: str) -> str:
return """
{
"name": "Mainak",
"age": 21,
"email": "mainak@example.com"
}
"""
result = guard(
schema=Customer,
llm_fn=llm,
prompt="Extract the customer's name, age and email.",
max_retries=3,
)
if result.success:
customer = result.value
print(customer.name)
else:
print("Validation failed:", result.error)
Repair Flow
Suppose the first response is:
{
"name": "Mainak",
"age": "twenty one",
"email": "mainak@example.com"
}
The schema requires:
age: int
Format Guard rejects the response and constructs a repair prompt containing the validation failure.
The next model response can then be:
{
"name": "Mainak",
"age": 21,
"email": "mainak@example.com"
}
The result records that the output was repaired:
result.success
# True
result.attempts
# 2
result.repaired
# True
result.flagged
# False
Fallbacks
from format_guard import guard
fallback = Customer(
name="Unknown",
age=0,
email="unknown@example.com",
)
result = guard(
schema=Customer,
llm_fn=llm,
prompt="Extract customer information.",
max_retries=2,
fallback=fallback,
)
if result.fallback_used:
print("Using safe fallback:", result.value)
Metrics
result.metrics.attempts
result.metrics.repairs
result.metrics.validation_failures
result.metrics.repair_rate
result.metrics.successful
Example:
Attempts: 2
Repairs: 1
Validation failures: 1
Repair rate: 50.00%
Successful: True
JSON Cleanup
The deterministic cleanup layer handles:
```json ... `````` ... ```- trailing commas before
}or]
It does not attempt to silently rewrite arbitrary natural-language output into data. Invalid semantic output should reach the repair loop instead.
Groq Integration
Install the optional Groq dependency:
python -m pip install "format-guard[groq]"
Set:
GROQ_API_KEY=your_key_here
Then:
from format_guard import guard
from format_guard.providers import GroqProvider
from pydantic import BaseModel
class Customer(BaseModel):
name: str
age: int
email: str
provider = GroqProvider(
model="openai/gpt-oss-120b",
)
result = guard(
schema=Customer,
llm_fn=provider,
prompt="Extract the customer name, age and email as JSON.",
max_retries=3,
)
print(result.value if result.success else result.error)
LangChain Structured Output
Install:
python -m pip install "format-guard[langchain]"
Then:
from format_guard.providers import LangChainGroqProvider
provider = LangChainGroqProvider(
model="openai/gpt-oss-120b",
)
customer = provider.generate_structured(
prompt="Extract customer information.",
schema=Customer,
)
print(customer)
This adapter is intentionally separate from the raw benchmark path. The benchmark measures the incremental effect of Format Guard rather than provider-native structured-output enforcement.
REST API
Install:
python -m pip install "format-guard[api]"
Start:
uvicorn format_guard.api:app --reload
Health:
GET /health
Validation:
POST /validate
Example request:
{
"prompt": "Extract customer information.",
"schema": {
"name": "string",
"age": "integer",
"email": "string"
},
"max_retries": 3,
"model": "openai/gpt-oss-120b"
}
Architecture
format_guard/
├── __init__.py
├── core.py
├── models.py
├── cleaner.py
├── repair.py
├── exceptions.py
├── api.py
├── providers/
│ ├── base.py
│ ├── mock.py
│ ├── groq.py
│ └── langchain_groq.py
└── benchmark/
├── models.py
├── providers.py
└── runner.py
The architecture separates core validation, provider adapters, the API layer, and benchmark/evaluation code.
Benchmark
The benchmark compares:
Raw LLM output
│
├── BEFORE
│ └── schema validation
│
└── AFTER
└── Format Guard
├── validation
├── repair
└── retry
Dataset
- 200 synthetic prompts
- 5 categories
- 40 prompts per category
- CRM
- Support
- Sales
- Meeting
- Onboarding
Live benchmark models
- GPT-OSS 120B
- GPT-OSS 20B
- Qwen 3.8 27B
The benchmark architecture is provider-agnostic and can be extended with additional adapters.
Benchmark Results
The completed live benchmark evaluated 200 prompts per model, for 600 model/prompt evaluations.
| Model | Before | After | Improvement | Avg. Retries | Repair Rate | Flagged | Extra Cost |
|---|---|---|---|---|---|---|---|
| GPT-OSS 120B | 5% | 100% | +95 pp | 0.95 | 95% | 0 | $0.028606 |
| GPT-OSS 20B | 9% | 100% | +91 pp | 0.91 | 91% | 0 | $0.013078 |
| Qwen 3.8 27B | 0% | 100% | +100 pp | 1.00 | 100% | 0 | $0.115355 |
Interpretation
On this benchmark, Format Guard increased schema-valid output from 0–9% to 100% for all three tested models.
The improvement figures are percentage-point changes, not universal claims about model accuracy.
Repairs introduce additional model calls and therefore additional token cost. Qwen 3.8 27B had the largest measured repair overhead in this run, while GPT-OSS 20B had the lowest.
Benchmark chart
Limitation
These results are benchmark-specific. They should not be interpreted as a universal statement about the tested models on arbitrary production prompts. The benchmark intentionally evaluates raw responses before provider-native structured-output enforcement.
Running the Benchmark
Generate the dataset:
python benchmarks\generate_dataset.py
Smoke test:
python benchmarks\smoke_test.py
Full checkpointed benchmark:
python benchmarks\run_full_benchmark.py
Generate report and chart:
python benchmarks\generate_report.py
Results:
benchmarks/
└── results/
├── groq_full_benchmark.json
└── format_guard_before_after.png
Testing
Current test status:
94 passed in 2.33s
Run:
python -m pytest -v
Coverage includes:
- public API
- validation
- retries
- repair prompts
- fallback behavior
- JSON cleanup
- provider adapters
- Groq integration
- LangChain integration
- benchmark engine
- benchmark dataset
- provider cost tracking
- observability metrics
- REST API
Public API
from format_guard import (
FormatGuard,
guard,
GuardMetrics,
ValidationResult,
)
Important ValidationResult fields:
result.success
result.value
result.attempts
result.repaired
result.error
result.raw_output
result.final_output
result.flagged
result.fallback_used
result.metrics
Use Cases
- Database ingestion
- CRM updates
- Spreadsheet automation
- REST/API payload generation
- Document field extraction
- Agent tool-call validation
- Structured classification pipelines
Typical boundary:
LLM → Pydantic validation → application
Design Principles
- Validation belongs outside the model. Prompting alone should not be the application's only safety mechanism.
- Fail explicitly. Failed output should be represented as failure rather than silently converted into questionable data.
- Retry with context. Repair attempts receive the original request, invalid output, validation error, and schema.
- Keep providers replaceable. The core guard should not depend on one LLM vendor.
- Measure reliability. Repair frequency, attempts, and cost should be observable.
- Prefer deterministic preprocessing. Simple JSON cleanup happens before another model call.
Project Structure
FormatGuard/
├── benchmarks/
│ ├── data/
│ │ └── customer_prompts.json
│ ├── results/
│ │ ├── groq_full_benchmark.json
│ │ └── format_guard_before_after.png
│ ├── generate_dataset.py
│ ├── generate_report.py
│ ├── run_full_benchmark.py
│ └── smoke_test.py
├── examples/
│ ├── benchmark_demo.py
│ ├── groq_example.py
│ └── langchain_structured_example.py
├── src/
│ └── format_guard/
│ ├── __init__.py
│ ├── api.py
│ ├── cleaner.py
│ ├── core.py
│ ├── exceptions.py
│ ├── models.py
│ ├── repair.py
│ ├── benchmark/
│ │ ├── __init__.py
│ │ ├── models.py
│ │ ├── providers.py
│ │ └── runner.py
│ └── providers/
│ ├── __init__.py
│ ├── base.py
│ ├── groq.py
│ ├── langchain_groq.py
│ └── mock.py
├── tests/
├── .env.example
├── .gitignore
├── pyproject.toml
├── README.md
└── LICENSE
Environment
.env.example:
GROQ_API_KEY=
Never commit .env or API keys.
PyPI
The distribution name is:
format-guard
The Python import name is:
format_guard
After configuring the package metadata:
python -m pip install --upgrade build twine
python -m pytest -v
python -m build
python -m twine check dist/*
Then publish:
python -m twine upload dist/*
For GitHub-hosted production releases, PyPI Trusted Publishing through GitHub Actions is preferred over storing a long-lived API token locally.
Roadmap
Completed
- Pydantic schema validation
- JSON cleanup
- Automatic repair loop
- Bounded retries
- Safe fallback
- Failure flagging
- Metrics
- Provider abstraction
- Groq provider
- LangChain structured-output adapter
- FastAPI
/validate - 200-prompt synthetic benchmark
- Checkpointed benchmark runner
- Token/cost tracking
- 3-model live benchmark
- Before/after visualization
- 94 automated tests
Next
- Add more independent LLM providers
- Reach the roadmap target of 6+ LLMs
- Add richer failure categories
- Add latency measurements
- Add confidence intervals
- Add GitHub Actions CI
- Publish stable releases to PyPI
- Integrate Format Guard into larger agent workflows
Contributing
git clone https://github.com/Mainak156/FormatGuard.git
cd FormatGuard
python -m venv .fguard
.\.fguard\Scripts\Activate.ps1
python -m pip install -e ".[dev]"
python -m pytest -v
Please add tests for behavioral changes.
License
MIT License. See LICENSE.
Author
Mainak Sen
AI/ML Developer focused on LLM applications, agent reliability, evaluation, and production-oriented AI systems.
GitHub: https://github.com/Mainak156
LinkedIn: https://www.linkedin.com/in/techmainak001
Format Guard turns unreliable LLM formatting into a validated application contract.
Metadata
Release files for format-guard 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| format_guard-0.1.0.tar.gz | 30.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| format_guard-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 51.2 kB
Release files / format_guard-0.1.0.tar.gz
| Download URL | format_guard-0.1.0.tar.gz |
|---|---|
| Size | 30.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
60ab2f4945b1311d2de31b9a9be08bbfa0d7cc188d6fdab950f90ceead06a8f5
|
|
BLAKE2b-256 checksum How to use checksums |
55f0c76d4d6dabebc4f9c89da899fe742beeea48fc173a45f59a067dfd2cd73e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.8
|
Release files / format_guard-0.1.0-py3-none-any.whl
| Download URL | format_guard-0.1.0-py3-none-any.whl |
|---|---|
| Size | 20.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3abc399f4e59d8d745fc0391bd59b7a226b46b7b80ced56cb0df1b609d501d41
|
|
BLAKE2b-256 checksum How to use checksums |
55a2498697afc1cf74d950864defba0ab9e27dce29e96d47b9ff8043f640c63d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.8
|