Target, Evaluate, Improve: A self-improving loop for agentic systems
Project description
TEI Loop
Target, Evaluate, Improve — a self-improving loop for agentic systems.
TEI wraps any Python agent as a black box, evaluates its output across 4 dimensions using assertion-based LLM judges, and automatically applies targeted fixes when failures are detected.
Install
pip install tei-loop
# With your preferred LLM provider:
pip install 'tei-loop[openai]' # OpenAI
pip install 'tei-loop[anthropic]' # Anthropic
pip install 'tei-loop[google]' # Google Gemini
pip install 'tei-loop[all]' # All providers
Quick Start
import asyncio
from tei_loop import TEILoop
# Your existing agent — any function that takes input and returns output
def my_agent(query: str) -> str:
# ... your agent logic ...
return result
async def main():
loop = TEILoop(agent=my_agent)
# Evaluate only (baseline measurement)
result = await loop.evaluate_only("your test query")
print(result.summary())
# Full TEI loop (evaluate + improve + retry)
result = await loop.run("your test query")
print(result.summary())
# Before/after comparison
comparison = await loop.compare("your test query")
print(f"Baseline: {comparison['baseline'].baseline_score:.2f}")
print(f"With TEI: {comparison['improved'].final_score:.2f}")
asyncio.run(main())
CLI
# Evaluate your agent (baseline)
tei evaluate my_agent.py --query "test input" --verbose
# Run full improvement loop
tei improve my_agent.py --query "test input" --max-retries 3
# Before/after comparison
tei compare my_agent.py --query "test input"
# Generate config file
tei init
TEI CLI looks for a function named agent, run, or main in your Python file.
How It Works
1. Target
Define what success looks like. TEI evaluates across 4 dimensions:
| Dimension | What it checks |
|---|---|
| Target Alignment | Did the agent pursue the correct objective? |
| Reasoning Soundness | Was the reasoning logical and non-contradictory? |
| Execution Accuracy | Were the right tools called with correct parameters? |
| Output Integrity | Is the output complete, accurate, and consistent? |
2. Evaluate
TEI runs 4 LLM judges in parallel (~3-5 seconds). Each judge produces verifiable assertions, not subjective scores. Every claim is backed by evidence from the agent's output.
3. Improve
When a dimension fails, TEI applies the targeted fix strategy:
| Failure | Fix Strategy |
|---|---|
| Target drift | Re-anchor to original objective |
| Flawed reasoning | Regenerate plan with failure context |
| Execution errors | Correct tool calls and parameters |
| Output issues | Repair factual errors and fill gaps |
The loop retries automatically. Each cycle is sharper because the last was diagnosed.
Two Modes
Runtime mode (default): Per-query, 1-3 retries, fixes individual failures in seconds.
Development mode: Across many queries, proposes permanent prompt improvements.
# Development mode
dev_results = await loop.develop(
queries=["query1", "query2", "query3", ...],
max_iterations=50,
)
print(f"Avg improvement: {dev_results['avg_improvement']:+.2f}")
Configuration
TEI auto-detects your LLM provider from environment variables (OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY). No new accounts or API keys needed.
loop = TEILoop(
agent=my_agent,
eval_llm="gpt-5.2", # Smartest for evaluation
improve_llm="gpt-5.2-mini", # Cost-effective for fixes
max_retries=3,
verbose=True,
)
Or via .tei.yaml:
tei init # generates config file
Cost
| Scenario | Cost |
|---|---|
| Agent passes all dimensions | ~$0.005 |
| One improvement cycle | ~$0.025 |
| Full 3-retry loop | ~$0.07 |
TEI shows the cost estimate before running.
Works With
TEI wraps any Python callable. No framework lock-in:
- LangGraph agents
- CrewAI crews
- Custom Python functions
- FastAPI endpoints
- Any callable that takes input and returns output
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file tei_loop-0.1.0.tar.gz.
File metadata
- Download URL: tei_loop-0.1.0.tar.gz
- Upload date:
- Size: 35.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f4094fb29e25feaf96d5949aea316c40e40edb321fe71a5ab246c4f313fe1422
|
|
| MD5 |
be056f0e9cc41e826ca4812350fc44af
|
|
| BLAKE2b-256 |
a32d47a28dccb42d019b42ef660862329638888d54fbbf368565a1866e7ad5b4
|
File details
Details for the file tei_loop-0.1.0-py3-none-any.whl.
File metadata
- Download URL: tei_loop-0.1.0-py3-none-any.whl
- Upload date:
- Size: 45.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.9.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
095e22004ba4ebccf68ab8258dc868bb3a417b73fce8ee839cde2b0b8881a38d
|
|
| MD5 |
0e5198c8ef66db28dd0a6ce35919dd4b
|
|
| BLAKE2b-256 |
68a3cf911e32c4a355207bde395114a43eeb9feb1e009c6d9ff732096121284b
|