Meta-Harness
Agent Output Quality Evaluation and Experience Tracking System
Overview
Meta-Harness is an enterprise-grade Agent quality assurance system, based on the Stanford paper Meta-Harness: End-to-End Optimization of Model Harnesses.
Core Features
| Feature | Description |
|---|---|
| Auto Evaluation | Multi-dimensional scoring after each Agent output |
| Experience Storage | SQLite-based structured storage for all task executions |
| Smart Indexing | Auto-mark high/low score experiences |
| Statistics | Success rate, tool effectiveness analysis |
Evaluation Dimensions
| Dimension | Weight | Description |
|---|---|---|
| Correctness | 30% | Syntax, logic correctness |
| Completeness | 20% | Requirements coverage |
| Efficiency | 15% | Time/space complexity |
| Maintainability | 15% | Code structure |
| Security | 10% | No injection risks |
| Test Coverage | 10% | Has test cases |
Installation
# PyPI (recommended)
pip install meta-harness
# From source
pip install .
Quick Start
1. Evaluate Output
from meta_harness import quick_evaluate
# Evaluate Agent output
result = quick_evaluate("Implement user login", login_code)
print(f"Score: {result.overall_score}")
print(f"Dimensions: {result.scores}")
print(f"Feedback: {result.feedback}")
2. Record Experience
from meta_harness import ExperienceTracker
# Create tracker (auto-creates DB at ~/.meta_harness/)
tracker = ExperienceTracker()
# Record execution experience
record = tracker.record(
task="Implement user login",
output=login_code,
evaluation={"overall_score": 85},
tools_used=["code", "file_writer"],
success=True,
duration_seconds=30
)
print(f"Record ID: {record.id}")
3. Statistics
# Get overall statistics
stats = tracker.get_stats(days=30)
print(f"Total: {stats['total']}")
print(f"Success Rate: {stats['success_rate']}%")
print(f"Average Score: {stats['avg_score']}")
# Tool effectiveness analysis
tool_stats = tracker.analyze_tool_effectiveness()
for tool, stat in tool_stats.items():
print(f"{tool}: {stat['success_rate']} success rate")
Advanced Features
Batch Evaluation
from meta_harness import batch_evaluate
pairs = [
("Task 1", "Output 1"),
("Task 2", "Output 2"),
("Task 3", "Output 3"),
]
results = batch_evaluate(pairs)
for r in results:
print(f"{r.task}: {r.overall_score}")
Experience Search
# Search similar tasks
similar = tracker.search_similar("user auth", limit=5)
for r in similar:
print(f"- {r.task} (score: {r.evaluation.get('overall_score', 'N/A')})")
# Get low score records (need improvement)
low_score = tracker.get_low_score_records(threshold=60)
Data Export
# Export to JSON
tracker.export_json("backup.json", days=30)
# Archive old records
tracker.archive_old(days=90)
# Cleanup (keep only recent 1000)
tracker.cleanup(keep_recent=1000)
CoPaw Integration
For integration with CoPaw Agent framework, see INTEGRATION.md
Project Structure
meta-harness/
├── pyproject.toml
├── README.md # English (this file)
├── README_zh.md # 中文
├── LICENSE
├── CONTRIBUTING.md
├── CHANGELOG.md
├── src/
│ └── meta_harness/
│ ├── __init__.py
│ ├── evaluator/
│ └── tracker/
├── tests/
└── docs/
└── INTEGRATION.md
Dependencies
- Python >= 3.10
- SQLAlchemy >= 2.0
Optional:
memorycoreclaw- For memory system integration
License
Related Links
- MemoryCoreClaw - Human-like memory system
- CoPaw - Agent framework
⭐ If you find this useful, please star!
Metadata
Release files for meta-harness 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| meta_harness-1.0.0.tar.gz | 16.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| meta_harness-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 28.4 kB
Release files / meta_harness-1.0.0.tar.gz
| Download URL | meta_harness-1.0.0.tar.gz |
|---|---|
| Size | 16.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d3c0aaff0f90051a5ec1b50c0b534a2080237ac8bd05bdd4bbdfc93326abc392
|
|
BLAKE2b-256 checksum How to use checksums |
3442bd24d7313953a677c390e58607ef4d42ef890537ca008311bb25418887df
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.11
|
Release files / meta_harness-1.0.0-py3-none-any.whl
| Download URL | meta_harness-1.0.0-py3-none-any.whl |
|---|---|
| Size | 12.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ab7faee961e3a6a5463dae3ed4bb9e7f126d0570428369594ffa316fa0f23b3c
|
|
BLAKE2b-256 checksum How to use checksums |
5317005dc741d89c1788f0c9c15adf836bfc7006d5b4408e8f8fe1cdbff27e19
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.11
|