Alfred AutoResearch — Standalone Python SDK for autonomous experiment loops — try ideas, keep what works, discard what doesn't.
Project description
AutoResearch SDK
Standalone Python SDK for autonomous experiment loops.
Try ideas, keep what works, discard what doesn't, never stop.
Inspired by karpathy/autoresearch and the OpenClaw AutoResearch plugin.
Features
- 🔄 Experiment Loop Engine — init → run → log → improve → repeat
- 📊 Statistical Confidence — MAD-based noise floor to distinguish real improvements from noise
- 📝 Ideas Backlog — Failed experiments automatically generate ideas for future exploration
- 📈 State Persistence — JSONL results log, survives restarts
- 🛠️ CLI — Full command-line interface with rich output
- 🐍 Pure Python — Zero external dependencies required (only
rich+click) - 📦 OpenClaw Compatible — Same JSONL format as the OpenClaw AutoResearch plugin
Installation
pip install autoresearch
Quick Start
CLI
# Initialize experiment
autoresearch init --name "optimize-latency" --metric latency_ms --direction lower --unit ms
# Run baseline
autoresearch run -c "./benchmark.sh"
# Log result
autoresearch log -d "Baseline measurement"
# Make changes, run again
autoresearch run -c "./benchmark.sh"
autoresearch log -d "Added caching layer"
# Discard bad idea
autoresearch run -c "./benchmark.sh"
autoresearch log -s discard -d "Thread pool didn't help" -i "Try async I/O instead"
# Check status
autoresearch status
Python API
from autoresearch import AutoResearch
# Initialize
ar = AutoResearch(
name="optimize-latency",
metric="latency_ms",
direction="lower", # lower is better
unit="ms",
)
ar.init()
# Run baseline
result = ar.run("./benchmark.sh")
print(f"Passed: {result.passed}, Duration: {result.duration_seconds:.1f}s")
ar.log("Baseline measurement")
# Run experiment
result = ar.run("./benchmark.sh")
if result.passed:
ar.log("Added Redis caching")
else:
ar.log("Crashed", status="crash")
# Discard with idea
result = ar.run("./benchmark.sh")
ar.discard_with_idea("Thread pool overhead", "Try event loop instead")
# Check confidence
print(f"Confidence: {ar.confidence}")
print(f"Best: {ar.best_metric} vs Baseline: {ar.baseline_metric}")
# Export results
import json
print(json.dumps(ar.get_status(), indent=2))
How It Works
The Loop
SCAN → HYPOTHESIS → EXPERIMENT → MEASURE → INTEGRATE → repeat
- SCAN — Identify optimization opportunities
- HYPOTHESIS — Form a testable hypothesis
- EXPERIMENT — Make the change and run the benchmark
- MEASURE — Parse metrics, compare to baseline
- INTEGRATE — Keep if better, discard if worse, log idea for future
Confidence Scoring
AutoResearch uses Median Absolute Deviation (MAD) as the noise floor:
| Confidence | Meaning |
|---|---|
| ≥ 2.0x | Improvement is likely real ✅ |
| 1.0-2.0x | Above noise but marginal ⚠️ |
| < 1.0x | Within noise, re-run to confirm ❓ |
This prevents false positives from natural variance.
Output Format
Your benchmark script should output METRIC name=value lines:
#!/bin/bash
set -euo pipefail
# Run your benchmark
result=$(./my_benchmark)
# Output metrics
echo "METRIC latency_ms=$result"
echo "METRIC throughput_qps=$(echo '10000 / ' $result | bc)"
Persistence
All results are stored in autoresearch.results.jsonl:
{"type": "config", "name": "optimize-latency", "metricName": "latency_ms", "metricUnit": "ms", "bestDirection": "lower"}
{"run": 1, "commit": "abc1234", "metric": 142.5, "metrics": {"latency_ms": 142.5}, "status": "keep", "baseline": true, "description": "Baseline", "timestamp": 1713888000, "segment": 0, "confidence": null}
Ideas Backlog
Discarded experiments append to autoresearch.ideas.md:
- Try async I/O instead of thread pool
- Consider connection pooling for database
- Pre-compute lookup table for hot path
Architecture
autoresearch/
├── src/autoresearch/
│ ├── __init__.py # Public API
│ ├── core/
│ │ ├── engine.py # Experiment loop engine
│ │ ├── confidence.py # Statistical confidence (MAD)
│ │ └── metrics.py # METRIC name=value parser
│ ├── cli.py # Rich CLI interface
│ ├── utils/
│ └── metrics/
├── tests/
│ └── test_autoresearch.py
├── pyproject.toml
└── README.md
Comparison with OpenClaw AutoResearch Plugin
| Feature | OpenClaw Plugin | Python SDK |
|---|---|---|
| Experiment loop | ✅ (via tools) | ✅ (API + CLI) |
| Confidence scoring | ✅ (MAD) | ✅ (MAD, same algorithm) |
| Ideas backlog | ✅ | ✅ |
| State persistence | ✅ (JSONL) | ✅ (JSONL, compatible) |
| Git integration | ✅ (auto checkout) | ⚠️ (commit tracking only) |
| Session management | ✅ (OpenClaw sessions) | ❌ (standalone) |
| Rich output | ✅ | ✅ (click + rich) |
| OpenClaw required | ✅ | ❌ |
| Python importable | ❌ | ✅ |
License
MIT — Use it however you want.
Author
Jose Manuel Sabarís García (@llllJokerllll)
Acknowledgments
- karpathy/autoresearch — Original concept
- OpenClaw AutoResearch Plugin — Confidence algorithm and JSONL format
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file alfred_autoresearch-1.1.0.tar.gz.
File metadata
- Download URL: alfred_autoresearch-1.1.0.tar.gz
- Upload date:
- Size: 14.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
804e143443bcd7e08e85435430c00e49e2e1ce2c07e6c0b6ff6f6d739354794a
|
|
| MD5 |
86ff8589637d54aa5f624bcc768fdf48
|
|
| BLAKE2b-256 |
2c966b5a009fc58f319d9032983c1afd37000ee53779ad8603e6b314d74a0353
|
File details
Details for the file alfred_autoresearch-1.1.0-py3-none-any.whl.
File metadata
- Download URL: alfred_autoresearch-1.1.0-py3-none-any.whl
- Upload date:
- Size: 15.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
92d358aac422f61504a91b00fca6ce828e2ef68506fce337f27ec7672502ae52
|
|
| MD5 |
9fdf1d86fed178fa7c9a063ea90b92be
|
|
| BLAKE2b-256 |
75a2ed9419d6f58205d669ea7e26203db7c002192e10ae104115ca29002e7da7
|