The open-source prompt engineering toolkit. Test, compare, and evaluate prompts across any LLM.
Project description
PromptRacer
The open-source prompt engineering toolkit. Test, compare, and evaluate your prompts across any LLM — from your terminal or Python code.
Free & open-source alternative to LangSmith, PromptLayer, and Humanloop.
Why PromptRacer?
- One
pip install— no web app, no database, no account needed - Multi-provider — OpenAI, Anthropic, Google Gemini, Ollama (local models)
- Compare models side-by-side — latency, cost, and quality in one table
- LLM-as-judge evaluation — automated scoring with customizable criteria
- Version your prompts — save as YAML, track changes with git
- Template variables — reusable prompts with
{{placeholders}} - CLI + Python API — use it however you want
Install
pip install promptracer
# With specific providers:
pip install promptracer[openai] # OpenAI
pip install promptracer[anthropic] # Anthropic (Claude)
pip install promptracer[google] # Google Gemini
pip install promptracer[all] # All providers
Ollama works out of the box — no extra install needed (just have Ollama running locally).
Quick Start
Python API
from promptracer import Prompt, compare, evaluate
# Create a prompt with template variables
p = Prompt("Translate to {{lang}}: {{text}}")
p.set_vars(lang="English", text="Hola mundo")
# Run against a single model
result = p.run("openai/gpt-4o")
print(result.response) # "Hello world"
print(result.latency) # 0.82
print(result.cost) # 0.0003
# Compare across models
results = compare(p, models=[
"openai/gpt-4o",
"anthropic/claude-sonnet-4-6",
"gemini/gemini-2.5-flash",
"ollama/llama3",
])
results.print_table()
# ┌──────────────────────────────┬──────────────┬─────────┬─────────────────┬─────────┐
# │ Model │ Response │ Latency │ Tokens (in/out) │ Cost │
# ├──────────────────────────────┼──────────────┼─────────┼─────────────────┼─────────┤
# │ openai/gpt-4o [fastest] │ Hello world │ 0.52s │ 24/3 │ $0.0001 │
# │ anthropic/claude-sonnet-4-6 │ Hello world │ 0.61s │ 22/3 │ $0.0001 │
# │ gemini/gemini-2.5-flash │ Hello world │ 0.48s │ 20/3 │ $0.0000 │
# │ ollama/llama3 [cheapest] │ Hello world! │ 1.20s │ 18/4 │ FREE │
# └──────────────────────────────┴──────────────┴─────────┴─────────────────┴─────────┘
# Evaluate quality with LLM-as-judge
eval_result = evaluate(result, criteria="accuracy", judge="openai/gpt-4o-mini")
eval_result.print()
# ╭──────────── PromptRacer Eval ────────────╮
# │ Score: 9/10 │
# │ Criteria: accuracy │
# │ Reasoning: Accurate translation... │
# ╰────────────────────────────────────────╯
Save & Load Prompts as YAML
# Save — version control with git!
p.save("prompts/translate.yaml")
# Load
p2 = Prompt.load("prompts/translate.yaml")
# prompts/translate.yaml
template: "Translate to {{lang}}: {{text}}"
name: translate
system: You are a professional translator
vars:
lang: English
text: Hola mundo
Prompt Versioning
p = Prompt("Translate: {{text}}")
p.update("Translate accurately: {{text}}")
p.update("Translate accurately and naturally: {{text}}")
# View history
p.history() # [{'version': 1, ...}, {'version': 2, ...}, {'version': 3, ...}]
# Compare versions
old, new = p.diff(1, 3)
CLI
# Set your API keys
export OPENAI_API_KEY=sk-...
export ANTHROPIC_API_KEY=sk-ant-...
export GEMINI_API_KEY=AI...
# Run a prompt
promptracer run "What is Python?" -m openai/gpt-4o
# Compare models
promptracer compare "Explain quantum computing" \
-m "openai/gpt-4o,anthropic/claude-sonnet-4-6,ollama/llama3"
# With template variables
promptracer compare "Translate to {{lang}}: {{text}}" \
-v lang=French -v text="Good morning" \
-m "openai/gpt-4o,gemini/gemini-2.5-flash"
# Evaluate a response
promptracer eval "Write a haiku about coding" \
-m openai/gpt-4o \
-j openai/gpt-4o-mini \
-c "creativity and adherence to haiku format"
# Create a prompt template
promptracer init my-prompt
Async Support
import asyncio
from promptracer import Prompt
from promptracer.compare import acompare
async def main():
p = Prompt("Explain {{topic}} simply")
p.set_vars(topic="quantum computing")
# All models run concurrently
results = await acompare(p, models=[
"openai/gpt-4o",
"anthropic/claude-sonnet-4-6",
"ollama/llama3",
])
results.print_table()
asyncio.run(main())
Supported Models
| Provider | Prefix | Example | API Key |
|---|---|---|---|
| OpenAI | openai/ |
openai/gpt-4o |
OPENAI_API_KEY |
| Anthropic | anthropic/ |
anthropic/claude-sonnet-4-6 |
ANTHROPIC_API_KEY |
| Google Gemini | gemini/ or google/ |
gemini/gemini-2.5-flash |
GEMINI_API_KEY |
| Ollama | ollama/ |
ollama/llama3 |
None (local) |
Contributing
Contributions are welcome! Please open an issue or submit a PR.
# Dev setup
git clone https://github.com/Unlucko/promptracer.git
cd promptracer
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
# Run tests
pytest
# Lint
ruff check src/
License
MIT License — use it however you want.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file promptracer-0.1.0.tar.gz.
File metadata
- Download URL: promptracer-0.1.0.tar.gz
- Upload date:
- Size: 13.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b561cbc092ab02d32bf7eadaee203fbea738ca77e546febdf929c21a92922e9b
|
|
| MD5 |
164916e85dc8e4d7e87bc3b53ba12655
|
|
| BLAKE2b-256 |
1486c0321f01956753d859f901ff3e0119651ac849b9a87bb834ba78c3e94dec
|
File details
Details for the file promptracer-0.1.0-py3-none-any.whl.
File metadata
- Download URL: promptracer-0.1.0-py3-none-any.whl
- Upload date:
- Size: 16.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
443374cd8c77a310acf3a76460fb3b8eb4e78fb701671ea57a557235765b0c2a
|
|
| MD5 |
d088191767f6860aff4b305afdbd1550
|
|
| BLAKE2b-256 |
b9b028683e02dc5dfed289dfa24f31f955704f8d4b284f2cc7052c97417fd9c1
|