Assayer
Send a prompt to multiple language models in parallel and compare their outputs in the terminal. Useful for evaluating which model handles a given task better, measuring semantic similarity between responses, or running an LLM-as-judge evaluation - without leaving the shell.
Installation
pip install assayer
Similarity scoring requires the optional score extra:
pip install "assayer[score]"
Python 3.11 or newer is required.
Supported Providers
- OpenAI: All GPT models.
- Anthropic: Claude models (Opus 4.7, Sonnet 4.6, Haiku 4.5, and earlier).
- Google Gemini: Gemini 2.x and 3.x models.
- Ollama: Local models running on your machine.
Configuration
Assayer looks for API keys in environment variables or a configuration file at ~/.assayer/config.json.
Environment Variables
export OPENAI_API_KEY="your-key"
export ANTHROPIC_API_KEY="your-key"
export GEMINI_API_KEY="your-key"
Configuration File
{
"OPENAI_API_KEY": "sk-...",
"ANTHROPIC_API_KEY": "sk-ant-...",
"GEMINI_API_KEY": "..."
}
Use assayer models check to verify your configuration.
Quickstart
assayer run "Explain recursion in one sentence." --models gpt-4o,claude-haiku-4-5-20251001
Commands
run
assayer run "prompt" --models gpt-4o,claude-sonnet-4-5
assayer run --prompt-file prompt.txt --models gpt-4o,ollama/llama3.2
assayer run "prompt" --models gpt-4o,claude-sonnet-4-5 --score
assayer run "prompt" --models gpt-4o,claude-sonnet-4-5 --judge gpt-4o --judge-criteria "clarity,brevity"
assayer run "prompt" --models gpt-4o,claude-sonnet-4-5 --output results.json
assayer run "prompt" --models gpt-4o,claude-sonnet-4-5 --output results.csv
assayer run "prompt with {var}" --models gpt-4o --var key=value
| Flag | Description |
|---|---|
--models |
Comma-separated model identifiers (required) |
--prompt-file |
Path to a .txt file instead of an inline prompt |
--var |
KEY=VALUE template variable, repeatable |
--system |
System prompt applied to all models |
--temperature |
Sampling temperature |
--max-tokens |
Maximum output tokens |
--score |
Show pairwise similarity matrix |
--judge |
Model to use as judge |
--judge-criteria |
Comma-separated criteria for the judge |
--output |
Save results to .json or .csv |
--timeout |
Per-model timeout in seconds (default: 30) |
models
assayer models list # list all supported model identifiers
assayer models check # check which API keys are configured
assayer models check ollama # check if Ollama is running and list local models
config
assayer config set OPENAI_API_KEY sk-...
assayer config show
Keys are saved to ~/.assayer/config.json. Environment variables take precedence.
Providers
OpenAI
export OPENAI_API_KEY=sk-...
Supported models: gpt-5.5, gpt-5.5-pro, gpt-5.4, gpt-5.4-pro, gpt-5.4-mini, gpt-5.4-nano, gpt-5.2, gpt-5, gpt-5-mini, gpt-5-nano, gpt-4.1, gpt-4.1-mini, gpt-4.1-nano, gpt-4o, gpt-4o-mini, o3, o3-mini, o4-mini
Anthropic
export ANTHROPIC_API_KEY=sk-ant-...
Supported models: claude-opus-4-7, claude-sonnet-4-6, claude-haiku-4-5-20251001, claude-opus-4-6, claude-sonnet-4-5, claude-opus-4-5
Google Gemini
export GEMINI_API_KEY=...
Supported models: gemini-3.1-pro-preview, gemini-3.1-flash-lite, gemini-3-flash-preview, gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.0-flash, gemini-2.0-flash-lite
Ollama (local)
No API key needed. Start Ollama and use the ollama/ prefix:
ollama serve
assayer run "prompt" --models ollama/llama4-scout,ollama/llama3.2,ollama/qwen3
Scoring
--score embeds all outputs using all-MiniLM-L6-v2 (runs locally, no API call) and displays a pairwise cosine similarity matrix. Values range from 0 (unrelated) to 1 (identical meaning).
LLM-as-judge
--judge <model> sends all outputs to the specified model and asks it to pick a winner. Use --judge-criteria to focus the evaluation:
assayer run "Write a sorting algorithm." \
--models gpt-4o,claude-sonnet-4-5 \
--judge gpt-4o \
--judge-criteria "correctness,readability"
If the judge call fails, a warning is printed to stderr and the run continues normally.
Export
--output results.json saves full results as JSON. --output results.csv saves as CSV. The file format is determined by the extension.
Contributing
Contributions are welcome. See CONTRIBUTING.md for setup instructions, code style, and the PR process.
License
MIT - see LICENSE for details.
Metadata
Release files for assayer 1.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| assayer-1.0.1.tar.gz | 19.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| assayer-1.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 36.8 kB
Release files / assayer-1.0.1.tar.gz
| Download URL | assayer-1.0.1.tar.gz |
|---|---|
| Size | 19.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e3e1a923095648c7cebf0ad5c9fba526cb626cc0456dcf92d2551585e228563a
|
|
BLAKE2b-256 checksum How to use checksums |
2c4c7c0b4e5fa9aec2635968bfab7498599ef91f1699f4d08b7ed67d82bde163
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.0
|
Release files / assayer-1.0.1-py3-none-any.whl
| Download URL | assayer-1.0.1-py3-none-any.whl |
|---|---|
| Size | 17.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bfa3f361d37e3e115cf2bc3b9e57cbd0855ab50f0981ea884d1ace9dbf2f7ea3
|
|
BLAKE2b-256 checksum How to use checksums |
8aa1a0deb7fe3d68b896d6ecd16a1e15b45496686ed63ccaaeaab8157dde9777
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.0
|