Skip to main content

answer-engine-benchmark

Asks ChatGPT, Gemini, Perplexity and Claude the questions your buyers ask, with web search on, several times each. Then it counts how often each answer cites your site, what it cites instead, and what it says about you. Every answer is kept in a JSONL file, so you can check any number in the report against the text behind it.

Docs: https://synapsereality.io/open-source/answer-engine-benchmark/

pip install answer-engine-benchmark
export OPENAI_API_KEY=... GEMINI_API_KEY=... PERPLEXITY_API_KEY=... ANTHROPIC_API_KEY=...
aeb run questions/template.yaml --out runs/first \
  --set brand="Acme Analytics" --set domain=acme.example \
  --set category="invoice software" --set audience="small accounting firms"

Why repeat every question

The same question can come back with different sources a minute later. One answer is one sample. The default is 3 runs per question per engine, and every rate in the report is over those runs. aeb noise re-asks a sample of questions later and tells you whether a change you see is bigger than the noise.

The question set

questions/template.yaml holds 24 questions in four groups:

group the buyer is your name in the question
discovery looking for a provider no
comparison learning how to choose no
brand asking about you yes
trust checking you are safe to buy from yes

Discovery and comparison tell you whether engines find you when nobody asked for you. Brand and trust tell you what they say when somebody does.

Fill the four vars in the file, or pass them with --set. A run won't start while any of them still holds its example value. Add your own questions under any group, or new groups. Write them the way a buyer types, and never make a competitor the subject of a question.

The file also sets:

  • ours: domains that count as your citation. acme.example covers its subdomains. github.com/acme covers only paths under it, so a citation of someone else's GitHub repo is not yours.
  • brand_terms: names that count as a mention.
  • stale_markers: phrases that describe you wrongly or out of date, such as an old product or a wrong founding year. One is flagged only within 200 characters of a brand term, so another company "founded in 2015" doesn't count against you.

How an answer is scored

Cited means one of the answer's URLs is yours. The URLs are the sources the engine attached plus any link in the text. Mentioned means a brand term appears in the text. An answer can mention you without citing you, and that difference is worth watching. Position is the rank of your first source among the distinct sites the answer cited.

Failed calls are recorded as errors and left out of every rate. A rate limit never shows up as "not cited".

Gemini returns its sources as Google redirect links. They are resolved to the real pages before scoring, with a cache next to the results. Otherwise every Gemini answer would look like it cited Google.

Engines

engine key default model how it searches
openai OPENAI_API_KEY gpt-5-mini Responses API web_search tool
gemini GEMINI_API_KEY gemini-3.5-flash Google Search grounding
perplexity PERPLEXITY_API_KEY perplexity/sonar Responses API web_search tool
claude ANTHROPIC_API_KEY claude-sonnet-5 Messages API web search tool
claude-cli none, uses your claude login whatever the CLI picks Claude Code's WebSearch and WebFetch tools

Keys are read from the environment only. They are sent in request headers and removed from any error message before it is written, so they never reach the results file. An engine with no key is skipped.

claude-cli is for people who pay for a Claude subscription and have no API key. It runs the Claude Code CLI (claude -p) once per call, with only the web search and fetch tools switched on. It never joins a run on its own: name it with --engine claude-cli. If claude is not on your PATH, the run stops before the first call and says so. Log in once by running claude. On macOS, where the login sits in the Keychain, run claude setup-token and export CLAUDE_CODE_OAUTH_TOKEN instead. Its rows record a cost of 0 and "subscription": true. The calls use up your plan's usage limits and never show on an API bill. A call gets 600 seconds by default. Change that with --timeout, which works for every engine (the API engines default to 180). --model claude-cli=sonnet is passed on to the CLI. Its sources are the search results and fetched pages from the tool calls, plus any link in the answer text.

The CLI's answers are not the API's answers. Claude Code adds its own system prompt. The model and search limits come from your account and plan. Two people can get different results from the same question file. Say which one you used when you publish numbers.

Change a model with --model claude=claude-opus-5. Pick the models your buyers actually use in the apps, or say in the report which ones you used.

Perplexity's API doesn't search unless you ask it to. Without the search tool it still answers, fluently, with no citations. This adapter always turns search on.

Commands

aeb check QUESTIONS [--set ...]           # validate, show which keys are set and how many calls a run needs
aeb run QUESTIONS --out DIR [--set ...]   # ask everything, write DIR/answers.jsonl and DIR/report.md
aeb run QUESTIONS --out DIR --dry-run     # first question of each group, once: a cheap smoke test
aeb report DIR/answers.jsonl [QUESTIONS]  # rebuild the report, re-scoring if you changed the question file
aeb noise QUESTIONS --baseline DIR/answers.jsonl --out DIR2 --sample 5

--anonymise on run and report replaces every domain that isn't yours with "Source A", "Source B" and so on. Use it before you share a report outside your company. --max-calls (default 500) stops a run that would make more calls than you expected.

answers.jsonl gets one line per answer, appended as it arrives, so a stopped run keeps what it already paid for. Each line holds the question, engine, model, run number, the full answer text, every URL, token usage, cost, and the score fields.

What it costs

A full run of the template is 24 questions x 4 engines x 3 runs = 288 calls. Our own run of 360 answers cost $1.39 in API fees without Claude (details in examples/). Adding Claude costs more, because web search on the Claude API is billed per search on top of tokens. Each row carries its cost. Perplexity reports the real cost. The others are estimated from list prices in engines.py, so check them against your bills.

Measure as a stranger

Run the benchmark from an account and machine that has never been told who you are. We learned this from a run through a coding assistant's CLI instead of an API. The CLI passed the operator's own git identity to the model, and many brand answers then told the reader the company was probably their own. The API adapters here send only the question and a one-line instruction to cite sources.

The claude-cli engine guards against the same leak. Each call runs in a new, empty temporary folder with no git repository above it. HOME and the config folder are temporary too, and git is told to read no config at all. Only an allow-list of environment variables gets through (PATH, locale, proxy and CA settings), so GIT_*, USER and any CLAUDE* or ANTHROPIC* variables are dropped. Your CLAUDE.md files, memory, settings, hooks and MCP servers are never loaded. Only your login is copied in, and if the CLI refreshes it during the call, the new one is written back. What cannot be hidden is the account: the answers still come from your Claude login, on your plan, so results depend on the local account.

Example

examples/synapse-launch-week-2026-09.md is a real run on one company's own site, with the published figures only. 105 of 156 answers that named the company cited its site. Of the 204 that didn't name it, 1 did. It is an example of the output, not part of the question set.

Install from source

git clone https://github.com/synapsereality/answer-engine-benchmark
cd answer-engine-benchmark
pip install .
aeb --help

Python 3.10 or later. Needs requests and PyYAML.

Tests

pip install -e ".[test]"
pytest

46 tests. They mock every API and stand in a fake claude for the CLI, so they spend nothing and need no keys or login.

Licence

MIT. Made by Synapse.

Metadata

Release files for answer-engine-benchmark 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for answer-engine-benchmark 0.1.2
File Size Uploaded
answer_engine_benchmark-0.1.2.tar.gz 34.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for answer-engine-benchmark 0.1.2
File Interpreter ABI Platform
answer_engine_benchmark-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 59.9 kB

Release files / answer_engine_benchmark-0.1.2.tar.gz

Download URL answer_engine_benchmark-0.1.2.tar.gz
Size 34.1 kB
Tags Source
SHA-256 checksum
How to use checksums
efae3474a2f57511a5608f764a8b4267962f72d097e8162e34fe8638d4015d15
BLAKE2b-256 checksum
How to use checksums
7b2cc502b5a1d839af0bba4a6df3bc6ba2f9575a2aa616c9a0fb30df5cae6ca7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release files / answer_engine_benchmark-0.1.2-py3-none-any.whl

Download URL answer_engine_benchmark-0.1.2-py3-none-any.whl
Size 25.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7ad59299189043b847cafd7ab39757121b16f3294d88a68df3e50a8f09fa0249
BLAKE2b-256 checksum
How to use checksums
c16682b2e9b8a144b5873f48062ce833d9e771aa5b67fc0f570ef0870a83c4d6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page