Skip to main content

BenchTrend

Ask which benchmarks researchers evaluate on, how usage changes across conference editions, and which newly introduced benchmarks other researchers adopt. Answers carry computed counts, dataset coverage and original paper evidence. Usage frequency describes adoption; it does not measure benchmark quality.

Python 3.10+ is required. The runtime uses only Python's standard library; no GPU, search engine, VS Code extension or extraction pipeline is needed.

Install

For a copy-and-paste macOS/Linux installation that includes data and starts the terminal conversation, use the README quick start. For an existing AI client login, use the complete MCP setup below.

From the public repository, or from a checkout:

uv tool install git+https://github.com/hyunyoungnam/BenchTrend
# From a checkout:
uv tool install .
# Alternatively, in a virtual environment:
python -m pip install .

If the browser installation (install.sh) is already on this machine, it has put a bellwether command on PATH, and uv will stop with "Executables already exist". Add --force; the installed package provides the same bellwether command.

An installable wheel can be shared without the source checkout:

uv tool install ./benchtrend-0.2.0-py3-none-any.whl

The wheel includes code, not the research corpus. It runs from any working directory. Installed runtime data normally lives in ~/.benchtrend; a source checkout uses its existing data. An existing ~/.bellwether snapshot is reused when there is no new installation directory. Set BENCHTREND_HOME or pass --home DIR to choose a different data directory. Code updates preserve data and conversations.

This version is prepared for package distribution. uv tool install benchtrend and pip install benchtrend should be advertised only after publishing this package to PyPI; these commands are not a promise that a public release exists.

Windows PowerShell — terminal conversation

Paste this block into PowerShell. It installs uv if needed, installs the tested source revision and published data, then starts BenchTrend. Choose the provider and model at the prompt, then enter a hidden API key.

if (-not (Get-Command uv -ErrorAction SilentlyContinue)) {
    powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
}
$env:PATH = "$env:USERPROFILE\.local\bin;$env:PATH"
uv tool install https://github.com/hyunyoungnam/BenchTrend/archive/refs/tags/v0.2.0.tar.gz
if ($LASTEXITCODE -eq 0) {
    benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e
    if ($LASTEXITCODE -eq 0) { benchtrend }
}

Install for Claude Code or Codex

These macOS/Linux blocks assume the selected AI client is already installed and signed in. Each installs BenchTrend and its data, registers the MCP server, and opens a new client session. No BenchTrend API key is required. Ask the client to use BenchTrend for benchmark questions.

For Codex:

if ! command -v uv >/dev/null 2>&1; then
  curl -LsSf https://astral.sh/uv/install.sh | sh
fi
export PATH="$HOME/.local/bin:$PATH"
uv tool install https://github.com/hyunyoungnam/BenchTrend/archive/refs/tags/v0.2.0.tar.gz &&
benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz \
  --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e &&
benchtrend mcp --connect codex &&
codex

For Claude Code:

if ! command -v uv >/dev/null 2>&1; then
  curl -LsSf https://astral.sh/uv/install.sh | sh
fi
export PATH="$HOME/.local/bin:$PATH"
uv tool install https://github.com/hyunyoungnam/BenchTrend/archive/refs/tags/v0.2.0.tar.gz &&
benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz \
  --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e &&
benchtrend mcp --connect claude &&
claude

On Windows, use the PowerShell installation above, replacing the final benchtrend with benchtrend mcp --connect codex or benchtrend mcp --connect claude, then launch that client.

Install benchmark data

Use a separately distributed snapshot or BenchTrend data bundle:

# The published bundle (GitHub release data-20261007, 16 MB):
benchtrend data install --url https://github.com/hyunyoungnam/BenchTrend/releases/download/data-20261007/benchtrend-data.tar.gz \
    --sha256 d19dcc2c47cc13738188176b5b87534e7e3eb0f0d11c86731bf41e0a1acca50e
# Or a local copy:
benchtrend data install --file ./benchtrend-data.tar.gz
benchtrend data status

--file also accepts the exported benchmark_snapshot.json. The bundle manifest verifies its snapshot checksum; --sha256 additionally verifies the entire download. Installation validates schema and stages an atomic replacement. Failed downloads, invalid snapshots and checksum failures retain existing data. Updating data uses the same install command. Conversations retain their original snapshot identity; continuing a conversation after a data update asks you to start a new one or reinstall the original snapshot.

On the collection machine, build and package reviewed data separately:

PYTHONPATH=src python3 -m bellwether benchmarks
PYTHONPATH=src python3 -m benchtrend data bundle --out dist/benchtrend-data.tar.gz

The bundle contains the self-contained benchmark snapshot and a manifest, plus an adjacent .sha256 file. No raw HTML, embeddings, browser assets or Meilisearch database is required for either terminal conversations or MCP. The corpus and generated bundles remain Git-ignored.

Start a conversation

Choose OpenAI or Anthropic and supply an API key using the environment:

export OPENAI_API_KEY='YOUR_API_KEY'
benchtrend init --provider openai
benchtrend

For Anthropic, set ANTHROPIC_API_KEY and choose --provider anthropic. --model MODEL_ID selects a model supported by your account; the compatibility defaults are gpt-5-mini and claude-opus-5-5. benchtrend init stores the provider, model and language, never the key. If no key is set, an interactive session can prompt for a hidden key used only for that process. On Windows PowerShell use $env:OPENAI_API_KEY = 'YOUR_API_KEY'.

With no configuration or data, an interactive first launch asks for a local data file or HTTPS URL, then provider/model and a session key if needed.

$ benchtrend
> Which benchmarks are researchers using in robotics?
> Which newly introduced ones are used by other authors?
> Show the evidence for the first one.

Answers use the question's language by default; --language ko or en fixes the prose language. Quotes preserve the authors' wording. Progress goes to stderr, and final answers appear after quote and figure verification. The prompt retains field, venue, period and role through the saved tool history.

Commands inside a conversation:

Command Action
/sources Show the last answer's original quotes, paper links and verification marks
/new Start a new conversation with the same model
/chats List saved conversations
/resume ID Continue a saved conversation
/status Inspect installed data, API key presence, and whether Claude Code / Codex have the MCP server registered and connecting
/help List conversation commands
/exit Exit
benchtrend chats
benchtrend --resume                 # newest saved conversation
benchtrend --resume CHAT_ID
benchtrend ask 'Which benchmarks are most used at ICML?' --json
benchtrend ask 'And which are new?' --resume CHAT_ID --json

Ctrl+C cancels a request without saving an unfinished turn. Quotes that fail matching and figures absent from tool results are marked in the answer. A matched quote or number does not prove the surrounding interpretation.

Use within Codex or Claude Code

You can use the same installed package entirely through an existing AI client. No OpenAI/Anthropic API key is needed by the BenchTrend MCP server; your client manages its own model authentication.

benchtrend mcp --connect codex
# Or:
benchtrend mcp --connect claude

These commands register a benchtrend MCP entry using the client's official CLI. Claude registration uses user scope; Codex uses its normal configuration scope. Existing entries for other servers are retained. Use --dry-run to inspect the exact registration command first. An existing entry named benchtrend follows the client's own replacement/error behavior. Restart the client session, then ask it to use BenchTrend for benchmark questions.

For other local MCP clients:

benchtrend mcp --config

This prints a configuration using an absolute Python executable and data directory. The MCP host starts benchtrend mcp automatically as a local stdio process. Running it manually waits for JSON-RPC requests; it is not a chat prompt.

The six exposed tools are benchmark_scope, benchmark_usage, benchmark_trend, new_benchmarks, benchmark_adoption, and benchmark_evidence. All are read-only and use the same invocation/validation path as the independent terminal conversation. Server instructions explain coverage, evaluation/training separation and evidence requirements. External clients write their own final answers; BenchTrend's automatic final answer verification runs in its own conversation interface.

Local data stays on your machine. Questions and returned tool evidence are sent to your chosen model provider during inference. API authentication for the independent terminal is separate from ChatGPT/Claude subscription login.

Build a release

Code versions, data snapshots, and PyPI publication are explained in the release guide.

python -m pip install build
python -m build
python -m pip install dist/benchtrend-0.2.0-py3-none-any.whl

The CI workflow tests Python 3.10 and 3.14, builds distribution artifacts, installs the wheel into a fresh environment and checks an MCP conversation from outside the source checkout. Publishing to PyPI and hosting a data bundle are separate release steps; neither happens during a build.

Provider adapters follow the official OpenAI function calling guide and Anthropic tool-call guide. OpenAI requests use Responses with store=false, preserving reasoning items between calls. Anthropic requests use Messages with paired tool-use/results. Only BenchTrend's read-only tools are available to these API sessions.

Metadata

Release files for benchtrend 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for benchtrend 0.2.0
File Size Uploaded
benchtrend-0.2.0.tar.gz 98.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for benchtrend 0.2.0
File Interpreter ABI Platform
benchtrend-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 187.3 kB

Release files / benchtrend-0.2.0.tar.gz

Download URL benchtrend-0.2.0.tar.gz
Size 98.5 kB
Tags Source
SHA-256 checksum
How to use checksums
c1819386d9c061dea3dabf209e178b0b605b12be8da7dfbcd85c82ff90146c18
BLAKE2b-256 checksum
How to use checksums
e0072018864e5a1e3ae8eaf937b0ede8c71518e550bd06bbb1e869129a9614ad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.9 {"installer":{"name":"uv","version":"0.12.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / benchtrend-0.2.0-py3-none-any.whl

Download URL benchtrend-0.2.0-py3-none-any.whl
Size 88.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
17d01b069f04ca3b639ccdb87e9afe1d0e7b5c2a221b3c59bd5e1f372a767017
BLAKE2b-256 checksum
How to use checksums
8f67ae4db54d5296787eee0fefb5023802dbb3332d5766b7d7e07fe7d3c1241f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.9 {"installer":{"name":"uv","version":"0.12.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page