BenchTrend
Find which benchmarks researchers evaluate on, how their use changes, and which new benchmarks other researchers adopt. Ask in a terminal conversation, or use BenchTrend inside Claude Code or Codex.
Answers include usage counts, the papers covered, and original paper evidence. Usage frequency describes adoption, not benchmark quality.
Install · User guide · Releases · License
Ask a question
Two shortened answers selected after testing four English prompts through MCP with the 2026-10-07 data snapshot. Counts were cross-checked against the returned tool results; model wording can vary.
You: What benchmarks are researchers using in robotics lately?
BenchTrend: Among 509 analysed papers labelled robotics, LIBERO appears in 80 evaluation papers (introducing paper), SIMPLER in 41 (introducing paper), and Meta-World in 29. This covers ICLR 2026, ICML 2026, and NeurIPS 2025. Unlabelled papers, including CoRL, are outside this field filter.
You: How has CIFAR-10 usage changed at ICLR from 2024 to 2026?
BenchTrend: Its stated evaluation use fell from 94.3 to 64.8 to 38.9 per 1,000 analysed papers: 178 of 1,888 in 2024, 197 of 3,038 in 2025, and 165 of 4,237 in 2026. The counts use papers with parsed full text; usage in unparsed papers is unknown.
You can also ask which newly introduced benchmarks other authors use, then
ask for the evidence behind a result. In the standalone terminal, /sources
shows the original sentences and paper links. Quotes and figures are checked;
anything the checker cannot verify is marked in the answer.
Answers include repository, dataset or homepage links and reviewed introducing papers where available. Locations are our mapping; a repository link does not guarantee a direct data download. Missing links mean no confirmed location.
Get started
Terminal conversation — macOS / Linux
With Python 3.10+ in your environment, install and start BenchTrend:
python -m pip install benchtrend
benchtrend
Alternative: install with uv
Choose this instead of pip. With uv installed:
uv tool install benchtrend
benchtrend
uv tool install creates a separate environment for the app. pip installs
into your active Python environment; a virtual environment is recommended.
The first launch downloads the 16 MB research-data snapshot and checks its checksum automatically. Later launches reuse your installed data. This snapshot contains benchmark usage and paper evidence; external benchmarks' underlying questions, labels, and images are separate downloads.
Choose OpenAI or Anthropic at the prompt, accept or change the model, and
enter your API key when asked. The key is hidden and used only for that
session. Existing OPENAI_API_KEY / ANTHROPIC_API_KEY environment variables
also work. API access uses your provider's API billing, separately from a
ChatGPT or Claude subscription.
Once installed, just run benchtrend to return. No GPU is needed.
Windows setup, saved conversations, and data updates.
Use your existing Claude Code or Codex login
BenchTrend can supply the same data to your AI client's conversation through MCP. The client handles model access through its own login. BenchTrend's MCP server needs no separate API key.
After installing BenchTrend, open a conversation in your preferred client. Claude Code or Codex must already be installed and signed in:
| Command | Model access |
|---|---|
benchtrend |
OpenAI or Anthropic API key, with separate API billing |
benchtrend claude |
Your existing Claude Code login |
benchtrend codex |
Your existing Codex login |
These commands download the data if needed and open a new session with BenchTrend tools available. Ask the client to use BenchTrend for your benchmark question. For a fresh installation using this route, follow the complete MCP setup. Your client writes its answers; BenchTrend's final-answer checks run in the standalone terminal interface.
Data coverage
The current snapshot covers 28 editions of 10 conferences, with 58,616 parsed paper-edition observations and 9,332 benchmark and dataset entries. Of those entries, 8,768 have reviewed introduction claims inside this corpus (review rubric).
| Venue | Editions | Venue | Editions |
|---|---|---|---|
| ICML | 2024 · 2025 · 2026 | ICCV | 2023 · 2025 |
| ICLR | 2024 · 2025 · 2026 | ECCV | 2024 · 2026 |
| NeurIPS | 2023 · 2024 · 2025 | ACL | 2024 · 2025 · 2026 |
| CVPR | 2024 · 2025 · 2026 | EMNLP | 2023 · 2024 · 2025 |
| AAAI | 2024 · 2025 · 2026 | CoRL | 2023 · 2024 · 2025 |
Counts come from explicitly stated evaluation or training use in parsed arXiv HTML. Each answer reports its scope and denominator. Missing full text or an unstated role does not establish that a benchmark was unused. Field filters cover papers with matching labels and can exclude unlabelled venues.
“New” means a reviewed introduction claim inside this corpus, which starts in 2023; it does not establish the first public release. Earlier observed use, self-use, use by other authors, and uncertain attribution are recorded separately. Recent introductions have less time to accumulate adoption.
BenchTrend is a research prototype. Released snapshots have fixed identities, and conversations record the snapshot used for their answers. Data rules and tool details.
Releases and development
Code and data are released separately. v0.2.1
is the code — wheel, source distribution and checksums, the same files
published on PyPI — and data-20261007
is the snapshot. Update code with python -m pip install --upgrade benchtrend
or uv tool upgrade benchtrend, using the method you installed with;
data updates are a separate benchtrend data install. See
how releases work.
The earlier paper-reading interface remains available as an evidence surface:
browser guide. Its compatibility command is bellwether.
Paper-view principles are in docs/paper-view.md.
BenchTrend's code is licensed under Apache-2.0.
Metadata
Release files for benchtrend 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| benchtrend-0.2.1.tar.gz | 102.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| benchtrend-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 191.9 kB
Release files / benchtrend-0.2.1.tar.gz
| Download URL | benchtrend-0.2.1.tar.gz |
|---|---|
| Size | 102.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e41b58845e07889bd46ad8d1dfc4c184e3dd40e2c6a4ec0f5a53cad65d87f99b
|
|
BLAKE2b-256 checksum How to use checksums |
3bb16e3ab0e59f7af2e466f88c00dd30a8f69f98eef391bd3af8e6653ecd38b4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.9 {"installer":{"name":"uv","version":"0.12.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / benchtrend-0.2.1-py3-none-any.whl
| Download URL | benchtrend-0.2.1-py3-none-any.whl |
|---|---|
| Size | 89.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6dc05446c91ff9c0efa0d5c1df5efcc2b3c331211f856a84be920e873f958c4b
|
|
BLAKE2b-256 checksum How to use checksums |
0400c216409f26798a4c68634fed9c357fb1273478a41f454b78ed6072792b9a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.9 {"installer":{"name":"uv","version":"0.12.9","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|