Bench'd Harness
The neutral benchmark harness for AI memory systems. Every score is independently run, cryptographically signed, and verifiable by anyone.
Leaderboard | Docs | Methodology | Submit Results
Quick Start
pip install benchd-harness
# Generate signing keys
benchd keys generate --out ./keys
# Set your LLM API key (for the judge)
export OPENROUTER_API_KEY=sk-or-...
# Run LongMemEval against your MCP-compatible memory system
benchd run -a mcp -b longmemeval-v1 --judge --key ./keys/private.key \
--adapter-config '{"endpoint": "http://localhost:3000/mcp"}'
# Submit results to the leaderboard
benchd submit ./runs/run_xxx/manifest.signed.json
MCP Systems: Zero-Code Testing
If your memory system exposes an MCP server with ingest and query tools, you don't need to write any adapter code:
benchd run -a mcp -b longmemeval-v1 --judge \
--adapter-config '{"endpoint": "http://localhost:3000/mcp"}'
The MCP adapter auto-discovers your tools and maps them to Bench'd's interface.
Available Benchmarks
| Benchmark | Slug | Questions | What it tests |
|---|---|---|---|
| LongMemEval | longmemeval-v1 |
500 | Recall, temporal reasoning, knowledge updates |
| LoCoMo | locomo-v1 |
1,540 | Multi-session conversational memory |
| Smoke | smoke-memory-v0 |
10 | Quick sanity check |
Built-in Adapters
| Adapter | System | Install |
|---|---|---|
mcp |
Any MCP server | Built-in |
mem0-local |
Mem0 OSS | pip install benchd-harness[mem0] |
langchain-memory |
LangChain | pip install benchd-harness[langchain] |
llamaindex-memory |
LlamaIndex | pip install benchd-harness[llamaindex] |
llm-baseline |
Raw LLM (no memory) | pip install openai |
echo |
Test adapter | Built-in |
Writing a Custom Adapter
from benchd_harness.adapters.base import BaseAdapter
class MyAdapter(BaseAdapter):
@property
def name(self) -> str:
return "my-system"
def setup(self) -> None:
self.client = MyMemoryClient()
def ingest(self, turns: list[dict]) -> None:
for turn in turns:
self.client.add(role=turn["role"], content=turn["content"])
def recall(self, query: str) -> str:
return self.client.search(query).text
def reset(self) -> None:
self.client.clear()
Register in benchd_harness/adapters/__init__.py and run with benchd run -a my-system.
Commands
| Command | Description |
|---|---|
benchd run |
Run a benchmark against a memory system |
benchd submit |
Submit signed results to benchd.ai |
benchd verify |
Verify a signed manifest |
benchd keys generate |
Generate Ed25519 signing keys |
benchd list |
List available adapters and benchmarks |
Signing & Verification
Every run produces an Ed25519-signed manifest containing all inputs, outputs, scores, and failure traces. Anyone can verify:
benchd verify ./runs/run_xxx/manifest.signed.json
Current Results (May 2026)
| # | System | LongMemEval | Status |
|---|---|---|---|
| 1 | LlamaIndex | 59.0% | Verified |
| 1 | LangChain | 59.0% | Verified |
| 3 | LLM Baseline | 57.6% | Verified |
| 4 | Mem0 OSS | 32.4% | Verified |
Full results at benchd.ai/leaderboard.
Links
- Website: benchd.ai
- Leaderboard: benchd.ai/leaderboard
- Docs: benchd.ai/docs
- Submit: benchd.ai/submit
Metadata
Release files for benchd-harness 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| benchd_harness-0.2.0.tar.gz | 67.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| benchd_harness-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 155.0 kB
Release files / benchd_harness-0.2.0.tar.gz
| Download URL | benchd_harness-0.2.0.tar.gz |
|---|---|
| Size | 67.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9c8b988d36a11825c1330de7b105a96fa583dff2e20e8dc0e05b8244e536f23e
|
|
BLAKE2b-256 checksum How to use checksums |
947435099771caffe572c93f1165e191589216d8f598c56e3dee10b3e74cfb35
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.13
|
Release files / benchd_harness-0.2.0-py3-none-any.whl
| Download URL | benchd_harness-0.2.0-py3-none-any.whl |
|---|---|
| Size | 87.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ab3738968df77891fa3276549add62c6494dfe70c1a6abe1aee2d2e301707eb7
|
|
BLAKE2b-256 checksum How to use checksums |
faaab7a39a72e9d86bc4e0475a24de1ca7e14968f3a3acbc3a8fc492279f85f5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.13
|