june-bench
A pip-installable, reproducible benchmark suite for memory / QA systems — June + pluggable competitors — over LoCoMo, LongMemEval, HotpotQA/2Wiki/MuSiQue, and FinanceBench, with the same data and the same scorer.
pip install june-bench
june-bench list
june-bench run --system echo --dataset smoke --split smoke # offline, no key, no download
Reproduce the June vs Cognee head-to-head
One command runs both systems over the same HotpotQA open-pool, the same answer model, and the same judge, and prints a side-by-side with the metered API cost:
pip install "june-bench[cognee,june-api]" # bundles cognee + fastembed
june-bench reproduce-h2h --key <YOUR_ACCESS_KEY> --questions 100
- Access key — June's endpoint is hardware-limited (not yet funded), so runs are key-gated. Request one at access@januraine.ai; the reply includes your key and this exact command.
- Same-embedder by default — Cognee automatically embeds with
bge-large-en-v1.5, the commodity open model June's dense lane uses, so it's a same-embedder matched run out of the box (nothing to export). This embedder is a disclosed benchmark parameter, not June's moat; pass--embedder <id>to swap it. - You bring an OpenRouter key (prompted) — it pays for both systems' gpt-4o answers (~$21 for the chain-of-thought tier at n=100); the host never holds or pays for it.
- Cognee runs locally (needs RAM + a one-time ~1.3 GB fastembed download); June answers over its endpoint. The command batches the pool upload, blocks the $90 Opus-on-Cognee path, and meters real cost.
june-bench reproduce runs the June-only HotpotQA number the same way; reproduce-retrieval scores
June's recall@k/nDCG/MRR. All three are plain-language and need no JUNE_BENCH_* env vars.
A benchmark is run(system, dataset) → records → score. Two typed ports are the only extension
points:
System— the thing benchmarked.JuneApiSystem(default; a thin HTTP client to June's/v1/answer, so no June source is shipped),JuneLocalSystem([june-local]extra; a source-protected compiled wheel),CogneeSystem([cognee]extra), or any future system as one adapter.Dataset— what it runs on. The four benchmarks behind a registry.
The scorer is the canonical SQuAD/HotpotQA EM/F1 + selective-accuracy/coverage/cost — Cognee-comparable. Tiny smoke fixtures ship in the wheel (offline wiring proof); full splits are fetched, sha-verified, from a pinned release. No score is ever baked into the package — every result row records dataset + scorer + system + model + cost, so a published number is reproducible by a stranger.
Every result row records dataset + scorer + system + model + cost, and no score is baked into the package — so a published number is reproducible by a stranger, with the exact command above.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file june_bench-0.0.31.tar.gz.
File metadata
- Download URL: june_bench-0.0.31.tar.gz
- Upload date:
- Size: 2.1 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1f82c667334c93b94ffd18b93ebedc4a32fc7821c82ecd115e4883df253af931
|
|
| MD5 |
d763901ce5e55d4c8f0ea08969012032
|
|
| BLAKE2b-256 |
68776e5c7e915e951a6bec07fddea75b0a6556333fd9987de25eee17ad78881c
|
Provenance
The following attestation bundles were made for june_bench-0.0.31.tar.gz:
Publisher:
publish-bench.yml on Junemind/june-brain
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
june_bench-0.0.31.tar.gz -
Subject digest:
1f82c667334c93b94ffd18b93ebedc4a32fc7821c82ecd115e4883df253af931 - Sigstore transparency entry: 2084017957
- Sigstore integration time:
-
Permalink:
Junemind/june-brain@bdbd234486c66efa8be20a564129e3f5d2033f1e -
Branch / Tag:
refs/tags/bench-v0.0.31 - Owner: https://github.com/Junemind
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-bench.yml@bdbd234486c66efa8be20a564129e3f5d2033f1e -
Trigger Event:
push
-
Statement type:
File details
Details for the file june_bench-0.0.31-py3-none-any.whl.
File metadata
- Download URL: june_bench-0.0.31-py3-none-any.whl
- Upload date:
- Size: 2.1 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f7d315836e709020c1a2b116f16af9e77cecd5d98bc2e63a80d135b4c4c1d380
|
|
| MD5 |
b5418d47676ca30294b03732ee4c85bc
|
|
| BLAKE2b-256 |
b111970d4c80395e051c08a3c14261a0af1709b1ad900ce7cc7f275a42133c1a
|
Provenance
The following attestation bundles were made for june_bench-0.0.31-py3-none-any.whl:
Publisher:
publish-bench.yml on Junemind/june-brain
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
june_bench-0.0.31-py3-none-any.whl -
Subject digest:
f7d315836e709020c1a2b116f16af9e77cecd5d98bc2e63a80d135b4c4c1d380 - Sigstore transparency entry: 2084018006
- Sigstore integration time:
-
Permalink:
Junemind/june-brain@bdbd234486c66efa8be20a564129e3f5d2033f1e -
Branch / Tag:
refs/tags/bench-v0.0.31 - Owner: https://github.com/Junemind
-
Access:
private
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish-bench.yml@bdbd234486c66efa8be20a564129e3f5d2033f1e -
Trigger Event:
push
-
Statement type: