Skip to main content

june-bench

A pip-installable, reproducible benchmark suite for memory / QA systems — June + pluggable competitors — over LoCoMo, LongMemEval, HotpotQA/2Wiki/MuSiQue, and FinanceBench, with the same data and the same scorer.

pip install june-bench
june-bench list
june-bench run --system echo --dataset smoke --split smoke    # offline, no key, no download

Reproduce the June vs Cognee head-to-head

One command runs both systems over the same HotpotQA open-pool, the same answer model, and the same judge, and prints a side-by-side with the metered API cost:

pip install "june-bench[cognee,june-api]"     # bundles cognee + fastembed
june-bench reproduce-h2h --key <YOUR_ACCESS_KEY> --questions 100
  • Access key — June's endpoint is hardware-limited (not yet funded), so runs are key-gated. Request one at access@januraine.ai; the reply includes your key and this exact command.
  • Same-embedder by default — Cognee automatically embeds with bge-large-en-v1.5, the commodity open model June's dense lane uses, so it's a same-embedder matched run out of the box (nothing to export). This embedder is a disclosed benchmark parameter, not June's moat; pass --embedder <id> to swap it.
  • You bring an OpenRouter key (prompted) — it pays for both systems' gpt-4o answers (~$21 for the chain-of-thought tier at n=100); the host never holds or pays for it.
  • Cognee runs locally (needs RAM + a one-time ~1.3 GB fastembed download); June answers over its endpoint. The command batches the pool upload, blocks the $90 Opus-on-Cognee path, and meters real cost.

june-bench reproduce runs the June-only HotpotQA number the same way; reproduce-retrieval scores June's recall@k/nDCG/MRR. All three are plain-language and need no JUNE_BENCH_* env vars.

A benchmark is run(system, dataset) → records → score. Two typed ports are the only extension points:

  • System — the thing benchmarked. JuneApiSystem (default; a thin HTTP client to June's /v1/answer, so no June source is shipped), JuneLocalSystem ([june-local] extra; a source-protected compiled wheel), CogneeSystem ([cognee] extra), or any future system as one adapter.
  • Dataset — what it runs on. The four benchmarks behind a registry.

The scorer is the canonical SQuAD/HotpotQA EM/F1 + selective-accuracy/coverage/cost — Cognee-comparable. Tiny smoke fixtures ship in the wheel (offline wiring proof); full splits are fetched, sha-verified, from a pinned release. No score is ever baked into the package — every result row records dataset + scorer + system + model + cost, so a published number is reproducible by a stranger, with the exact command above.

Honest-measurement notes

Both systems get the same documents, questions, gold answers, scorer, and embedder. Databases are reset before runs so results are never contaminated by prior state. Costs are measured from provider billing deltas, not estimated.

Protocol notes (read before comparing numbers)

june-bench runs a matched-pair protocol: identical evidence pool, answer model, and judge for every system, scored with EM/F1 plus a fixed LLM judge. Two deliberate differences from the official benchmark settings: the default mode pools QA over the corpus (the official LoCoMo/LongMemEval settings are per-conversation), and default runs use 100-question slices. This makes results directly comparable between systems run here — and NOT comparable to published leaderboard numbers, which use different protocols. Compare systems, not leaderboards.

Note on difficulty: pooling is the harder direction. The official settings give each question its own haystack (e.g. ~40 sessions in LongMemEval_S); the pool is the union of every conversation in the run, so each question faces strictly more distractors — including cross-conversation confusables the official design never tests. Both systems face the same pool.

Dataset licenses

Full splits are fetched from their official sources, sha-verified (see june-bench fetch). The small bundled reproduce/smoke fixtures are subsets of HotpotQA (CC BY-SA 4.0), LongMemEval (MIT), LoCoMo (CC BY-NC 4.0), and FinanceBench (CC BY-NC 4.0) — see DATA_LICENSES.md for attribution and modification notes.

Links

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

june_bench-0.0.32.tar.gz (2.1 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

june_bench-0.0.32-py3-none-any.whl (2.1 MB view details)

Uploaded Python 3

File details

Details for the file june_bench-0.0.32.tar.gz.

File metadata

  • Download URL: june_bench-0.0.32.tar.gz
  • Upload date:
  • Size: 2.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for june_bench-0.0.32.tar.gz
Algorithm Hash digest
SHA256 7d0b0531affde62cbc12c704f5c5af216240f993bf516591af45dcd092735817
MD5 0d6990c21bf0698814040cf66bfd44aa
BLAKE2b-256 f0552624d1c686a8b0bd25d97060ccfb39fc8b77accf1b35618d075bfa4f8ba2

See more details on using hashes here.

File details

Details for the file june_bench-0.0.32-py3-none-any.whl.

File metadata

  • Download URL: june_bench-0.0.32-py3-none-any.whl
  • Upload date:
  • Size: 2.1 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for june_bench-0.0.32-py3-none-any.whl
Algorithm Hash digest
SHA256 0d2b43ac59d8c6fdf3218bad768f2161e7db1fa00457491da909aac3e140ef9e
MD5 f4715c3f7477af0658460bf7730fb925
BLAKE2b-256 b1e1a989780dd56729b3d8a5561d7f5e03817b6c283297c1a20b54136ebb9fab

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page