Skip to main content

Benchmark and validate AI memory systems

Project description

Memory Harness

Benchmark and validate AI memory systems. Detect regressions, leakage, and shortcuts.

PyPI

==================================================
 MEMORY BENCHMARK REPORT
==================================================
 Accuracy@1:      83.3%  pass
 Accuracy@3:     100.0%  pass
 Cross-tenant:     0.0%  pass
 Collision:       16.7%  pass
 Confidence:      83.3%  pass
==================================================
 SCORE: 90.0/100  GRADE: A
==================================================

Install

pip install memory-harness httpx

Quick Start

1. Create your dataset (data.jsonl)

{"type":"store","item_id":"doc1","tenant_id":"acme","text":"Customer bought 3 widgets"}
{"type":"store","item_id":"doc2","tenant_id":"acme","text":"Support ticket: login issue"}
{"type":"store","item_id":"doc3","tenant_id":"globex","text":"New user signup from France"}
{"type":"store","item_id":"doc4","tenant_id":"globex","text":"User upgraded to premium"}
{"type":"query","query_id":"q1","tenant_id":"acme","text":"customer purchase","expected_item_id":"doc1"}
{"type":"query","query_id":"q2","tenant_id":"acme","text":"login problem support","expected_item_id":"doc2"}
{"type":"query","query_id":"q3","tenant_id":"globex","text":"new customer france","expected_item_id":"doc3"}
{"type":"query","query_id":"q4","tenant_id":"globex","text":"plan upgrade","expected_item_id":"doc4"}

2. Run benchmark

memorybench dataset -d data.jsonl --provider-endpoint https://your-memory-api.com --n-probe 16

3. Get your score

SCORE: 90.0/100  GRADE: A
PASS (threshold: 70)

Metrics

Metric What it measures Target
Accuracy@1 Exact match rate ≥70%
Accuracy@k Correct item in top-k ≥90%
Cross-tenant Data leakage between tenants <5%
Collision Different queries → same result <20%
Confidence Clear winner (margin) ≥80%

Grading

Grade Score CI Exit
A 90-100 0 (pass)
B 80-89 0 (pass)
C 70-79 0 (pass)
D 60-69 1 (fail)
F <60 1 (fail)

CI Integration

Add to .github/workflows/memory-audit.yml:

name: Memory Audit
on: [push, pull_request]

jobs:
  audit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: '3.11'
      
      - name: Install
        run: pip install memory-harness httpx
      
      - name: Run Memory Benchmark
        run: |
          memorybench dataset \
            -d tests/memory_data.jsonl \
            --provider-endpoint ${{ secrets.MEMORY_API_URL }} \
            --n-probe 16 \
            --pass-threshold 70
      
      - name: Upload Report
        uses: actions/upload-artifact@v4
        if: always()
        with:
          name: memory-report
          path: dataset_report.*

Setup:

  1. Add MEMORY_API_URL to repository secrets
  2. Create tests/memory_data.jsonl with your test data
  3. Push — CI fails if score < 70

Dataset Format

Store items (what to remember):

{"type":"store","item_id":"unique_id","tenant_id":"namespace","text":"content"}

Query items (retrieval tests):

{"type":"query","query_id":"q1","tenant_id":"namespace","text":"search query","expected_item_id":"unique_id"}

Validate before running

memorybench validate -d data.jsonl

CLI Reference

memorybench --version                    # Version
memorybench validate -d FILE             # Validate dataset
memorybench dataset -d FILE [OPTIONS]    # Run benchmark

Options

Flag Default Description
-d, --dataset required JSONL file
--provider-endpoint - Memory API URL
--n-probe 16 Pattern dimension
--pass-threshold 70 Minimum score
-a, --adapter text hash, text, embedding
-o, --output dataset_report.json Report file

Provider API

Your memory endpoint must implement:

POST /reset   {"seed": int}
POST /store   {"pattern": [[float]], "cue": [[float]], "learn_steps": int}
POST /recall  {"cue": [[float]], "steps": int} → {"pattern": [[float]]}

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

memory_harness-1.2.2.tar.gz (13.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

memory_harness-1.2.2-py3-none-any.whl (17.7 kB view details)

Uploaded Python 3

File details

Details for the file memory_harness-1.2.2.tar.gz.

File metadata

  • Download URL: memory_harness-1.2.2.tar.gz
  • Upload date:
  • Size: 13.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.3

File hashes

Hashes for memory_harness-1.2.2.tar.gz
Algorithm Hash digest
SHA256 f79925d6672aa99b4dc580c58cd5f9a4f6585f7a73084a2438c8cbca181c49b6
MD5 609692e04ece2ab801506f3074a23004
BLAKE2b-256 b4fe4a7b49b7530d8af3d7974f4774e2f69aae7c35a59f4f2e2643d5a6be7d0c

See more details on using hashes here.

File details

Details for the file memory_harness-1.2.2-py3-none-any.whl.

File metadata

  • Download URL: memory_harness-1.2.2-py3-none-any.whl
  • Upload date:
  • Size: 17.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.3

File hashes

Hashes for memory_harness-1.2.2-py3-none-any.whl
Algorithm Hash digest
SHA256 336b2a2611ef7635b8614db39105eef049e8b452d0cae5ae35b23f77dd18d6de
MD5 79112c3fadab424d039a11f54e25ed5f
BLAKE2b-256 c9dbc30095a512144283ccde4086dd58d693e9911c36ab675363d720a26a62b6

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page