Skip to main content

Benchmark and validate AI memory systems

Project description

Memory Harness

Benchmark and validate AI memory systems. Detect regressions, leakage, and shortcuts.

PyPI

==================================================
 MEMORY BENCHMARK REPORT
==================================================
 Accuracy@1:      83.3%  pass
 Accuracy@3:     100.0%  pass
 Cross-tenant:     0.0%  pass
 Collision:       16.7%  pass
 Confidence:      83.3%  pass
==================================================
 SCORE: 90.0/100  GRADE: A
==================================================

Install

pip install memory-harness httpx

Quick Start

1. Create your dataset (data.jsonl)

{"type":"store","item_id":"doc1","tenant_id":"acme","text":"Customer bought 3 widgets"}
{"type":"store","item_id":"doc2","tenant_id":"acme","text":"Support ticket: login issue"}
{"type":"store","item_id":"doc3","tenant_id":"globex","text":"New user signup from France"}
{"type":"store","item_id":"doc4","tenant_id":"globex","text":"User upgraded to premium"}
{"type":"query","query_id":"q1","tenant_id":"acme","text":"customer purchase","expected_item_id":"doc1"}
{"type":"query","query_id":"q2","tenant_id":"acme","text":"login problem support","expected_item_id":"doc2"}
{"type":"query","query_id":"q3","tenant_id":"globex","text":"new customer france","expected_item_id":"doc3"}
{"type":"query","query_id":"q4","tenant_id":"globex","text":"plan upgrade","expected_item_id":"doc4"}

2. Run benchmark

memorybench dataset -d data.jsonl --provider-endpoint https://your-memory-api.com --n-probe 16

3. Get your score

SCORE: 90.0/100  GRADE: A
PASS (threshold: 70)

Metrics

Metric What it measures Target
Accuracy@1 Exact match rate ≥70%
Accuracy@k Correct item in top-k ≥90%
Cross-tenant Data leakage between tenants <5%
Collision Different queries → same result <20%
Confidence Clear winner (margin) ≥80%

Grading

Grade Score CI Exit
A 90-100 0 (pass)
B 80-89 0 (pass)
C 70-79 0 (pass)
D 60-69 1 (fail)
F <60 1 (fail)

CI Integration

Add to .github/workflows/memory-audit.yml:

name: Memory Audit
on: [push, pull_request]

jobs:
  audit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: '3.11'
      
      - name: Install
        run: pip install memory-harness httpx
      
      - name: Run Memory Benchmark
        run: |
          memorybench dataset \
            -d tests/memory_data.jsonl \
            --provider-endpoint ${{ secrets.MEMORY_API_URL }} \
            --n-probe 16 \
            --pass-threshold 70
      
      - name: Upload Report
        uses: actions/upload-artifact@v4
        if: always()
        with:
          name: memory-report
          path: dataset_report.*

Setup:

  1. Add MEMORY_API_URL to repository secrets
  2. Create tests/memory_data.jsonl with your test data
  3. Push — CI fails if score < 70

Dataset Format

Store items (what to remember):

{"type":"store","item_id":"unique_id","tenant_id":"namespace","text":"content"}

Query items (retrieval tests):

{"type":"query","query_id":"q1","tenant_id":"namespace","text":"search query","expected_item_id":"unique_id"}

Validate before running

memorybench validate -d data.jsonl

CLI Reference

memorybench --version                    # Version
memorybench validate -d FILE             # Validate dataset
memorybench dataset -d FILE [OPTIONS]    # Run benchmark

Options

Flag Default Description
-d, --dataset required JSONL file
--provider-endpoint - Memory API URL
--n-probe 16 Pattern dimension
--pass-threshold 70 Minimum score
-a, --adapter text hash, text, embedding
-o, --output dataset_report.json Report file

Provider API

Your memory endpoint must implement:

POST /reset   {"seed": int}
POST /store   {"pattern": [[float]], "cue": [[float]], "learn_steps": int}
POST /recall  {"cue": [[float]], "steps": int} → {"pattern": [[float]]}

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

memory_harness-1.3.0.tar.gz (13.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

memory_harness-1.3.0-py3-none-any.whl (19.1 kB view details)

Uploaded Python 3

File details

Details for the file memory_harness-1.3.0.tar.gz.

File metadata

  • Download URL: memory_harness-1.3.0.tar.gz
  • Upload date:
  • Size: 13.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.3

File hashes

Hashes for memory_harness-1.3.0.tar.gz
Algorithm Hash digest
SHA256 deef03e5e1fb4de839cbc9c3a1320fefb282690b576cd954e9865609f444da06
MD5 f05ad7f0580541af4d95ca462144e4b6
BLAKE2b-256 4cee5ef8beee193ba991eeecdb99d2877e167ff591bc41efed0b29ddb9534bcc

See more details on using hashes here.

File details

Details for the file memory_harness-1.3.0-py3-none-any.whl.

File metadata

  • Download URL: memory_harness-1.3.0-py3-none-any.whl
  • Upload date:
  • Size: 19.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.3

File hashes

Hashes for memory_harness-1.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 18c26b379a112ae5dea220c7516f0b239fc2310babd9fb464a063beac5833055
MD5 ec113ce7ea1eea9bb09f03bf6908292e
BLAKE2b-256 7a2003bc86569b94d15d0c1e758c4498211e9d74761d597f9e3dde4fda1e07b3

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page