Skip to main content

Quantum Memory Graph ⚛️🧠

PyPI version PyPI downloads LongMemEval License: MIT Python 3.9+ Docker

Relationship-aware memory for AI agents. Knowledge graphs + quantum-optimized subgraph selection.

Every memory system treats memories as independent documents — search, rank, stuff into context. But memories aren't independent. They have relationships. "The team chose React" becomes 10x more useful paired with "because of ecosystem maturity" and "FastAPI handles the backend."

🏆 #1 on LongMemEval (ICLR 2025 Benchmark)

Tested on the official LongMemEval benchmark — verified submission.

Method R@1 R@5 R@10 NDCG@10
OMEGA (prev SOTA) — 89.2% 94.1% 87.5%
Mastra OM — 91.0% 95.2% 89.1%
QMG v1.1 (published #1) — 95.8% 98.85% 93.2%
QMG v1.2 — chunked retrieval pipeline 🏆 90.6% 98.6% 99.4% 94.26%
QMG v1.3 — +BM25 hybrid retrieval 🥇 90.0% 99.0% 99.8% 93.32%

Competitor comparison (same benchmark):

System R@5 Source
QMG v1.3 (BM25 hybrid) 99.0% This repo — benchmarks/longmemeval_bm25_hybrid_results.json
QMG v1.2 (chunked gte-large) 98.6% This repo
Mem0 (Apr 2026) 94.8% mem0.ai/research
Mastra OM 91.0% LongMemEval #46
OMEGA (prev SOTA) 89.2% LongMemEval paper

Benchmark run: 500 questions, chunked gte-large embeddings (500-char blocks, 100-char overlap, mean-of-top-3 session scoring). Verified on DGX Spark GB10 (CUDA, ~53 min).

Chunking technique: Each session split into overlapping 500-char chunks → gte-large embedding → per-session score = mean of top-3 chunk scores → rank by score. This recovers the v7 methodology that achieved our original #1, now verified with a clean reproducible pipeline.

BM25 hybrid (v1.3): Keyword matching (BM25) fused with embedding scores at 70/30 ratio using stopword-filtered tokenization. Provides +0.4% R@5 lift at the ceiling — significant when every miss counts. The rank_bm25 package is optional — falls back to embedding-only if not installed.

See: benchmarks/run_longmemeval_chunked_staged.py and benchmarks/run_longmemeval_hybrid.py for exact benchmark code. benchmarks/longmemeval_bm25_hybrid_results.json for full per-question results.

Install

# Python (all platforms)
pip install quantum-memory-graph

# macOS
brew tap Dustin-a11y/qmg && brew install quantum-memory-graph

# Docker
docker pull ghcr.io/dustin-a11y/quantum-memory-graph:latest

# Node.js (thin wrapper)
npm install -g qmg

# Conda (pending — PR #33723)
conda install -c conda-forge quantum-memory-graph

Quick Start

from quantum_memory_graph import store, recall

# Store memories — automatically builds knowledge graph
store("Project Alpha uses React frontend with TypeScript.")
store("Project Alpha backend is FastAPI with PostgreSQL.")
store("FastAPI connects to PostgreSQL via SQLAlchemy ORM.")
store("React components use Material UI for styling.")
store("Team had pizza for lunch. Pepperoni was great.")

# Recall — graph traversal + QAOA finds the optimal combination
result = recall("What is Project Alpha's full tech stack?", K=4)

for memory in result["memories"]:
    print(f"  {memory['text']}")
    print(f"    Connected to {len(memory['connections'])} other selected memories")

Output: Returns React, FastAPI, PostgreSQL, and SQLAlchemy memories — connected, complete, no noise. The pizza memory is excluded because it has no graph connections to the tech stack cluster.

How It Works

Query: "What's the tech stack?"
        │
        ▼
┌─────────────────────┐
│  1. Hybrid Search     │  BM25 keyword + embedding cosine (70/30 fusion)
│     Find neighbors   │  Discovers memories connected to relevant ones
└────────┬────────────┘
         │ 14 candidates
         ▼
┌─────────────────────┐
│  2. Subgraph Data    │  Extract adjacency matrix + relevance scores
│     Build problem    │  Encode relationships as optimization weights
└────────┬────────────┘
         │ NP-hard selection
         ▼
┌─────────────────────┐
│  3. QAOA Optimize    │  Quantum approximate optimization
│     Find best K      │  Maximizes: relevance + connectivity + coverage
└────────┬────────────┘
         │ K memories
         ▼
┌─────────────────────┐
│  4. Return with      │  Each memory includes its connections
│     relationships    │  to other selected memories
└─────────────────────┘

Why Quantum?

Optimal subgraph selection is NP-hard. Given N candidate memories, finding the best K that maximize relevance, connectivity, AND coverage has exponential classical complexity. QAOA provides polynomial-time approximate solutions that beat greedy heuristics — this is the one problem where quantum computing has a genuine algorithmic advantage over classical approaches.

Architecture

Three Layers

  1. Knowledge Graph (graph.py) — Memories are nodes. Relationships are weighted edges based on:

    • Semantic similarity (embedding cosine distance)
    • BM25 keyword matching (70/30 hybrid fusion)
    • Entity co-occurrence (shared people, projects, concepts)
    • Temporal proximity (memories close in time)
    • Source proximity (same conversation/document)
  2. Subgraph Optimizer (subgraph_optimizer.py) — QAOA circuit that maximizes:

    • α × relevance (individual memory scores from hybrid BM25+embedding)
    • β × connectivity (edge weights within selected subgraph)
    • γ × coverage (topic diversity across selection)
  3. Pipeline (pipeline.py) — Unified store() and recall() interface.


## API Server

```bash
pip install quantum-memory-graph[api]
python -m quantum_memory_graph.api

Endpoints:

  • POST /store — Store a memory
  • POST /recall — Graph + QAOA recall
  • POST /store-batch — Batch store
  • GET /stats — Graph statistics
  • GET / — Health check

Advanced Usage

Custom Graph

from quantum_memory_graph import MemoryGraph, recall
from quantum_memory_graph.pipeline import set_graph

# Tune similarity threshold for edge creation
graph = MemoryGraph(similarity_threshold=0.25)
set_graph(graph)

# Store and recall as normal

Tune QAOA Parameters

result = recall(
    "query",
    K=5,
    alpha=0.4,       # Relevance weight
    beta_conn=0.35,   # Connectivity weight  
    gamma_cov=0.25,   # Coverage/diversity weight
    hops=3,           # Graph traversal depth
    top_seeds=7,      # Initial seed nodes
    max_candidates=14, # Max qubits for QAOA
)
def my_recall(memories, query, K):
    # Your recall implementation
    return selected_indices  # List[int]

results = run_benchmark(my_recall, K=5)
print(f"Coverage: {results['avg_coverage']*100:.1f}%")

IBM Quantum Hardware

For production workloads, run QAOA on real quantum hardware:

pip install quantum-memory-graph[ibm]
export IBM_QUANTUM_TOKEN=your_token

Validated on ibm_fez and ibm_kingston backends.

Requirements

  • Python ≥ 3.9
  • sentence-transformers
  • networkx
  • qiskit + qiskit-aer
  • numpy

License

MIT License — Copyright 2026 Coinkong (Chef's Attraction)

Links

Metadata

Release files for quantum-memory-graph 1.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for quantum-memory-graph 1.3.0
File Size Uploaded
quantum_memory_graph-1.3.0.tar.gz 67.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for quantum-memory-graph 1.3.0
File Interpreter ABI Platform
quantum_memory_graph-1.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 154.4 kB

Release files / quantum_memory_graph-1.3.0.tar.gz

Download URL quantum_memory_graph-1.3.0.tar.gz
Size 67.5 kB
Tags Source
SHA-256 checksum
How to use checksums
1256b30266f61b00c0b86eaae77d5407af211340564af05f11120e5c860fdab8
BLAKE2b-256 checksum
How to use checksums
0202020f3d4041c03ce0d08f7df2a89f76ea52ddc0a3f99e915b5c3f1ba59bb5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.15

Release files / quantum_memory_graph-1.3.0-py3-none-any.whl

Download URL quantum_memory_graph-1.3.0-py3-none-any.whl
Size 87.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e51fb0ea1f9fda4c2177ead89496f4a615f5db628f5d093edb12de2834a9ffc7
BLAKE2b-256 checksum
How to use checksums
56228c62d7aa9b531dd47e66eb69ac312688e745a29cd19371d9061889ae2bad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.15

Release history Release notifications | RSS feed

This release

1.3.0 This release

2 release files

1.2.2

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.4.0

1 release file

0.3.0

1 release file

0.2.0

1 release file

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page