Skip to main content

Memoria 1.0.0

Local-first, LLM-agnostic memory infrastructure with parallel hybrid retrieval, multi-signal ranking, declarative type routing, temporal retrieval, and a plugin-based architecture.

Memoria provides a complete memory substrate without requiring a cloud service, API key, GPU, or even an LLM.

  • CPU-only · ~4 GB RAM capable · No cloud dependency · No API keys required
  • LLM optional · MIT licensed · Python
  • CLI · TUI · GUI · API · MCP
  • Pluggable retrieval and extension architecture

Memoria 1.0.0 is the first official public release.

Documentation

Full documentation, architecture notes, configuration reference, adapter documentation, retrieval details, and benchmark methodology:

https://kitzkatz.github.io/memoria/


What Is Memoria?

Memoria is memory infrastructure first — built to store, retrieve, rank, route, and manage information independently of whichever language model happens to consume it.

Rather than treating memory as a thin vector-search layer, Memoria combines multiple retrieval and ranking signals, with temporal retrieval running as a separately evaluable path that merges in at the end:

                    ┌─────────────────────┐
                    │       Query         │
                    └──────────┬──────────┘
                               │
                 ┌─────────────┴─────────────┐
                 │                           │
          Base Retrieval              Temporal Retrieval
                 │                           │
        ┌────────┼────────┐                  │
        │        │        │                  │
      FAISS    BM25    Graph             Temporal
        │        │        │                  │
        └────────┴────────┘                  │
                 │                           │
               Fusion                        │
                 │                           │
            Base Ranking                     │
                 │                           │
                 └─────────────┬─────────────┘
                               │
                         Final RRF Fusion
                               │
                         Ranked Results

The system remains useful without an LLM. An LLM can be connected later as a consumer, generator, or optional reasoning layer.


Results

LongMemEval-S

Evaluated against LongMemEval-S on 500 questions: 470 were retrieval-evaluable, 30 were intentional abstentions. Of the 470, 468 returned results and 2 were retrieval failures.

Metric Result
Recall@1 89.8%
Recall@3 96.4%
Recall@5 97.9%
Recall@10 98.9%
Recall@30 99.6%
Recall@50 99.6%
Session NDCG@10 0.9257

Full benchmark configuration, reproduction steps, and dataset handling are documented separately.

Synthetic Benchmark

A separate 4,632-question synthetic benchmark:

Metric Result
Questions 4,632
Queries returning results 99.46%
Recall@1 32.60%
Recall@3 39.98%
Recall@5 52.03%
Recall@10 78.76%
Average query time ~122.3 ms

These two benchmarks measure different things — full methodology and interpretation live in the benchmark documentation.


Benchmark Environment

The published numbers above were run on deliberately modest hardware, not a dedicated workstation:

Intel Celeron N4020 @ 1.10 GHz · ~3.7 GiB RAM · no GPU · Debian Linux · CPU-only

Memory Usage

Measured peak resident memory on this system:

Workload Peak RSS
LongMemEval — cached 2.65 GiB
LongMemEval — uncached 2.65 GiB
Average individual query 580 MiB

These are real measurements from the benchmark environment, not estimates.


Architecture

Memoria is organized around independent subsystems rather than a monolithic retrieval pipeline.

                         MemorySystem
                              │
          ┌───────────────────┼───────────────────┐
          │                   │                   │
       Storage             Routing             Query
          │                   │                   │
      Database          Type / Signal       Processing
      Vector Store          Routing              │
          │                   │                  │
          └───────────────────┼──────────────────┘
                              │
                         Blackboard
                              │
                         Scheduler
                              │
              ┌───────────────┼───────────────┐
              │               │               │
            FAISS           BM25           Graph
              │               │               │
           Phrase         Attribute        Temporal
              │               │               │
              └───────────────┴───────────────┘
                              │
                           Fusion
                              │
                           Ranking
                              │
                         Final Results

Major subsystems: FAISS vector retrieval, BM25 lexical retrieval, graph retrieval, phrase retrieval, attribute retrieval, temporal retrieval, multi-signal fusion, ranking/reranking, declarative type routing, signal routing, blackboard scheduling, persistent storage, plugin hooks, and external adapters.


Retrieval

Memoria runs multiple retrieval strategies in parallel rather than relying on a single embedding search.

FAISS — Semantic vector retrieval using local embeddings.

BM25 — Lexical retrieval for exact terms, identifiers, names, and textual matches embeddings may underweight.

Graph — Relationship-oriented retrieval through stored graph structure.

Phrase — Phrase-sensitive matching for exact or near-exact textual relationships.

Attribute — Metadata and attribute-based retrieval.

Temporal — Runs independently of the base retrieval path, identifying temporal intent and producing its own ranking signal without calling the base Fusion worker. The result merges with the base retrieval/ranking path through a final fusion stage — keeping temporal retrieval independently measurable and ablatable rather than folded into the core pipeline.


Ranking

Retrieval candidates are combined and ranked using multiple signals rather than vector similarity alone: semantic similarity, lexical relevance, graph relationships, phrase matches, attribute matches, temporal relevance, any additional registered signals, and reranking.

Individual signals can be evaluated independently and extended through plugins.


Blackboard and Scheduler

Memoria uses a blackboard-based execution model to coordinate independent retrieval and processing workers. Workers publish results to shared execution state while the scheduler controls task execution and completion policies.

This supports parallel execution, completion policies, time budgets, worker isolation, extensible scheduling, and deterministic orchestration.


Storage

Persistent memory storage is kept separate from retrieval implementation: relational storage, vector indexes, metadata, graph relationships, retrieval-specific indexes, and query/memory management. Retrieval workers can evolve without the underlying memory representation becoming tied to one strategy.


Integrations

Obsidian

A source-ingestion adapter for Obsidian vaults — markdown notes, frontmatter, wikilinks, note metadata, and structured ingestion into Memoria. Tested against real vault data and a dedicated adapter test suite.

MCP

Memoria includes a Model Context Protocol server:

memoria-mcp

Available operations: memory_search, memory_store, memory_store_many, memory_fetch, memory_update, memory_delete — letting MCP-compatible clients and agents use Memoria as an external memory system.


Plugin System

A configuration-driven plugin architecture spanning analysis, evaluation, feedback, ingestion, lifecycle, query, ranking, retrieval, routing, scheduler, and storage.

Plugins can register retrieval workers, ranking signals, rerankers, query processors, type/signal routers, storage backends, vector stores, database backends, extractors, entity recognizers, benchmark adapters, feedback recorders, and scheduler policies.

Loading supports automatic discovery, explicit activation, allowlists, denylists, runtime enable/disable, local plugins, and entry-point plugins. The denylist takes precedence when both allow and deny rules apply.

Plugin Generator

memory plugin create

Interactively creates a plugin module, optional tests, optional README documentation, and the corresponding hook configuration.


Interfaces

Python API · CLI · TUI · GUI · HTTP API · MCP

CLI

memory --help
memory info
memory store "Memoria stores this locally."
memory recall "What does Memoria store?"

Installation

PyPI

python -m pip install kitzkatz-memoria

From Source

git clone https://github.com/KitzKatz/Memoria.git
cd memoria/memoria
python -m pip install -e .

No external LLM or API key required for core memory and retrieval functionality.


Quick Start

memory store "Memoria is a local-first memory system."
memory recall "What is Memoria?"
memory info
memoria-mcp
memory plugin create

Configuration

Layered: built-in defaults → user configuration → environment variables.

Covers database, embeddings, retrieval, ranking, scheduling, temporal retrieval, plugins, interfaces, logging, and API services. Environment variables use the MEMORY_ prefix.

PLUGIN_ENABLED = true
PLUGIN_DIR = "plugins"
PLUGIN_AUTO_LOAD = true

PLUGIN_ENABLED_PLUGINS = []
PLUGIN_DISABLED_PLUGINS = []
export MEMORY_PLUGIN_AUTO_LOAD=true

See the configuration documentation for the complete settings reference.


Benchmarking

Memoria includes infrastructure for evaluating retrieval, latency, ranking, and memory-system behavior — LongMemEval, LoCoMo, synthetic workloads, benchmark analysis, result formatting, batch loading, memory extraction, and question extraction.

Full reproduction commands, dataset preparation, cache behavior, and analysis tooling are documented separately rather than duplicated here.


Performance

On the benchmark system described above, a representative query:

Stage Average
Query processing ~0.24 ms
Embedding ~79.8 ms
Retrieval ~101.9 ms
Scheduler wait ~35.2 ms
Database ~30.9 ms
Ranking ~0.28 ms
Response construction ~3.4 ms
Total ~208.5 ms

Configured retrieval deadline: ~125 ms. Actual latency varies with workload, cache state, memory size, candidate counts, and hardware — these numbers come from the Celeron N4020 / ~3.7 GiB RAM / CPU-only system above, not a GPU-equipped dev machine.


Project Structure

memoria/
├── blackboard/
├── benchmark/
│   ├── github/
│   └── obsidian/
├── cache/
├── core/
├── db/
├── graph/
├── ingestion/
├── memory/
├── memoria_mcp/
├── plugins/
├── ranking/
├── retrieval/
├── routing/
├── routes/
├── shared/
├── system/
├── tools/
├── cli.py
├── tui.py
├── demo.py
├── pyproject.toml
└── README.md

The repository also contains benchmark datasets, test infrastructure, development tooling, and documentation.


Requirements

  • Python, Linux recommended
  • ~4 GiB RAM recommended, CPU-only supported, GPU optional
  • No cloud service, no API key, LLM optional

Tested Hardware

Intel Celeron N4020 @ 1.10 GHz · ~3.7 GiB RAM · no GPU · Debian Linux

Despite the modest hardware, the full LongMemEval workload reached a measured peak RSS of 2.65 GiB, with an average individual query peaking around 580 MiB.


License

MIT License.


Memoria

Memory infrastructure first. Local. Composable. Measurable. Extensible. Independent of any particular LLM.

Metadata

Release files for kitzkatz-memoria 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for kitzkatz-memoria 1.0.0
File Size Uploaded
kitzkatz_memoria-1.0.0.tar.gz 253.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for kitzkatz-memoria 1.0.0
File Interpreter ABI Platform
kitzkatz_memoria-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 509.2 kB

Release files / kitzkatz_memoria-1.0.0.tar.gz

Download URL kitzkatz_memoria-1.0.0.tar.gz
Size 253.0 kB
Tags Source
SHA-256 checksum
How to use checksums
74d16824fed50b6c2610f7821f06b7b71d79b6926f7ba272e6d85474044fa589
BLAKE2b-256 checksum
How to use checksums
9c1723791fb3f2d895ae352e8e8ec843e3ae5d36af47aac35ec011befb41ac17
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release files / kitzkatz_memoria-1.0.0-py3-none-any.whl

Download URL kitzkatz_memoria-1.0.0-py3-none-any.whl
Size 256.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d290c120085b64bc00f77339ed6b748474ade33823ffef367d2042ac3faeadb6
BLAKE2b-256 checksum
How to use checksums
c5ee26ca7866bd0c5fabd242a8dad30896f1506970e509c97ab799ccc90dde96
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page