Skip to main content

Vesma — memory & knowledge server for AI agents

Vesma

A memory & knowledge server for AI agents
named after the Vesper of memory, built for AI agents that need to remember

PyPI npm Python License: Apache-2.0 Version

🇬🇧 English · 🇷🇺 Русский

Quick start · Features · What it is · Connect a harness · Architecture · Docs


AI agents forget everything when a session ends. Vesma gives them a place to lay it down — structured, searchable, governed by contract — so what they learn does not vanish with the closing of a window.

  • Local-first. One process on your machine. SQLite + a bundled embedding model; nothing leaves the host, no API keys, works offline.
  • One server, any harness. VS Code Copilot, Claude Code, Cursor, OpenCode, Codex, Windsurf, ZCode, pi, Hermes — the same MCP wire, one line each.
  • The agent learns to use it. Not just tools: always-on instructions, a skill pack, and a memory-first prompt mode, deployed into your harness in one command.

🚀 Quick start

Three commands from an empty machine to an agent that remembers — and knows when to look.

1 · Install the server

pip install vesma

One package, everything included: the memory server, the vesma CLI, the REST API, and the MCP server your agent harness talks to. The embedding model ships inside — search works fully offline, no API keys, nothing downloaded.

⚠️ Names. The product and the CLI are vesma (pip install vesma, PyPI slot project/vesma). The pre-rebrand packages remain live until deprecation: pip install mnemos-memory-server installs the same server under the legacy name (its legacy CLI spelling was mnemos, now an alias). The bare pip install mnemos is an unrelated project — do not use it.

Or take the prebuilt image — Docker, Podman, or a Kubernetes cluster

The same server ships as a public image (ghcr.io/vesmaro/vesmaro) — pull and run: no build, no login. Full guides live in the admin docs.

☸️ Kubernetes / K3s — Helm chart with ingress
helm install vesmaro deploy/helm/vesmaro \
  --namespace vesmaro --create-namespace \
  --set auth.totpMasterKey="$(openssl rand -hex 32)" \
  --set ingress.className=traefik \
  --set 'ingress.hosts[0].host=vesma.example.com'
kubectl -n vesmaro rollout status deploy/vesmaro

K3s ships Traefik and local-path storage, so the commands work as-is — chart, values and TLS setup: Kubernetes deployment.

🐳 Docker / Docker Compose
cd deploy/docker
cp .env.example .env                    # put your TOTP_MASTER_KEY=$(openssl rand -hex 32) inside
docker compose up -d                    # podman-compose works too
curl -fsS http://localhost:8787/health  # → {"status":"ok"}

Full guide: container deployment.

🦭 Podman — systemd quadlet or kube play
# systemd user service (preferred for a long-running host)
# legacy-named asset — unit file stays mnemos.container until the deploy wave renames it
cp deploy/podman/quadlet/mnemos.container ~/.config/containers/systemd/
# add the TOTP key to ~/.vesmaro.env (both env spellings), then:
podman pull ghcr.io/vesmaro/vesmaro:4.3.0
# quadlet derives the unit name from the file name — the unit is mnemos.service for now
systemctl --user daemon-reload && systemctl --user start mnemos
curl -fsS http://localhost:8787/health

Both Podman recipes (quadlet + podman kube play): deploy/podman/README.md · full guide: container deployment.

2 · Connect your harness — and teach it to use memory

vesma integration setup

One pass: detects the agent harnesses on your machine, registers the Vesma MCP server in each supported one (VS Code Copilot, Cursor, ZCode, OpenCode, pi, Hermes, and everything reading the ~/.agents standard — Claude Code, Codex and friends), and deploys the behavioral pack — always-on instructions and memory skills, so the agent recalls at session start, checkpoints before its context gets compacted, and treats memory as a priority instead of forgetting the tools exist.

Running a harness that reads nothing standard? One paste block per harness: Connect Vesma to any harness.

3 · Verify — then try it

vesma doctor

PASS / WARN / FAIL per check: store, config, MCP transport, harness registration (--fix repairs the common warnings). Then give it a memory:

vesma add "First memory — Vesma remembers across sessions" \
  --tags project:vesma,agent:me,vesma:learning
vesma search "remembers across sessions"

That is the whole loop: write, find, never lose it — and the agent knows when to look.

📘 Want every detail? The extended guide covers all install variants (uv tool, pipx, CLI-only, external LLM extras, installer script, container), per-harness connection walkthroughs, configuration, and troubleshooting: Getting Started — the complete first run.


✨ Features

One local server — and a connected agent harness gets the full memory stack.

Area What you get
Universal connectivity MCP server (38 tools, stdio) + REST API — any MCP-capable harness connects in one line (tools · HTTP)
Ready integrations VS Code Copilot, Claude Code, Cursor, Codex, Windsurf, OpenCode, ZCode, pi, Hermes Agent — one-line MCP presets for all of them, native deploy targets for most, multi-harness doctor (vesma doctor)
Skill pack 14+ memory skills deployed into your harnesses
Flexible memory Hybrid search (full-text + vector, rank fusion) over the bundled offline model mnema-embed-v1, tag contract, per-agent / per-project memory, context-filter profiles, CCR compression — 70–90% token savings, originals kept
Context assembly assemble_context: search → compress → filter → secret scan → cache align → token budget, per-block provenance
Context bridge on_context_rewrite — when the harness compacts history, the lossless original stays available on demand
Lifecycle hooks pre_llm_call context injection, on_session_start, post_tool_call auto-compression of tool outputs
Publication v3.0.0 Entries visible immediately after save, background refinement with seamless swap, quarantine with neutral retraction
Self-protection Injection / secret detectors on input and publication, every output scanned, full per-entry audit
Auto-pipeline Background processor: clustering, deduplication, quality gate, publication. Entries awaiting refinement sit at pipeline_state=pending until the processor runs — in CLI-only deployments (no daemon) start it with vesma processor start; vesma doctor reports the pending-queue depth

Autonomy for an arbitrary harness and LLM-driven enrichment are partial — the full, honest map lives in docs/en/features.md.


🧩 What Vesma is

A single-tenant, local-first memory server for AI agents. One in-process core, three equivalent control surfaces, and a storage layer you can read with your own eyes.

Capability What it gives you
🔎 Hybrid search Vector similarity + SQLite FTS5 full-text over every memory
🧪 Knowledge pipeline raw → processing → processed → published lifecycle with a state machine
🧠 Per-agent recall A focused recall surface scoped to each agent's project context
⚙️ Policy engine Schedule and trigger automation over the memory store
🧹 Context filter Five-stage noise stripper for logs / stdout before anything hits a model
🗜️ Reversible compression (CCR) Compress large content with zero data loss — originals cached in SQLite, retrievable via hash marker
🧷 CacheAligner Relocate dynamic content (timestamps, UUIDs, session ids, tokens) to the tail so provider KV caches (Anthropic cache_control, OpenAI prefix caching) hit across requests
🪶 Output token reduction Optional verbosity / effort params on mnemos_add / mnemos_search / mnemos_recall_context (MCP tool names — unchanged wire contract) steer the caller's output style — backward compatible, defaults are a no-op
📂 Path-scoped rules Ingest project rules and apply them by file path
🗂️ Obsidian vault A markdown mirror humans can browse, edit, and grep

SQLite for metadata, a local numpy + SQLite vector index for recall, and an Obsidian-compatible vault for the humans in the loop.

Where this is going. The shipped, measured rung is stores and finds. The next rungs — collapse with checks (session → project → cross-project synthesis: automatic by default, but never unconditional in authority — every derivative traces back to its sources, nothing pins without an operator) and, later, builds understanding — open only as pre-registered experiments prove them (ADR-0025). The invariant on that road is zero silent losses: a fact is either retained, or its loss is visible. And the shape of the ambition is a nervous system, not a conductor — memory that surfaces the right thing at the right moment, never one that conducts the agent.


🤝 Connect any harness

Vesma works with every MCP-capable agent harness. Three integration levels — pick the strongest one your harness supports:

Harness Native deploy target One-line MCP preset Adapter template
VS Code Copilot copilot (+ prompts via generic-copilot) mcp-setup.sh ✓
Claude Code via agents preset ✓
Cursor cursor preset ✓
Codex via agents preset ✓
Windsurf — preset ✓
OpenCode — preset ✓
ZCode zcode — ✓
Any AGENTS.md-standard harness agents — ✓
pi pi (bridge extension, also on npm as pi-mnemos) preset ✓
Hermes Agent hermes (native in-process MemoryProvider plugin) — —
  • Native targets — vesma integration setup --target <name> deploys the behavioral pack and registers the MCP server in one pass (integration guide).
  • One-line presets — integrations/mcp-presets.md: every harness above, copy-paste ready.
  • Adapter template — integrations/adapter-template.md: Connect / Expose / Configure + acceptance checklist for any harness that speaks MCP stdio.
  • Hermes Agent runs Vesma in-process: pip install vesma in the Hermes environment, then vesma integration setup --target hermes (details).

The shared contract is the tag schema — project:<slug>, agent:<slug>, and at least one mnemos:<subtype> (tag namespace — unchanged wire contract) — that every memory entry must carry.


🏗️ Architecture

System diagram — clients → interfaces → core → storage
flowchart TB
    subgraph CLIENTS["Clients"]
        C1(["Agent harness\nstdio MCP"])
        C2(["CLI — vesma …"])
        C3(["HTTP API client"])
    end

    subgraph IFACE["Interface Layer"]
        MCP["mcp_server.py"]
        FAPI["api/main.py · FastAPI"]
        TYPER["cli/main.py · Typer"]
    end

    MGR(["MemoryManager\nmanager.py"])

    subgraph PROC["Processing Subsystems"]
        CF["Context Filter\nfilter/"]
        PP["Knowledge Pipeline\npipeline/"]
        RE["Recall Engine\nrecall/"]
        PE["Policy Engine\npolicy/"]
    end

    subgraph BG["Background Services"]
        WA["Watchers\nwatchers/"]
        AC["Auto-collect\nauto_collect.py"]
    end

    subgraph STORE["Storage Layer"]
        SQ[("SQLite\nFTS5 · traces · projects")]
        VS[("Vector Store\nnumpy + SQLite")]
        VLT[("Obsidian Vault\nmarkdown mirror")]
    end

    C1 -->|"stdio"| MCP
    C2 --> TYPER
    C3 --> FAPI
    MCP --> MGR
    TYPER --> MGR
    FAPI --> MGR
    MGR --> CF
    MGR --> PP
    MGR --> RE
    MGR --> SQ
    MGR --> VS
    MGR --> VLT
    CF -.->|"raw + clean"| SQ
    PP -->|"status transitions"| SQ
    PP -->|"published upsert"| VS
    RE -->|"FTS5 MATCH"| SQ
    RE -->|"cosine search"| VS
    PE -->|"schedule / trigger"| MGR
    WA -->|"file events"| MGR
    AC -.->|"checkpoint reminder"| MCP

A deeper walkthrough — data model, state machines, security boundaries, operational concerns — lives in architecture/overview.md.


🎛️ Three surfaces, one core

The same MemoryManager powers all three interfaces. Pick the one that fits your client.

Surface Use it when… Reference
MCP — vesma mcp-server You are an agent harness — the path every connected agent takes mcp-tools.md
CLI — vesma … You live in a shell, want fast ad-hoc add / search, or are scripting cron jobs cli-reference.md
HTTP — vesma serve You have a non-MCP client — a web dashboard, a mobile app, a CI runner http-api.md

The HTTP surface also exposes the A2A Sessions API — a persistent backend for multi-step agent conversations that survive restarts. See a2a-sessions.md.


📚 Documentation

Page What it covers
docs/README.md Documentation landing — language picker (EN / RU)
getting-started.md First run: install → first memory → first search → connect your harness
mcp-presets.md Connect Vesma to any harness — one-line MCP presets (VS Code, Claude Code, Cursor, OpenCode, Codex, Windsurf, pi, Hermes)
integration-guide.md The behavioral pack: instructions, skills, prompt mode, deploy targets, agent wiring, Hermes plugin
features.md What works out of the box, what is partial, what is planned
architecture/overview.md System shape, data model, state machines, security boundaries
cli-reference.md Every vesma subcommand with flags, defaults, examples
mcp-tools.md Every mnemos_* tool exposed to agent harnesses (tool names — unchanged MCP wire contract)
http-api.md Every HTTP endpoint (memory CRUD, workflow, hooks, A2A Sessions)
tag-contract.md The project: / agent: / mnemos: tag schema (namespace — unchanged data contract) enforced on every memory
security.md Threat model, SSRF guard, FTS5 escape, auth model
kubernetes-deployment.md Helm chart for K8s/K3s clusters: ingress, storage, TLS, TOTP secret
contrib/node-install/ vesmaro-node — one-command bare-metal node bundle: venv + mesh binary + units, with adopt / atomic pair upgrade / uninstall
runbooks/ Install, migrate, backup / restore, dependency updates, container deployment
adr/ Architectural decision records — the why behind the design
CHANGELOG.md Release notes — Keep a Changelog format
CONTRIBUTING.md Development setup, git workflow, quality gate

📖 The lore

In Hesiod's Theogony, Mnemosyne (Μνημοσύνη) is the Titaness of memory — she who, by Zeus, gave birth to the nine Muses and through them made the world's remembering possible. Her name is the root of mnemonic, and she is what every singer, poet, and philosopher prays to before they begin.

The lore section below preserves the memory of the pre-rebrand name — the product itself now sails under Vesma, carrying the same task: to make remembering possible for the things that think. AI agents, unmoored from any single conversation, lose everything that came before. Vesma gives them a place to lay it down — structured, searchable, governed by contract — so that what they learn does not vanish with the closing of a session. The Muses, after all, were not for the gods' benefit. They were for the songs.


⚖️ License & contributing

Apache-2.0 — see LICENSE and NOTICE. Source: github.com/vesmaro/vesmaro.

Contributions are welcome — CONTRIBUTING.md has the development setup, the branch and commit conventions, and the quality gate a change must pass.


🛠 CI / Релизы — кластерный конвейер release-pipeline

Этот проект подключён к общему кластерному конвейеру релизов (Korrnals/release-pipeline, K3s abyss-ai-agent, namespace release-pipeline). GitHub Actions не используется (биллинг аккаунта заблокирован) — конвейер и есть штатный путь релизов. Релизный артефакт подписывается трёхслойно: SHA256 → SBOM → cosign → GPG.

Релиз новой версии:

  1. VERSION → релизный коммит (конвенция репо) → тег vX.Y.Z → push.
  2. Обновить версию проекта в projects[] файла ~/.cache/release-pipeline-values.yaml и применить:
    helm upgrade --install release-pipeline \
      ~/LABs/Projects/Project-Umbra/release-pipeline/chart/release-pipeline \
      --kube-context abyss-ai-agent -n release-pipeline \
      -f ~/.cache/release-pipeline-values.yaml
    
  3. Наблюдение: kubectl --context abyss-ai-agent -n release-pipeline get jobs, логи: kubectl ... logs -f job/release-<proj>-<ver>.

Подробности (типы проектов, kaniko-контейнеры, teardown): README конвейера.

Metadata

Release files for vesma 5.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vesma 5.1.0
File Size Uploaded
vesma-5.1.0.tar.gz 22.8 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for vesma 5.1.0
File Interpreter ABI Platform
vesma-5.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 43.9 MB

Release files / vesma-5.1.0.tar.gz

Download URL vesma-5.1.0.tar.gz
Size 22.8 MB
Tags Source
SHA-256 checksum
How to use checksums
542795aeecf5d7409f1b4a13512f9691a081a0c59a8bba8da8f7177719057824
BLAKE2b-256 checksum
How to use checksums
8ef94a5abfed454a0c8131aeabf83f9960c7e58b3737b35d5cd30fb35690b7aa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / vesma-5.1.0-py3-none-any.whl

Download URL vesma-5.1.0-py3-none-any.whl
Size 21.0 MB
Tags Python 3
SHA-256 checksum
How to use checksums
ffdc6f08edf06bde5c1ee5f2f6a43d3cd42fd77a429067d3f9023cf2987e4c08
BLAKE2b-256 checksum
How to use checksums
941e8b1ba06c6d00f8c6edebd1214e1474870e992fe3f52f69c9b172b1a661bc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.10 {"installer":{"name":"uv","version":"0.12.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

5.2.0

2 release files

5.1.2

2 release files

5.1.1

2 release files

This release

5.1.0 This release

2 release files

5.0.0

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page