Datacron
Local MCP server to query and maintain a Markdown vault from Claude, Codex, Gemini, or another stdio MCP client, without sending the whole vault into the context.
English | Français
What can you do with Datacron?
Recover project context, prepare for a conversation, and keep track of commitments. Datacron gives your assistant durable memory in readable, editable Markdown. Your notes remain usable independently of the client you choose.
Offline library: build a browsable Markdown library for Obsidian or another local reader, with a home page and topic indexes. Review proposed rewrites, splits and archives before changing the source vault. See the offline library guide and the 2026.0913.02 release notes.
| Need | Example request to your assistant |
|---|---|
| Resume a project | "Where did we leave off? Find the decisions and next actions." |
| Prepare a meeting | "Summarize our recent conversations and open points, with sources." |
| Remember a person | "Who is this person, how have we interacted, and what should we follow up on?" |
| Track objectives | "Find the commitments and achievements relevant to my next review." |
| Preserve a reliable record | "Save this decision, link it to the project, and verify that it was stored." |
The assistant orchestrates these requests using the available tools and granted permissions. A shared protocol guides reading, people updates, and write verification. Ambiguous identities require clarification; storing a deadline does not schedule a reminder. Explore daily follow-up.
Start here: install · first session · user guide · MCP reference · privacy.
Installation
Windows: one double-click installer
The easiest way on Windows: download Datacron-Setup.exe from the
latest Release, double-click it,
and pick your vault. No Python, no terminal, no administrator rights; Datacron registers
itself with your AI clients automatically. Full guide:
Windows installation.
Python: from PyPI
python -m pip install datacron
datacron setup
From source
From a clone of the repository:
python -m pip install -e ".[dev]"
Or, to install only the application:
python -m pip install -e .
Runtime prerequisites:
- Python 3.11+
ripgrepavailable on thePATHforsearch_regex; without it that tool falls back to a slower scan of indexed chunk bodies- a folder of Markdown notes
- a supported stdio MCP client, such as Claude Desktop, Codex CLI, or Gemini CLI
First session
- Choose your notes folder with the installer or
datacron setup. - Reconnect Datacron in your MCP client to load the tools and instructions. The Claude Desktop chat does not present the server instructions: paste the session start line printed by setup into your Claude preferences (see setup).
- Ask: "Find the notes for my project and summarize its status with sources."
For memory sessions, session_context returns bounded context and the shared protocol.
prepare_follow_up prepares sourced updates; existing writers apply them according to
permissions. get_follow_up retrieves the latest structured revisions. Existing prose notes
remain readable and are not automatically converted.
The server operates locally. Your client may send returned excerpts to its model provider; see privacy and security.
For cached session contracts, archive ranking, resumable write tracking and conversation evaluation, see daily workflow improvements.
Quick start
The easy path - one command detects your AI clients, initializes the vault, indexes it, and registers Datacron everywhere:
datacron setup # interactive; add --yes for all defaults
See the installation guide for options (--client, --scope, writing,
durability). Or step by step:
datacron init /path/to/vault
datacron index --vault /path/to/vault
datacron status --vault /path/to/vault
datacron mcp install --client claude-desktop --vault /path/to/vault
The mcp install subcommand above is dedicated to Claude Desktop. For Codex CLI, Gemini CLI,
Antigravity, LM Studio, Cursor, and the other clients, use multi-client setup with
datacron setup --client <identifier> or auto-detection with --client all.
Add to LM Studio
LM Studio 0.3.17+ has one user configuration and no project scope. The preferred command is:
datacron setup --yes --vault "VAULT_PATH" --client lmstudio --scope user
For a Python installation where datacron-mcp is on PATH, the equivalent read-only
configuration can also be imported with this official deeplink:
The link imports this example. Open LM Studio's MCP editor and replace both
<YOUR_VAULT> placeholders before starting the server:
{
"mcpServers": {
"datacron": {
"command": "datacron-mcp",
"args": [],
"env": {
"DATACRON_VAULT_ROOT": "<YOUR_VAULT>",
"DATACRON_READ_PATHS": "<YOUR_VAULT>",
"DATACRON_DURABILITY": "best-effort"
}
}
}
}
The example does not enable write tools. CLI setup is safer for packaged installations because it writes the actual executable path automatically.
Restart the configured client or clients after installation.
To run the server manually:
datacron mcp serve --vault /path/to/vault
The direct script entry used by the installer is also available:
datacron-mcp
datacron-mcp reads the vault from DATACRON_VAULT_ROOT.
Configuration
datacron init creates .datacron/VAULT.yaml. That file can carry vault-local
configuration, notably query expansion:
query_expansion:
supervision: [monitoring]
sauvegarde: [backup]
restauration: [restore]
chiffrement: [encryption]
sécurité: [security]
validité: [validity]
certificat: [certificate]
Useful environment variables:
| Variable | Default | Role |
|---|---|---|
DATACRON_VAULT_ROOT |
unset | fallback after --vault; the current directory is accepted only when it contains .datacron/VAULT.yaml |
DATACRON_READ_PATHS |
empty | read allowlist; client setup sets it to the vault |
DATACRON_WRITE_PATHS |
empty | write allowlist; empty = write tools disabled |
DATACRON_MAX_RESULT_COUNT |
20 |
maximum number of results returned |
DATACRON_MAX_RESULT_TOKENS |
8000 |
token budget for search results |
DATACRON_REPAIR_MIN_INTERVAL_SECONDS |
30 |
minimum interval between repair-on-read sweeps; 0 = every read |
DATACRON_GET_NOTE_MAX_TOKENS |
25000 |
budget for get_note(format="full") |
DATACRON_SESSION_CONTEXT_SECTIONS |
{} (no section selected) |
JSON mapping of note paths to heading paths for bounded orientation excerpts |
DATACRON_CHUNK_MAX_TOKENS |
1024 |
target maximum chunk size |
DATACRON_RIPGREP_PATH |
rg |
ripgrep binary |
Path lists use the OS separator (: on Unix, ; on Windows).
Writing
Writes are deliberately OFF by default. Without DATACRON_WRITE_PATHS, write tools return a
clear error and create no file.
To enable writing to a specific subfolder:
$env:DATACRON_VAULT_ROOT = "C:\Notes"
$env:DATACRON_READ_PATHS = "C:\Notes"
$env:DATACRON_WRITE_PATHS = "C:\Notes\_memory"
datacron mcp serve --vault C:\Notes
datacron setup can also apply the allowlist machine-wide (user environment
variable, opt-in) so every MCP client inherits it; default: _memory, _drafts,
_journal. See the setup guide.
Available write tools:
create_note_ai: creates a typed Markdown note, without overwrite.append_journal: adds an entry under a heading of an existing note.set_frontmatter: updates lifecycle fields, therejectedoptions list, and the monotonelast_idcounter without modifying the Markdown body.patch_note_preamble: replaces or removes the Markdown preamble before the first recognized Markdown heading (ATX or Setext), with mandatory CAS control.patch_note_section: replaces the content under an existing heading with CAS control.delete_note_section: explicitly deletes an H2-H6 section (ATX or Setext) and its subtree.rename_note_section: renames only the title of an H2-H6 section (ATX or Setext).move_note_section: previews or commits an exact H2-H6 subtree move within a note, with mandatory CAS.revert_note: restores the exact bytes of a version kept in history.apply_organization_manifest: validates and then applies a local content-addressed bundle after confirmation bound to the exact admitted organization pre-state.
Guarantees:
- strict note confinement within
DATACRON_WRITE_PATHS; organization-batch note sources and targets must also stay inside the unchanged liveorganization.scopeand pass the live note-admission policy, including exclusions - two internal exact-CAS targets for an organization batch:
.datacron/VAULT.yaml, only to change the top-levelorganizationmapping without changingorganization.scope, and.datacron/ulids.json, only when Datacron derives the key migration required by amove_replace_exact - atomic overwrite via temporary file +
os.replace - content-addressed history before modifying an existing note
- synchronous
reconcile()after a normal write; immediate searchability is guaranteed only when reconciliation succeeds - local audit log
- for an organization manifest: crash-consistent recovery and atomic replacement of each file; simultaneous visibility across several paths is not guaranteed
Concurrent multi-machine mode is not supported for writes: keep a single-writer rule on the vault.
For apply_organization_manifest, also stop every other Datacron client and server during the
maintenance window. Before applying, keep a verified byte-exact backup outside the vault of the
affected notes and the complete .datacron directory until every post-commit check is green. Call
mode="validate" first, review the bounded hashes it returns, then reuse
the exact confirmation_token with mode="apply". The token binds the manifest and payloads, all
admitted Markdown notes inside organization.scope, the exact vault configuration and identity
sidecars, and the projected report. It deliberately does not bind unrelated note bytes outside
organization.scope. A change to any authenticated component invalidates the confirmation before
mutation. history_mode=full is required at validation time. If Datacron derives identity-sidecar
case-collision cleanup, also review identity_sidecar_case_canonicalization_count and its
content-free SHA-256 before applying; both proofs are token-bound and retained in the durable
receipt.
An existing replace_exact or move_replace_exact source must carry its id in frontmatter; an
identity available only from the sidecar is unsupported by this v1 schema. If the batch is already
durably committed but index reconciliation or the planner oracle fails, the response says so
explicitly (committed_index_incomplete or committed_report_mismatch) and the same call can be
retried with the same token.
An organization-batch blocker is reported by datacron ops inspect with a pending_batch_ reason
and both single-note repair actions unavailable; use the full offline rollback procedure in the
operational-health guide rather than repairing or quarantining one member.
Available capabilities
Datacron indexes a folder of Markdown notes, exposes a local MCP server, then returns the
relevant notes or chunks to the client instead of a full dump. The vault stays an ordinary
Markdown folder: Datacron only adds a .datacron/ sidecar for the index, logs, internal
ULIDs, history, and the operation journal. The one exception is datacron setup at project
scope, which is part of the default: it also writes each detected client's project config
into the vault root, such as .mcp.json, .cursor/mcp.json, .gemini/settings.json,
.agents/mcp_config.json, .codex/config.toml or .vscode/mcp.json. Those files carry
machine-local absolute paths, so a synced vault carries them to every machine. Pass
--scope user to keep the vault free of them.
| Surface | Current state |
|---|---|
| Vault reading | list_notes, get_note, resources datacron://vault/map, vault/info, policy/active |
| Search | SQLite FTS5/BM25, FR↔EN query expansion, temporal re-rank, ripgrep via search_regex |
| Local graph | Wikilinks and backlinks via get_backlinks |
| Writing | 9 confined note tools + 1 organization batch, journaled and disabled by default without DATACRON_WRITE_PATHS |
| MCP transport | Python MCP SDK v2 through MCPServer, local stdio only; modern 2026-07-28 protocol and legacy 2025-11-25 compatibility, with no HTTP listener |
| Index | datacron index incremental, datacron reindex full, conditional repair on read |
| Organization | Optional organization block in VAULT.yaml; datacron reorganize --dry-run measures the gap read-only, apply_organization_manifest applies |
| Evaluation | datacron eval over the real MCP pipeline: recall@k, MRR, nDCG, freshness, latency, and payload tokens |
| Guided setup | datacron setup: init + index + MCP registration in one command |
| Clients | Auto-detect and register via datacron setup --client all: Claude Desktop, Claude Code, Cursor, Gemini CLI, Antigravity, LM Studio, Codex CLI, Windsurf, VS Code |
| Daily memory | session_context, prepare_follow_up, get_follow_up: bounded context, sourced follow-up, and structured state |
| Memory protocol | Shared versioned server/client contract; protocol status checks distribution, not model behavior |
| Distribution | Windows installer (Datacron-Setup.exe), standalone executable (PyInstaller) with no Python required, or installation from source |
MCP Tools
Reading
| Tool | Description |
|---|---|
session_context |
Bounded session context and versioned common protocol. |
prepare_follow_up |
Prepare sourced follow-up plans without writing. |
get_follow_up |
Latest structured follow-up revisions with snapshot-bound pagination. |
list_notes |
returns a paginated list, filterable by folder, tags, and frontmatter key/value pairs, with ULID, title, tags, aliases, and dates |
get_note |
reads a note or an exact heading subtree, with pagination, chunk lookup, or a heading outline |
search_text |
runs a BM25 search on the FTS5 index with ranked snippets and stale notes demoted by default |
search_regex |
runs a regex search via ripgrep and resolves the found lines to indexed chunks |
get_backlinks |
returns chunks whose wikilinks target a ULID or a resolved alias |
Writing
| Tool | Description |
|---|---|
create_note_ai |
creates a new typed _memory note, confined to allowed paths, without overwrite and with a durable journal |
append_journal |
adds a Markdown entry under a heading, with confinement, exact history, and atomic write |
set_frontmatter |
updates allowed lifecycle fields, rejected, the monotone last_id counter, and updated, preserving the Markdown body |
patch_note_preamble |
replaces or removes the preamble before the first recognized Markdown heading (ATX or Setext), with mandatory CAS and suffix preservation |
patch_note_section |
replaces the content of an existing heading with CAS, exact history, and preservation of other sections |
move_note_section |
previews or commits an exact subtree move beneath an existing heading in the same note |
delete_note_section |
explicitly deletes an H2-H6 section (ATX or Setext) and its subtree, with optional CAS and exact history |
rename_note_section |
renames the title of an H2-H6 section (ATX or Setext) without modifying its content or subtree |
revert_note |
restores a note from its content-addressed history; the operation stays durable, reversible, and audited |
apply_organization_manifest |
validates a local content-addressed bundle containing at least one exact note operation and/or an exact organization configuration replacement, then applies its declared members and any required derived ULID-sidecar migration under CAS; application is journaled and crash-consistent |
Operational
| Tool | Description |
|---|---|
get_health |
returns the real state of index freshness, integrity, checksum, durability, and invariants |
get_write_progress |
Inspect multi-note write receipts, conflicts and current indexing without retrying writes. |
get_note_history |
lists the committed operation metadata of a note without reading historical content or modifying the journal |
audit_query |
queries operation metadata by period, tool, or note without modifying the journal or the vault |
Advisory (experimental)
| Tool | Description |
|---|---|
contradiction_scan |
live, deterministic, bounded scan of contradictions/refinements between sections; proposes and confirms an explicit CAS call read-only, without ever writing automatically |
MCP resources:
datacron://vault/mapdatacron://vault/infodatacron://policy/active
Search
search_text combines several signals:
- FTS5/BM25 for the base lexical score
- a heavier weight on the note title and heading trail than on the chunk body, so a note about a subject outranks a note that merely mentions it
- FR↔EN query expansion configured in
VAULT.yaml - conservative temporal re-rank:
- a note referenced in another note's
supersedesis strongly demoted confidence: lowandconfidence: needs_verificationapply a light penaltyinclude_superseded=truebrings historical notes back up
- a note referenced in another note's
- optional scope:
folder,tags, andfrontmatternarrow the searched notes with the same semantics aslist_notes; the response echoes the filters actually applied - optional grouping:
group_by_note=truekeeps the best chunk of each note and reports how many of its chunks matched, which cuts the returned tokens by about 40 percent on the eval corpus
search_regex stays literal: it applies neither query expansion nor temporal re-rank.
Historical search measurements - July 17, 2026
These measurements cover 19 questions and one configuration. They are not a benchmark of the current release or a guarantee for another vault.
Local measurement of the tool/impl pipeline actually received by the agent, 19 questions,
8k-token / 20-result configuration, July 17, 2026:
recall@5 0.89
recall@10 0.95
recall@20 0.95
MRR 0.73
nDCG@10 0.79
latency p50 57 ms
latency p95 276 ms
payload tokens 90567
On this historical set, tool-level recall@5 matched the BM25 store. Use datacron eval
with a suitable question set to measure behavior on your own notes.
Privacy and security
- Datacron does no telemetry.
- Datacron calls no cloud LLM.
- The MCP client, for example Claude, Codex, or Gemini, may send the chunks that Datacron returns to its provider. Datacron does not send it the full vault.
- Content returned to clients is wrapped in
<vault_content>...</vault_content>. - Results are bounded by count and by token budget.
- Filesystem access is confined by
DATACRON_READ_PATHSandDATACRON_WRITE_PATHS. - MCP operations are audited in the local logs.
CLI commands
datacron setup # guided path: init + index + client config
datacron setup --yes # all defaults, no prompts
datacron setup --client all --scope both --vault /path/to/vault
datacron setup --protocol # also install client memory rules
datacron protocol install --client all
datacron protocol status --client all --scope user
datacron init /path/to/vault
datacron status --vault /path/to/vault
datacron index --vault /path/to/vault
datacron reindex --vault /path/to/vault
datacron scrub-init --vault /path/to/vault
datacron scrub --vault /path/to/vault
datacron reorganize --vault /path/to/vault --dry-run # measure organization, read-only
datacron reorganize --vault /path/to/vault --dry-run --json # stable machine-readable report
datacron eval --questions examples/eval-questions.example.yaml --vault /path/to/vault
datacron eval --questions local/golden.yaml --vault /path/to/vault --save-baseline
datacron eval --questions local/golden.yaml --vault /path/to/vault --compare --json
datacron mcp serve --vault /path/to/vault
datacron mcp install --client claude-desktop --vault /path/to/vault # Claude Desktop only
datacron unregister --client all --scope both --vault /path/to/vault
datacron protocol uninstall --client all
Current limitations
- Lexical search only: no vector search or embeddings.
- No autonomous agent: the MCP client orchestrates.
- No GUI.
- No concurrent multi-machine writes.
- Client detection in
datacron setupis best-effort (a config directory or a binary on thePATH); an install in a non-standard location may be missed and can then be configured by hand.
Documentation
Full index: docs/en/index.md | Index français.
To get started:
- Installation and configuration guide
- Use Datacron with Ollama
- Frequently asked questions
- User guide
- Offline library and note consolidation
- Daily memory, people, and commitments
Technical references:
- Vault conventions (SPEC)
- Vault organization
- Read and reorganize note sections
- Architecture and public surface
- Security boundary
- Integrity scrubber
- Operational health and durability
- Freshness contract
Development
CI runs the invariants and the entire regression suite on Linux/Python 3.12 for changes limited to the READMEs, CHANGELOG, and Markdown pages under docs/fr/ or docs/en/. All other changes retain the six Linux/Windows and Python 3.11-3.13 combinations. Publications require the full matrix, as do empty or unverifiable diffs. ShellCheck, the dependency audit, and the required Quality gate remain active in both paths. The first push of a new branch also uses the full matrix because no previous comparison point is available.
python -m pip install -e ".[dev]"
ruff check .
ruff format --check .
mypy
pytest
License
Copyright 2026 Julien Bombled.
Licensed under the Apache License, Version 2.0.
Release files for datacron 2026.924.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datacron-2026.924.0.tar.gz | 437.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datacron-2026.924.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 911.1 kB
Release files / datacron-2026.924.0.tar.gz
| Download URL | datacron-2026.924.0.tar.gz |
|---|---|
| Size | 437.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fafeac9137b133f7bc1abd0b8803ef0b473af123bb3f2c57e38834cfe6a696c8
|
|
BLAKE2b-256 checksum How to use checksums |
f0915b0ba63732166e8315381443cc461641a167d52ad97f8335d8e0a02a8765
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency logRelease files / datacron-2026.924.0-py3-none-any.whl
| Download URL | datacron-2026.924.0-py3-none-any.whl |
|---|---|
| Size | 473.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3705232bfd4c7ca0099d70a55a9e7d8d216916df783bdeb15b1b01510b17f06f
|
|
BLAKE2b-256 checksum How to use checksums |
db7a87aa0e8e379a3f60c8fbf5e07ead7905422e6ba45e41d0d1fd7dcb83b4d0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency log