Skip to main content

mdcx

PyPI License DOI

Convert a document collection to verified Markdown, package it into a single encrypted file, and make it queryable by agents through the Model Context Protocol.

The problem

An agent answering questions about a document collection has two options. It can receive the documents in its context window, which is expensive and bounded by the window size. Or it can query a component that already knows where each item is.

Measuring one specific query — where the minimum pipe diameter to be modelled in 3D is stated — over a real collection of 99 documents and 180 MB, using the cl100k_base tokenizer:

Model tokens Local tokens
Reading the originals 2,265,488 2,265,327
Querying the package 435 2,688,861

The 435 comprise 20 for the question, 274 for the retrieved passage and 141 for the answer.

The first row costs the entire collection for a concrete reason: a PDF cannot be searched, it is a binary, and without prior conversion there is no way to know which of the 99 documents holds the answer. They all have to be extracted and read.

This is one measurement, not an average: the saving depends on how much text an answer requires. What does not vary is the shape of the change. The work does not disappear, it moves from the context window — which is billed and finite — to the CPU, which is not. That is why the local column rises rather than falls.

The three stages

Conversion. Each document is converted to Markdown and checked against the text the original actually exposes, read with a library independent from the engine that performed the conversion. Content the structured engine omits is appended verbatim rather than reported as lost.

Over the collection used during development — 99 documents, 1,144,553 reference words — 594 words were not recovered, a global coverage of 99.948%. Of the 184 documents exposing text, 116 came out at exactly 100% and none below 99.5%. The remaining four are scanned drawings containing no text at all in the file: they were read by optical character recognition and are marked as unverifiable, because no text original exists to measure them against.

Packaging. The corpus, its search index and the provenance of every passage fit into a single .mdcx file, encrypted with AES-256-GCM, whose header can be read without the key. From 8.8 MB of Markdown to 3.9 MB in one file.

Retrieval. A query returns the passages that answer it with their exact source. Over the 20 real queries used for tuning, the correct document appears within the top five results in 19 cases and within the top ten in all 20.

Installation

The package separates querying from conversion, because they have very different requirements.

Command Installs Size
pip install mdcx query and read .mdcx packages ~10 MB
pip install "mdcx[mcp]" the above plus the MCP server ~50 MB
pip install "mdcx[convert]" document conversion (Docling, PyTorch) ~1.4 GB
pip install "mdcx[all]" everything, including OCR ~1.5 GB

Conversion is what pulls in the heavy dependencies. Someone who receives an .mdcx file and only needs to query it installs neither Docling nor PyTorch.

Converting a collection

pip install "mdcx[convert]"
mdcx-convert --input ./Documents --output ./Documents_md

The output mirrors the input directory structure, adds a global index, and records for each file the coverage achieved against its original.

Packaging and querying

mdcx pack --output ./Documents_md --target corpus.mdcx --key "..."
mdcx info corpus.mdcx
mdcx search corpus.mdcx "where is the minimum diameter stated" --key "..."
mdcx export corpus.mdcx --target ./restored --key "..."

info reads the header without the key, so the issuer and the integrity of a file can be checked before opening it. export rebuilds the original folder: a format that cannot be left is a trap, however well intended.

Using it as an MCP server

The server requires Python and this package. It does not require the conversion stack, so the footprint is about 50 MB.

{
  "mcpServers": {
    "mdcx": {
      "command": "python",
      "args": ["-m", "mdcx.mcp_server"],
      "env": {
        "MDCX_FILE": "/path/to/corpus.mdcx",
        "MDCX_KEY": "package-key"
      }
    }
  }
}

Alternatively, with uv the server runs without a prior installation, which is the usual arrangement for Python MCP servers:

{
  "mcpServers": {
    "mdcx": {
      "command": "uvx",
      "args": ["--from", "mdcx[mcp]", "python", "-m", "mdcx.mcp_server"],
      "env": {
        "MDCX_FILE": "/path/to/corpus.mdcx",
        "MDCX_KEY": "package-key"
      }
    }
  }
}

Three tools are exposed. search returns the passages answering a question, each with its source document and portable path. info describes the corpus and the fidelity of its conversion. document returns a full document when passages are not enough.

The server verifies the package before it starts listening, so a wrong path or key is reported immediately rather than on the first query.

Tests

pip install pytest
python -m pytest tests/ -v

The suite covers hostile inputs: empty and corrupted files, names in other alphabets, malformed queries including SQL injection attempts, truncated and tampered packages, and compaction against content loss.

Language

Retrieval is lexical: a query matches words that appear in the documents. It therefore works in whatever language a corpus is written in — English, German, French, Spanish, Portuguese, Italian and any other the tokenizer segments — and it cannot cross between languages. A Spanish query finds nothing in an English corpus, however well indexed, because the words are not there.

The package records the predominant language of a corpus when it is built, and info reports it. When a query returns nothing and none of its terms appear in the index, the result says so and names the corpus language, so an empty answer can be told apart from material the corpus does not hold.

Earlier versions shipped a glossary of 83 Spanish terms and advertised Spanish queries against English documents. Those terms all came from piping engineering and project management; measured against a corpus of mathematics, biology and history the glossary contributed nothing, and the claim it backed failed on eleven of twelve queries. Removing it left retrieval unchanged even on the engineering corpus it was written for: 19 of 20 queries still find the right document in the top five results.

A glossary remains available for anyone who wants one, as a decision of theirs rather than an assumption of the package:

from mdcx import search
search.GLOSSARY = search.load_glossary("my-glossary.json")

The file maps a term to its equivalents: {"caneria": ["piping", "pipe"]}.

Paths

No output contains absolute paths. Every document is identified by a pseudopath beginning with @/, resolved against the folder or package containing it, so a corpus remains valid wherever it is stored: local disk, network share or cloud.

Signing

A package can be signed so that its issuer can be proven rather than merely declared. The signature covers the digest of the encrypted body, so it attests both origin and content, and is verified without the encryption key.

mdcx keygen
mdcx pack --output ./Documents_md --target corpus.mdcx --key "..." \
          --issuer "Acme Ltd" --signing-key <private-key>
mdcx verify corpus.mdcx --public-key <public-key>

Verification also requires the body to be intact: a signature covering only the recorded digest would otherwise accept a package whose contents had been replaced while its header was left untouched.

The issuer field alone is free text and proves nothing. Only a signature does.

Encryption

The package encrypts at rest and decrypts in memory when opened; nothing is written to disk in clear. This protects a file in transit. It is not the same as searching over encrypted data without ever decrypting it, which is a separate field with documented leakage attacks and per-query costs measured in seconds.

The key is derived with scrypt, which makes guessing slow: about 8 attempts per second, each requiring 32 MB of memory, which prevents parallelisation on a GPU. Even so, the real strength is the passphrase: a dictionary password falls in a day.

Authorship

Conceived and directed by Jorge Ellena G., programmed with the assistance of Claude (Anthropic).

Every decision in this package was made against measurements rather than convention: which conversion engine to use, which licence permits which, how to rank a search, which optimisations to accept and which to discard. Several were discarded precisely because they were measured — reducing the search candidate pool appeared to be ten times faster and in fact lowered accuracy from 19 to 17 out of 20 — and those measurements are recorded alongside the decisions they justify.

Citation

Archived on Zenodo with a permanent identifier. The concept DOI always resolves to the latest version:

https://doi.org/10.5281/zenodo.22015991

Licence

Apache 2.0. The software may be used, modified and sold, provided the copyright notice is retained.

PyMuPDF was deliberately avoided: its AGPL licence would require anyone using this software to publish their own under AGPL, including those offering it only as a network service.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mdcx-1.0.6.tar.gz (58.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mdcx-1.0.6-py3-none-any.whl (65.7 kB view details)

Uploaded Python 3

File details

Details for the file mdcx-1.0.6.tar.gz.

File metadata

  • Download URL: mdcx-1.0.6.tar.gz
  • Upload date:
  • Size: 58.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.10

File hashes

Hashes for mdcx-1.0.6.tar.gz
Algorithm Hash digest
SHA256 ca17a58a4bf059d730544ede903e79da0c55433d0d3d37fa9796acd714e1d5dd
MD5 446d0d1e9592b277fdd1339f45b2fdba
BLAKE2b-256 f4313a79cca93e3b62682fcd81c6b4d4266853888c965d5db0077847c1d1bc43

See more details on using hashes here.

File details

Details for the file mdcx-1.0.6-py3-none-any.whl.

File metadata

  • Download URL: mdcx-1.0.6-py3-none-any.whl
  • Upload date:
  • Size: 65.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.10

File hashes

Hashes for mdcx-1.0.6-py3-none-any.whl
Algorithm Hash digest
SHA256 11eb04e3496d8a9afe35135ec4b5eba7789396e91e485c47ca476701bb04717a
MD5 572c82de49fd4cf4c9803ee6c4575d11
BLAKE2b-256 11fbd8f397be802be25662e1e06876c49e4a6317bd912ca0f964845430da7318

See more details on using hashes here.

Release history Release notifications | RSS feed

1.22.0

2 files

1.21.1

2 files

1.21.0

2 files

1.20.0

2 files

1.19.0

2 files

1.18.0

2 files

1.17.0

2 files

1.16.0

2 files

1.15.1

2 files

1.15.0

2 files

1.14.0

2 files

1.13.0

2 files

1.12.0

2 files

1.11.0

2 files

1.10.0

2 files

1.9.0

2 files

1.8.0

2 files

1.7.2

2 files

1.7.1

2 files

1.7.0

2 files

1.6.9

2 files

1.6.8

2 files

1.6.7

2 files

1.6.6

2 files

1.6.5

2 files

1.6.4

2 files

1.6.3

2 files

1.6.2

2 files

1.6.1

2 files

1.6.0

2 files

1.5.1

2 files

1.5.0

2 files

1.4.0

2 files

1.3.5

2 files

1.3.4

2 files

1.3.3

2 files

1.3.2

2 files

1.3.1

2 files

1.3.0

2 files

1.2.2

2 files

1.2.1

2 files

1.2.0

2 files

1.1.0

2 files

1.0.7

2 files

This release

1.0.6 This release

2 files

1.0.5

2 files

1.0.4

2 files

1.0.3

2 files

1.0.2

2 files

1.0.1

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page