Privacy-preserving, performant document extraction (text + metadata) for the Knovas Semantix platform. Python reference implementation.
Project description
knovas-extract
Privacy-preserving, performant document extraction (text + metadata) for the Knovas Semantix platform. Python reference implementation of the cross-language knovas-extract spec.
Status: alpha (0.1.0.dev). API stable in shape, not yet in version. Not yet on PyPI.
What it does
Give it a document (PDF, DOCX, MSG, EML, HTML, RTF, MD, TXT). Get back a ExtractionResult: extracted text, document metadata (title, author, language, dates, page count), per-page text (when paginated), and heading-derived sections.
That's the whole library. It does not chunk, embed, upload to Semantix, or talk to any network. Those concerns belong in your application code.
Why
- Spec-first: the output contract is shared across every language implementation in the
knovas-extract-*family. Your Python output and a future Node/Go/Rust output are guaranteed equivalent (within documented tolerances). - Fast: uses the best native library per format (PyMuPDF for PDF, selectolax for HTML, …). PDF throughput ≈ 20–100 pages/sec on a single thread; lazy-imports keep cold start ≈ 50 ms.
- Safe by default: no network calls, no embedded-code execution, no path traversal, ZIP-bomb caps, XXE-hardened XML, typed errors only. See
SECURITY.mdfor the full posture.
Install
pip install knovas-extract # core only (TXT/MD/HTML/EML)
pip install 'knovas-extract[pdf]' # + PyMuPDF (AGPL — read NOTICE)
pip install 'knovas-extract[docx,msg,rtf]' # + DOCX/MSG/RTF
pip install 'knovas-extract[markdown]' # + emit_markdown=True (sanitized MD)
pip install 'knovas-extract[sentences]' # + emit_sentences=True (pysbd, MIT)
pip install 'knovas-extract[all]' # everything
License-sensitive embedders: install knovas-extract[minimal] for the permissive-only subset (no AGPL / GPL deps). Calling extract() on a format that needs an unavailable backend raises DependencyMissingError with the exact pip install command to fix it.
Quickstart
from knovas_extract import extract
result = extract("report.pdf")
print(result.content.text) # full extracted text, canonicalized
print(result.metadata.title) # e.g. "Q4 Earnings"
print(result.metadata.page_count) # e.g. 12
for page in result.content.pages: # per-page text (when paginated)
print(page.index, page.text[:80])
print(result.warnings) # e.g. ["page 7: unrecognized font"]
Bytes work too:
data = open("report.pdf", "rb").read()
result = extract(data, mime="application/pdf")
ExtractionResult.to_dict() round-trips through the spec's JSON Schema; pass it directly to anything expecting the contract shape.
Markdown output (opt-in, sanitized)
Pass emit_markdown=True to populate content.markdown with a whole-document Markdown rendering. Requires the [markdown] extra (adds markdownify); PDFs additionally require pymupdf4llm, which ships in the [pdf] extra.
result = extract("report.docx", emit_markdown=True)
print(result.content.markdown) # ATX headings, - bullets, **bold**, tables
Emitted Markdown is sanitized before conversion: <script>, <style>, <iframe>, <object>, <embed>, <applet>, <svg>, <math>, HTML comments, and event-handler attributes are stripped with their contents; <a href> URLs are gated by an allowlist of http, https, mailto, tel; <img> is emitted as alt-text only so downstream renderers cannot beacon on render. See SECURITY.md → Markdown emission for the full contract. RTF, and formats without recoverable structure, leave content.markdown = None with a warning explaining why.
Sentences with line + page + section citations (opt-in)
Pass emit_sentences=True to populate content.sentences with a
deterministic tokenization. Every Sentence carries an exact char
slice, a 1-based line window (into content.text), an optional page
number, and a back-pointer to the enclosing Section. Requires the
[sentences] extra (adds pysbd, MIT).
result = extract("report.pdf", emit_sentences=True)
for s in result.content.sentences:
print(f"[p{s.page_number or '-'}:L{s.line_start}] {s.text}")
# Exact retrieval:
assert result.content.text[s.char_start : s.char_end] == s.text
Guaranteed contracts (asserted on every result — never "best effort"):
content.text[s.char_start : s.char_end] == s.text— byte-for-byte exact retrieval.line_start/line_endare 1-based indexes intocontent.text; the window[line_start, line_end]containss.textas a substring.- Sentences appear in reading order; non-overlapping;
indexis monotonic 0-based. - When
content.pagesis present (PDF), every sentence has a non-nullpage_indexandpage_number(withpage_number == page_index + 1). When absent, both areNone— never fabricated. - When
content.sectionsis present,section_indexpoints to the innermost (most specific) enclosing section, orNonefor sentences before the first heading. - Same bytes → same sentences, byte-identical. No timestamps, no randomness.
For the full contract reference (including chunk-and-cite recipes and per-format fidelity), see docs/citations.md. For a complete walked-through example of what the library returns for a PDF, see docs/example-pdf-output.md.
Source path — where did this document come from?
Pass path= (or use a path-like input) and it flows to Source.path
verbatim so downstream citations can identify the source:
# Automatic — from a path-like input:
r = extract("reports/q4.pdf")
r.source.path # "reports/q4.pdf"
# Explicit — for bytes input:
data = open("reports/q4.pdf", "rb").read()
r = extract(data, mime="application/pdf", path="reports/q4.pdf")
r.source.path # "reports/q4.pdf"
Source.path is caller-supplied metadata — knovas-extract never
opens, canonicalizes, or resolves it. It is validated to reject
inputs that would corrupt logs / terminals: NUL bytes, ASCII control
characters (\n, \r, ANSI escapes), Unicode bidi-override characters
(CVE-2021-42574 Trojan Source), and lengths over
Limits.max_path_length (default 4096). Rejections raise ValueError
with a class-of-violation message that never contains the offending
payload — re-logging the exception is safe.
Full reference including per-format metadata gap-fill: docs/metadata-and-paths.md.
CLI
knovas-extract reports/q4.pdf --emit-sentences --emit-markdown --pretty
Emits the full ExtractionResult as JSON to stdout. --emit-sentences
and --emit-markdown are independent opt-ins.
Resource limits
Every extraction is bounded. Override the defaults per call:
from knovas_extract import extract, Limits
result = extract(
"huge.docx",
limits=Limits(
max_input_bytes=50 * 1024 * 1024, # 50 MiB cap
max_pages=1_000,
max_decompression_ratio=50,
max_text_bytes=10 * 1024 * 1024,
max_sentences=10_000, # explicit DoS cap on sentence array
max_metadata_value_length=1024, # per-scalar cap in Metadata.extra
max_xmp_bytes=256 * 1024, # PDF XMP metadata size cap
max_path_length=1024, # Source.path length cap
),
)
When a limit is crossed, you get a ResourceExhaustedError with .what / .limit / .observed attributes. Defaults (in Limits()) are conservative; tune them with your throughput budget in mind.
Errors
Every call either returns an ExtractionResult or raises a subclass of ExtractError:
| Exception | When |
|---|---|
UnsupportedFormatError |
MIME not registered (you can register a custom extractor via knovas_extract.dispatch.MIME_REGISTRY). |
CorruptDocumentError |
Bytes couldn't be parsed as the claimed format. |
EncryptedDocumentError |
Password-protected document, no password supplied. |
ResourceExhaustedError |
A Limits threshold was crossed. |
DependencyMissingError |
An optional extra isn't installed — exception tells you the exact install command. |
Plus one stdlib exception: ValueError from Source.path validation
(caller misuse, distinct from document corruption). See
docs/metadata-and-paths.md for the full policy.
No bare exceptions, no None, no Optional[ExtractionResult].
Security promises (enforced by CI)
- Never makes a network call. Asserted across every test via
pytest-socket. - Never executes embedded code. PDF JavaScript, DOCX macros, RTF object linking — all stripped, warning emitted.
- Never writes outside an explicit tmpdir. ZIP-slip paths are rejected.
- XML parsing is XXE-hardened. All XML goes through
defusedxml. - Releases are Sigstore-signed + ship SLSA L3 provenance. Verify before installing in production — see
RELEASING.md.
For untrusted inputs, run inside a sandbox. Copy-paste recipes for nsjail, bubblewrap, and rootless Docker in docs/sandboxing.md.
Spec conformance
This implementation conforms to spec_version = 1.2.0 of knovas/KnowledgeBase/clients/extraction/spec. The pinned spec sha is recorded in tests/spec/ (Git submodule). Every release runs the spec's golden corpus + adversarial corpus before tagging.
To run the golden tests locally against a sibling KnowledgeBase checkout:
export KNOVAS_EXTRACT_SPEC_DIR=/path/to/KnowledgeBase/clients/extraction/spec
hatch -e golden run run
Development
hatch env create # one-time
hatch run test # unit tests
hatch -e golden run run # corpus contract tests
hatch -e property run run # hypothesis robustness
hatch -e bench run run # benchmarks
hatch -e lint run all # ruff + mypy + pyright
hatch -e sec run all # bandit + pip-audit
CI runs the full matrix (3 Pythons × 3 OSes × every gate) on every PR.
Reporting vulnerabilities
See SECURITY.md. Please do not open public issues for security reports.
License
Apache-2.0 for knovas-extract itself. Several optional extras pull AGPL / GPL libraries — see NOTICE for the full third-party license inventory.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file knovas_extract-0.2.0.tar.gz.
File metadata
- Download URL: knovas_extract-0.2.0.tar.gz
- Upload date:
- Size: 126.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
be0f208d27ca25c98b07b028b92047cf0600bae3b9e986c8b228b12953d07d67
|
|
| MD5 |
f3784ab410b2b22dee3ec9af9d66b9c8
|
|
| BLAKE2b-256 |
178b8f6717bb6bf76df9742cad5757f1f22870f699486472ce15531426808b3c
|
Provenance
The following attestation bundles were made for knovas_extract-0.2.0.tar.gz:
Publisher:
release.yml on Seifeddini/knovas-extract-python
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
knovas_extract-0.2.0.tar.gz -
Subject digest:
be0f208d27ca25c98b07b028b92047cf0600bae3b9e986c8b228b12953d07d67 - Sigstore transparency entry: 2223930245
- Sigstore integration time:
-
Permalink:
Seifeddini/knovas-extract-python@258f47e58261070f6f9e8606d4cc6d0572b9f86d -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/Seifeddini
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@258f47e58261070f6f9e8606d4cc6d0572b9f86d -
Trigger Event:
push
-
Statement type:
File details
Details for the file knovas_extract-0.2.0-py3-none-any.whl.
File metadata
- Download URL: knovas_extract-0.2.0-py3-none-any.whl
- Upload date:
- Size: 77.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2a019e379e9aa366af5517fd9256f70013cb0e02f540f25b4172d81b6f37e44c
|
|
| MD5 |
b80288c17c9ae10050f40dc864185636
|
|
| BLAKE2b-256 |
bf719857fa5b4a76816e879241cd037a7b009ae15686aee388f22e48a7b865bf
|
Provenance
The following attestation bundles were made for knovas_extract-0.2.0-py3-none-any.whl:
Publisher:
release.yml on Seifeddini/knovas-extract-python
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
knovas_extract-0.2.0-py3-none-any.whl -
Subject digest:
2a019e379e9aa366af5517fd9256f70013cb0e02f540f25b4172d81b6f37e44c - Sigstore transparency entry: 2223930704
- Sigstore integration time:
-
Permalink:
Seifeddini/knovas-extract-python@258f47e58261070f6f9e8606d4cc6d0572b9f86d -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/Seifeddini
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@258f47e58261070f6f9e8606d4cc6d0572b9f86d -
Trigger Event:
push
-
Statement type: