Skip to main content

Knovaryn

Knovaryn — open-source, MCP-native training-data foundry

Turn permitted PDFs and documents into traceable, quality-gated SFT, DPO/preference, KTO, and evaluation datasets that any MCP-capable agent can build, review, and export — every example traced to its source.

Python License MCP Offline-first, no API keys

Documentation site →

Knovaryn — from documents to defensible training data


Status

Self-hosted open-source alpha (0.2.1). You run it on your own machine or infrastructure; there is no hosted/managed service. The core pipeline, provenance, quality gates, durable jobs, release integrity, and the CLI / REST / web console are stable within the 0.x line (breaking changes documented in the changelog); the MCP tool surface and export formats are experimental and may still be reshaped before 1.0. The bundled demo runs entirely offline on a deterministic fake provider; live model generation and certified judge profiles are explicit opt-ins that need real provider credentials. Publishing is never automatic — always an explicit, gated action.


1. Install

One command, no API keys, no network needed to get started:

pip install knovaryn

From source (for development):

git clone https://github.com/waalwalker1/knovaryn.git
cd knovaryn
uv sync --dev

Optional extras are opt-in: docling, docetl, litellm, s3, parquet, hub, mcp, ml. See the configuration reference.

2. 90-second offline demo

The demo runs the complete pipeline on bundled sample documents with a deterministic fake provider — no API keys, no network:

uv run knovaryn demo --examples 20 --json
uv run knovaryn doctor        # environment + storage health

In about 90 seconds you get a release bundle: dataset card, per-file manifest, quality/source/license/privacy reports, lineage, and a detached checksum. The full command lifecycle (project → source → run → review → dataset) is in the CLI reference.

3. Real-document quickstart

Point Knovaryn at documents you are permitted to use and drive the pipeline through the MCP tools (or REST / SDK):

  1. knovaryn_create_project — create a project.
  2. knovaryn_add_source — add permitted documents with declared licenses.
  3. knovaryn_estimate_run — dry-run cost estimate.
  4. knovaryn_start_pipeline — run SFT / preference / KTO / evaluation.
  5. knovaryn_validate_dataset — quality gates quarantine failures.
  6. knovaryn_create_dataset_version + knovaryn_export_dataset — ship it.

See PDF → SFT dataset, build DPO preference data, and grounded QA datasets.

4. MCP quickstart

Knovaryn is an MCP training-data server: a Model Context Protocol server any MCP-capable agent can drive. The authoritative tool catalogue is generated from the server registration — see MCP tools. Supported MCP SDK range: mcp>=1.28,<3 (an acceptance matrix exercises both the 1.x and 2.x lines over stdio and authenticated streamable HTTP).

pip install "knovaryn[mcp]"
knovaryn-mcp                     # stdio (default MCP host transport)
# or remote over Streamable-HTTP:
knovaryn-mcp --transport streamable-http --host 127.0.0.1 --port 8000

Connect your agent and ask it to build a dataset — see MCP training-data server and MCP clients.

5. What it produces

A Knovaryn release is an immutable, versioned bundle:

  • trainer-ready records in JSONL / Parquet and framework layouts (trl_sft, trl_preference, kto, sharegpt, alpaca, openai_chat, huggingface_layout, evaluation);
  • a dataset card, per-file manifest, quality/source/license/privacy reports;
  • a detached checksum + reproducible bundle, verified with knovaryn verify-release.

6. Traceability (a concrete example)

Every exported row carries source_document_ids, source_span_ids, and a content_hash. You can walk any example backward:

TrainingExample → chunk → SourceSpan → ParsedDocument → SourceDocument
                                                          → location + precision

Every span carries a machine-verifiable location precision, derived from what the parser actually recorded — never asserted by callers: exact_bbox (page + bounding boxes, Docling-backed documents), exact_page, page_range, section, chunk, or unknown. Markdown and plain-text sources honestly report section/chunk granularity instead of fabricating page numbers; lineage (API /v1/projects/{id}/examples/{eid}/lineage, MCP knovaryn_lineage) returns the per-span precision so consumers can verify provenance claims.

The export provenance gate resolves every document/span reference to an existing record in the same project and recomputes the content hash. A row that cannot be resolved blocks the export — Knovaryn never returns a "successful" export it cannot defend. Export content manifests record the derived precision of every cited span.

7. Quality and policy gates

Candidates are scored against a dated acceptance policy and pass fail-closed gates — failure means quarantine, never silent export:

  • schema · grounding · completeness/answerability · format
  • refusal · duplicate · contamination (train/val/test leakage)
  • semantic consistency — deterministic contradiction checks (offline-fast, default), optionally extended by a cited-evidence-only LLM judge (certified-semantic; deterministic contradictions are final, an unavailable judge yields unverified, never verified)
  • preference signal (DPO/KTO pairs) and information gain
  • privacy · license

Human review (knovaryn_review_example) records an approve/reject decision as a new immutable revision — it never mutates an example in place. The gate list is generated from validator registration: quality gates; the validation-profile registry is documented in validation profiles.

8. Supported stack

Concern Supported
Inputs PDF, Markdown, office/documents (via Docling), local files, archives; URL ingestion opt-in
Topologies SFT, DPO/preference, KTO, evaluation (grounded QA)
Exporters Generated list in the exporter reference: JSONL, Parquet, TRL, ShareGPT, Alpaca, OpenAI chat, Hugging Face layout, evaluation
Providers Offline deterministic fake provider; LiteLLM gateway (OpenAI-/Anthropic-/DeepSeek-compatible)
Deployment Local (SQLite + filesystem); team (PostgreSQL + S3-compatible); Docker / Compose / Kubernetes
Interfaces CLI, MCP server, REST + web console, Python SDK

9. Architecture

One application-services core behind a durable pipeline engine, exposed through four interfaces:

Source → Parse → Split → Chunk → Generate → Validate → Version → Export → Publish

Everything runs as durable jobs (leases, heartbeats, checkpoints, idempotency, budgets) so a crash resumes instead of redoing paid work. Diagrams (system architecture, pipeline flow, durable jobs, MCP session, security, value):

Architecture at a glance →

10. Security & privacy

  • Offline-first — no credentials required for the demo, tests, or first run.
  • Secrets from the environment only — never hard-coded, never logged, redacted at the display boundary.
  • Secure intake — path-traversal/symlink protection, verified archives, URL ingestion off by default, SSRF defenses, loopback HTTP binding, no shell-command MCP tools.
  • Publication is dry-run by default and gated on license approval — nothing is pushed anywhere without explicit action. See license & privacy.

11. Benchmarks (methodology & limitations)

Reproducible benchmark suites cover parse fidelity, lineage resolution, generation validity, groundedness, duplicate/leakage rate, pipeline throughput, crash recovery, and export compatibility. Semantic-judge quality is measured as false-accept / false-reject rates with Wilson confidence intervals against an adversarial corpus. Honest framing: offline fake-provider throughput measures framework overhead only — it is not synthetic-data generation throughput or model quality. Numbers are versioned with their methodology in benchmarks/ — reproduce them before relying on them.

12. Limitations

  • Alpha software: APIs, storage layout, and the MCP surface may change before 1.0; do not build unmanaged long-lived dependencies on 0.x internals.
  • The offline demo uses a deterministic fake provider — its output demonstrates the pipeline mechanics, not generation quality. Real datasets require real providers and your own review effort.
  • Location precision depends on what the parser records: plain-text/markdown sources cannot yield page or bounding-box precision, and spans are reported at the honest lower granularity instead.
  • Deterministic semantic checks catch specific contradiction classes only; they cannot prove entailment. Judge-based verification requires configured, reachable providers and inherits their limitations; unavailable judges fail closed to unverified rather than guessing.
  • Publication gating enforces license/privacy policy inside Knovaryn; it is not legal clearance. You are responsible for the rights to your sources.
  • Single-node SQLite mode is local-first and not multi-writer; team scale expects PostgreSQL plus object storage.

13. Comparison — when to choose Knovaryn

Compared factually with related tools (Synthetic Data Kit, Distilabel, Docling, Easy Dataset) in the peer landscape and comparisons: choose Knovaryn when you want a document-grounded, provenance-enforced, gate-and-export foundry driven over MCP — not just a parser or a composition SDK.

14. Documentation

15. Contributing & community

Please read CONTRIBUTING.md, the code of conduct, and SECURITY.md. See GOVERNANCE.md, ROADMAP.md, and CHANGELOG.md.

16. License & citation

Apache-2.0 for original code (see LICENSE). Dataset licensing is kept separate from code licensing and governed by the source-license registry. Third-party licenses: LICENSES-THIRD-PARTY.md.

If you use Knovaryn in research, please cite it — see CITATION.cff.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

knovaryn-0.2.1.tar.gz (1.5 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

knovaryn-0.2.1-py3-none-any.whl (271.8 kB view details)

Uploaded Python 3

File details

Details for the file knovaryn-0.2.1.tar.gz.

File metadata

  • Download URL: knovaryn-0.2.1.tar.gz
  • Upload date:
  • Size: 1.5 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for knovaryn-0.2.1.tar.gz
Algorithm Hash digest
SHA256 711fcbc3838cce7dea562212f93c913203b90b984d1541f486305cca9ace491f
MD5 a8a07f3e45574f730f7cfee1ccd460c9
BLAKE2b-256 05712749cd7fe128c355259fe961152a3f5ba30f5d156a56305176ce2b4b6a9f

See more details on using hashes here.

Provenance

The following attestation bundles were made for knovaryn-0.2.1.tar.gz:

Publisher: publish.yml on waalwalker1/knovaryn

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file knovaryn-0.2.1-py3-none-any.whl.

File metadata

  • Download URL: knovaryn-0.2.1-py3-none-any.whl
  • Upload date:
  • Size: 271.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for knovaryn-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 d64e6e6372fb9c3d2443d420769871e35e33449bb56fdb113e927016c7c9aa53
MD5 896f0f0205bb12e374fdbca13cde229d
BLAKE2b-256 d7bd7e4af8e8d777c670ade78abb719baeeaf011109a2804bcb42583121155ac

See more details on using hashes here.

Provenance

The following attestation bundles were made for knovaryn-0.2.1-py3-none-any.whl:

Publisher: publish.yml on waalwalker1/knovaryn

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.2.2

2 files

This release

0.2.1 This release

2 files

0.2.0

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page