Skip to main content

viparse

Vietnamese-first document loader for RAG.

PyPI Python License

Tiếng Việt: README.vi.md · Website: viparse.trizenx.com

One command turns any Vietnamese document — including legacy TCVN3/VNI/VISCII fonts, scanned PDFs and page images, and old .doc/.xls/.ppt files — into clean Unicode NFC Markdown/JSON, ready to push into a vector DB.

Why

Generic loaders parse the file but often emit garbled diacritics (legacy fonts) or wrong Unicode normalization. viparse handles exactly that Vietnamese layer: detect & convert legacy encodings to Unicode, enforce NFC, and offer diacritic-aware OCR.

Backbone principle: never hand-write a parser. Wrap well-maintained engines behind thin adapters; if an engine gets a CVE or is abandoned, swap the adapter without touching the rest. Heavy dependencies are lazy-imported via extras (viparse[ocr], viparse[office]).

Where viparse fits

viparse is not a general-purpose document loader and does not try to replace one. Tools like Unstructured, LlamaParse and docling cover a far wider matrix of formats, layout analysis and table reconstruction; viparse covers one layer they are not built around — the Vietnamese text layer.

The concrete gap: a document authored in a pre-Unicode Vietnamese font stores Latin bytes that only render as Vietnamese when the original .VnTime/VNI font is applied. A faithful extractor returns those bytes faithfully — and faithfully wrong:

extracted   B¸o c¸o tµi chÝnh quý II n¨m 2026
correct     Báo cáo tài chính quý II năm 2026

Embeddings built on the first string retrieve nothing. viparse detects the legacy encoding, maps it to real Vietnamese letters and enforces NFC, so the text that reaches your vector DB is the text a human would read.

Use it as the loader for a Vietnamese-heavy corpus, or as a normalization pass over text another loader produced — that second shape is one call:

import viparse

docs = [viparse.fix(doc.page_content) for doc in some_other_loader.load()]

fix takes text, not a path, so it composes with whatever already read the file. Text that is already Unicode, and text that is not Vietnamese, come back unchanged.

On accuracy claims. viparse-corpus publishes 0.986 diacritic accuracy over 96 hand-transcribed Vietnamese government documents in five real formats, against 0.019 for the same reader with conversion switched off. Both rows are scored on the same 96 documents by the same published command, one flag apart. That is a claim about viparse, not a comparison: no other tool has been run against that corpus. A head-to-head would need their output on the same files, and until that exists this section makes no claim about anyone else. The corpus, the metric and every raw result are public so it can be argued with — including a written account of the ways the number is weaker than it looks.

Status

Released and published to PyPI — the version badge above tracks the current release. docs/specs/ holds the full spec map (SPEC-0 … SPEC-8) and CHANGELOG.md the release notes.

Installation

Requires Python 3.11+. The core install is pure stdlib — every parser and OCR binary lives behind an extra:

pip install viparse                # core — pure stdlib, no parser/OCR binaries
pip install "viparse[office]"      # .docx / .xlsx / .pptx and legacy .doc / .xls / .ppt
pip install "viparse[pdf]"         # digital PDFs
pip install "viparse[rtf]"         # RTF
pip install "viparse[ocr]"         # scanned PDFs and page images (needs Tesseract)
pip install "viparse[langchain]"   # LangChain document adapter
pip install "viparse[llamaindex]"  # LlamaIndex document adapter
pip install "viparse[mcp]"         # MCP server, for agents
pip install "viparse[all]"         # every engine and adapter

mcp is deliberately not in all: all is about parsing capability, and installing every format handler should not start pulling in a server runtime.

Run viparse doctor to see which engines your installed extras enable.

Usage

import viparse

docs = viparse.load("tai_lieu_cu.pdf")  # list[Document], already NFC
docs = viparse.load("bang_luong.xlsx", output="markdown", encoding="auto")

load() takes the knobs that matter per call — output (text / markdown / json), encoding (override detection), ocr, normalize (NFC by default), max_bytes, plus optional cache and chunk objects. load_batch() accepts the same options plus a workers count and yields one list[Document] per source, so a large corpus streams instead of materialising at once.

from viparse import load_batch
from viparse.cache import DiskCache

for docs in load_batch(paths, output="markdown", workers=8, cache=DiskCache(".viparse-cache")):
    index(docs)

Chunking runs on the document's block structure rather than flat text, so a chunk never straddles a section boundary, a table row is never split in half, and a chunk that continues a table repeats its header row. On PDF there are no headings to work with — see What it does not do.

from viparse.integrations.langchain import to_langchain_documents
from viparse.integrations.llamaindex import to_llamaindex_documents
viparse ./docs/**/*.pdf -o md
viparse doctor        # list available engines per installed extras

Using it from an agent

pip install "viparse[mcp]"
viparse-mcp                        # stdio; also `python -m viparse.mcp`

Claude Desktop, Claude Code and anything else that speaks MCP:

{ "mcpServers": { "viparse": { "command": "viparse-mcp" } } }

Four tools. repair_garbled_vietnamese takes a string, not a path, because most of the time the agent already has the broken text in context and there is no file to point at. identify_vietnamese_encoding names the encoding without changing anything, and returns a preview so its answer can be judged rather than trusted. read_vietnamese_document is viparse.load over a path. viparse_version is for bug reports.

The agent skill

skills/garbled-vietnamese-text/SKILL.md is a Claude-style skill covering the same ground for agents that do not have the MCP server: how to recognise each encoding, how to convert, and the traps — never convert text that is already Unicode, a font name is not proof, one document can be two encodings, and detection needs a phrase rather than a four-character fragment.

Copy it into .claude/skills/ (or your agent's equivalent) to use it.

tests/test_skill.py executes every conversion the document claims, reading the examples out of the Markdown table rather than duplicating them. A skill whose examples do not run is worse than no skill: an agent follows it, gets a wrong answer, and has no reason to doubt the instruction.

If you change the tool descriptions, keep the symptom in them

This is the one thing about src/viparse/mcp/server.py that is not obvious.

An agent never thinks "I should use viparse" — it has never heard of it. It encounters B¸o c¸o tµi chÝnh in a file it just read and needs something that recognises that. So the descriptions are written around the symptom: the mojibake itself, the font names (.VnTime, VNI-Times), the encoding names. A description that says "parses Vietnamese documents" is invisible to an agent that does not know the product; one containing tµi chÝnh is found by pattern-matching the broken text.

tests/test_mcp.py asserts the symptoms are present, because that property is easy to lose in an edit that is only trying to tighten the wording.

What it does not do

viparse is measured on ordinary Unicode documents too, not only on the legacy corpus — the structure benchmark plants labelled paragraphs, headings and tables in generated .docx / .xlsx / .pptx / PDF files and counts what comes back.

document order completeness headings
.docx, .xlsx, .pptx 1.000 1.000 1.000
one-column PDF 1.000 1.000 0.000
two-column PDF 0.600 1.000 0.000

Nothing is ever lost — completeness is 1.000 everywhere. Both failures are failures of arrangement, which is the harder kind to notice: the text is all there, fluent, and in the wrong order.

A PDF has no headings. viparse does not infer them from font size, so every title in a PDF arrives as an ordinary paragraph and every chunk from a PDF carries an empty section. Section-aware chunking is real on .docx, .xlsx and .pptx; on PDF it is splitting on size.

A multi-column PDF is read across the page, not down the columns — paragraph 1 is followed by paragraph 19. Recovering the columns means detecting them, which is layout analysis, and viparse does not do layout analysis. For multi-column PDFs use a layout-aware loader — Unstructured, docling, LlamaParse — and pass its output through viparse.fix(). That composition is the intended one: they know where the text is, viparse knows what the bytes mean.

Not attempted at all: figures and embedded images, formula recovery, reading order for rotated or freeform layouts, and handwriting.

OCR is measured. viparse[ocr] reads a scanned PDF and a page image (.png / .jpg / .tif, including multi-page TIFF). Against the corpus:

subject documents diacritic
real scans 3 0.973
rendered pages 96 0.990
conversion path, for comparison 96 0.986

Three real scans is a floor under the rendered figures, not a benchmark — all three are single pages, hand-transcribed from the image before OCR was run on them. The gap between 0.973 and 0.990 is roughly what a real page costs. That row moves as more are transcribed; the corpus carries the live figure.

The remaining errors are almost entirely tone marks, in both directions — a hook invented on bare i, the tone dropped from // — which is exactly what this product exists to preserve.

An earlier version of this section called OCR the weakest path here and quoted 0.967 / 0.898. Those numbers came from a defect in the corpus scorer, not from viparse, and are withdrawn; the corpus repository carries the full account.

Architecture

One pipeline, four layers, each behind a Protocol so implementations stay swappable and testable with fakes:

viparse.load("file")
    │
    ├─ route      detect format from magic bytes, pick engines by priority
    ├─ extract    Engine     → RawExtraction   (raw text + encoding/font signals)
    ├─ normalize  Normalizer → NormalizedDoc   (legacy → Unicode, NFC)
    └─ structure  Renderer   → Document        (text / markdown / json, + chunks)

Pipeline holds an EngineRegistry, a Normalizer and a Renderer, all injected — the orchestrator itself depends on no parsing library. The registry returns every engine matching a content type ordered by priority, and that ordered list is the fallback chain: the orchestrator walks it until one engine succeeds.

Module Role
detect.py Magic-byte format detection (zip/OOXML, %PDF, OLE2 for legacy .doc/.ppt)
registry.py Priority-ordered engine registry and fallback chain
engines/ Thin adapters — docx, xlsx, pptx, pdf, rtf, ocr, legacy
normalize/ The moat: detector, tcvn3, vni, viscii, vps, frequency, cleanup
structure/renderer.py Blocks → text / markdown (GFM tables) / versioned json
chunk.py Section-aware, table-row-atomic chunking
safety.py Size ceiling and zip-bomb guard, applied before any parser sees a file
cache.py Opt-in content-hash cache keyed by hash + options + schema version
pipeline.py The orchestrator, error policy and metrics hooks

When it hands text back unconverted, it says so

Two real cases return mojibake and nothing else would tell you. An RTF font table lists the fonts a document declares rather than the fonts applied to text, so the RTF engine emits no signal by design; and a PDF can embed subsetted fonts that expose no legacy name. In both, the output reads as Vietnamese-shaped nonsense — nothing errors, nothing is empty, the length is right — which is the hardest failure to notice.

So a document that comes back unconverted is scored the way encoding="auto" would have scored it, and if that would have found a legacy encoding you get a warning naming it and the fix:

text looks like tcvn3 and was returned unconverted;
pass encoding="tcvn3" or encoding="auto" to convert it

It never converts anything, and it reuses auto's guards — so a Spanish document does not get advice that would corrupt it.

How encoding detection decides

The primary signal is the font name the extraction engine carries out of the document: a .Vn* font implies TCVN3, a VNI* font implies VNI. That signal is high confidence, so it runs by default.

A content-frequency heuristic — trial-convert, then score against a Vietnamese character model — also ships, but is opt-in. A character model cannot reliably separate legacy Vietnamese from other diacritic-heavy Latin text, so running it unconditionally would corrupt documents it has no business touching. Text already in Unicode passes through untouched either way.

Contributing

See CONTRIBUTING.md. The workflow is deliberately narrow:

  • One task = one branch = one commit = one PR. Branch vip-<id>-<short-slug>, commit and PR title VIP-<id> <short imperative>.
  • main is protected: a PR must pass quality on Python 3.11 / 3.12 / 3.13, plus build and audit.
  • Keep a PR scoped to its task; unrelated work gets its own task.

Specs live in docs/specs/ — a change that alters behaviour should say which SPEC section it implements.

License

MIT © 2026 Đinh Minh Trí (Kayden)

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

viparse-0.1.28.tar.gz (198.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

viparse-0.1.28-py3-none-any.whl (106.1 kB view details)

Uploaded Python 3

File details

Details for the file viparse-0.1.28.tar.gz.

File metadata

  • Download URL: viparse-0.1.28.tar.gz
  • Upload date:
  • Size: 198.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for viparse-0.1.28.tar.gz
Algorithm Hash digest
SHA256 dab1b26e4f437d5f5ad0c17fba1c02c3b9fee2d5c4b9930e8a6115bf814b26d3
MD5 bdd60c1a9cccfaa976fea525db737d75
BLAKE2b-256 1f764ed698c14cc2ef6906e22fe1585cb82c519b8711ef46024f950276f00990

See more details on using hashes here.

Provenance

The following attestation bundles were made for viparse-0.1.28.tar.gz:

Publisher: publish.yml on TrizenX/viparse

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file viparse-0.1.28-py3-none-any.whl.

File metadata

  • Download URL: viparse-0.1.28-py3-none-any.whl
  • Upload date:
  • Size: 106.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for viparse-0.1.28-py3-none-any.whl
Algorithm Hash digest
SHA256 8c46d39476e719a39dad6d19b0c044bd3eac24b06bc021c4e97f27fca4075d23
MD5 07a76e2a521a2ac00ae37ec38407b584
BLAKE2b-256 5afdb4eb2c1203f001d3eacd4f4dc3f4aeaf9ec4824d0037165a2c2aa66b003d

See more details on using hashes here.

Provenance

The following attestation bundles were made for viparse-0.1.28-py3-none-any.whl:

Publisher: publish.yml on TrizenX/viparse

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.1.31

2 files

0.1.30

2 files

0.1.29

2 files

This release

0.1.28 This release

2 files

0.1.27

2 files

0.1.26

2 files

0.1.25

2 files

0.1.24

2 files

0.1.23

2 files

0.1.22

2 files

0.1.21

2 files

0.1.20

2 files

0.1.19

2 files

0.1.18

2 files

0.1.17

2 files

0.1.16

2 files

0.1.15

2 files

0.1.14

2 files

0.1.13

2 files

0.1.12

2 files

0.1.11

2 files

0.1.10

2 files

0.1.9

2 files

0.1.8

2 files

0.1.7

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page