Skip to main content

viparse

Vietnamese-first document loader for RAG.

PyPI Python License

Website: viparse.trizenx.com

One command turns any Vietnamese document — including legacy TCVN3/VNI/VISCII fonts, scanned PDFs, and old .doc/.xls files — into clean Unicode NFC Markdown/JSON, ready to push into a vector DB.

Why

Generic loaders parse the file but often emit garbled diacritics (legacy fonts) or wrong Unicode normalization. viparse handles exactly that Vietnamese layer: detect & convert legacy encodings to Unicode, enforce NFC, and offer diacritic-aware OCR.

Backbone principle: never hand-write a parser. Wrap well-maintained engines behind thin adapters; if an engine gets a CVE or is abandoned, swap the adapter without touching the rest. Heavy dependencies are lazy-imported via extras (viparse[ocr], viparse[office]).

Where viparse fits

viparse is not a general-purpose document loader and does not try to replace one. Tools like Unstructured, LlamaParse and docling cover a far wider matrix of formats, layout analysis and table reconstruction; viparse covers one layer they are not built around — the Vietnamese text layer.

The concrete gap: a document authored in a pre-Unicode Vietnamese font stores Latin bytes that only render as Vietnamese when the original .VnTime/VNI font is applied. A faithful extractor returns those bytes faithfully — and faithfully wrong:

extracted   B¸o c¸o tµi chÝnh quý II n¨m 2026
correct     Báo cáo tài chính quý II năm 2026

Embeddings built on the first string retrieve nothing. viparse detects the legacy encoding, maps it to real Vietnamese letters and enforces NFC, so the text that reaches your vector DB is the text a human would read.

Use it as the loader for a Vietnamese-heavy corpus, or as a normalization pass over text another loader produced. A published head-to-head benchmark on diacritic accuracy is planned for v0.2 — until it exists, this section deliberately makes no accuracy claims against those tools.

Status

Released and published to PyPI — the version badge above tracks the current release. docs/specs/ holds the full spec map (SPEC-0 … SPEC-8) and CHANGELOG.md the release notes.

Installation

Requires Python 3.11+. The core install is pure stdlib — every parser and OCR binary lives behind an extra:

pip install viparse                # core — pure stdlib, no parser/OCR binaries
pip install "viparse[office]"      # .docx / .xlsx and legacy .doc / .xls
pip install "viparse[pdf]"         # digital PDFs
pip install "viparse[rtf]"         # RTF
pip install "viparse[ocr]"         # scanned PDFs (needs the Tesseract binary)
pip install "viparse[langchain]"   # LangChain document adapter
pip install "viparse[llamaindex]"  # LlamaIndex document adapter
pip install "viparse[all]"         # every engine and adapter

Run viparse doctor to see which engines your installed extras enable.

Usage

import viparse

docs = viparse.load("tai_lieu_cu.pdf")  # list[Document], already NFC
docs = viparse.load("bang_luong.xlsx", output="markdown", encoding="auto")

load() takes the knobs that matter per call — output (text / markdown / json), encoding (override detection), ocr, normalize (NFC by default), max_bytes, plus optional cache and chunk objects. load_batch() accepts the same options plus a workers count and yields one list[Document] per source, so a large corpus streams instead of materialising at once.

from viparse import load_batch
from viparse.cache import DiskCache

for docs in load_batch(paths, output="markdown", workers=8, cache=DiskCache(".viparse-cache")):
    index(docs)

Chunking runs on the document's block structure rather than flat text, so a chunk never straddles a section boundary and a table row is never split in half.

from viparse.integrations.langchain import to_langchain_documents
from viparse.integrations.llamaindex import to_llamaindex_documents
viparse ./docs/**/*.pdf -o md
viparse doctor        # list available engines per installed extras

Architecture

One pipeline, four layers, each behind a Protocol so implementations stay swappable and testable with fakes:

viparse.load("file")
    │
    ├─ route      detect format from magic bytes, pick engines by priority
    ├─ extract    Engine     → RawExtraction   (raw text + encoding/font signals)
    ├─ normalize  Normalizer → NormalizedDoc   (legacy → Unicode, NFC)
    └─ structure  Renderer   → Document        (text / markdown / json, + chunks)

Pipeline holds an EngineRegistry, a Normalizer and a Renderer, all injected — the orchestrator itself depends on no parsing library. The registry returns every engine matching a content type ordered by priority, and that ordered list is the fallback chain: the orchestrator walks it until one engine succeeds.

Module Role
detect.py Magic-byte format detection (zip/OOXML, %PDF, OLE2 for legacy .doc)
registry.py Priority-ordered engine registry and fallback chain
engines/ Thin adapters — docx, xlsx, pdf, rtf, ocr, legacy
normalize/ The moat: detector, tcvn3, vni, viscii, vps, frequency, cleanup
structure/renderer.py Blocks → text / markdown (GFM tables) / versioned json
chunk.py Section-aware, table-row-atomic chunking
safety.py Size ceiling and zip-bomb guard, applied before any parser sees a file
cache.py Opt-in content-hash cache keyed by hash + options + schema version
pipeline.py The orchestrator, error policy and metrics hooks

How encoding detection decides

The primary signal is the font name the extraction engine carries out of the document: a .Vn* font implies TCVN3, a VNI* font implies VNI. That signal is high confidence, so it runs by default.

A content-frequency heuristic — trial-convert, then score against a Vietnamese character model — also ships, but is opt-in. A character model cannot reliably separate legacy Vietnamese from other diacritic-heavy Latin text, so running it unconditionally would corrupt documents it has no business touching. Text already in Unicode passes through untouched either way.

Contributing

See CONTRIBUTING.md. The workflow is deliberately narrow:

  • One task = one branch = one commit = one PR. Branch vip-<id>-<short-slug>, commit and PR title VIP-<id> <short imperative>.
  • main is protected: a PR must pass quality on Python 3.11 / 3.12 / 3.13, plus build and audit.
  • Keep a PR scoped to its task; unrelated work gets its own task.

Specs live in docs/specs/ — a change that alters behaviour should say which SPEC section it implements.

License

MIT © 2026 Đinh Minh Trí (Kayden)

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

viparse-0.1.9.tar.gz (133.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

viparse-0.1.9-py3-none-any.whl (81.3 kB view details)

Uploaded Python 3

File details

Details for the file viparse-0.1.9.tar.gz.

File metadata

  • Download URL: viparse-0.1.9.tar.gz
  • Upload date:
  • Size: 133.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for viparse-0.1.9.tar.gz
Algorithm Hash digest
SHA256 6271dbe92139bfe437d8622fecaf5ef0983f4cd45f6c497ea675a342794efeb5
MD5 ed9a8c8e5201e4216784e7ac26f21a9d
BLAKE2b-256 e38c3af1aced749c479501f02bd4a180e0e3a3324c2dd4bcc99da77bbd9cd5a7

See more details on using hashes here.

Provenance

The following attestation bundles were made for viparse-0.1.9.tar.gz:

Publisher: publish.yml on TrizenX/viparse

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file viparse-0.1.9-py3-none-any.whl.

File metadata

  • Download URL: viparse-0.1.9-py3-none-any.whl
  • Upload date:
  • Size: 81.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for viparse-0.1.9-py3-none-any.whl
Algorithm Hash digest
SHA256 46a9c6bd7a5c9ae8c7eb7b5f99a2a80717a26de60c537b138ae2dc8c8dd4c71b
MD5 45c4c8350d254371ce707a58b7d6c0a6
BLAKE2b-256 5200335b4eb0b16eea34c7cc256d6e228dcf8fe0bac916e139c95c6607813d76

See more details on using hashes here.

Provenance

The following attestation bundles were made for viparse-0.1.9-py3-none-any.whl:

Publisher: publish.yml on TrizenX/viparse

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.1.31

2 files

0.1.30

2 files

0.1.29

2 files

0.1.28

2 files

0.1.27

2 files

0.1.26

2 files

0.1.25

2 files

0.1.24

2 files

0.1.23

2 files

0.1.22

2 files

0.1.21

2 files

0.1.20

2 files

0.1.19

2 files

0.1.18

2 files

0.1.17

2 files

0.1.16

2 files

0.1.15

2 files

0.1.14

2 files

0.1.13

2 files

0.1.12

2 files

0.1.11

2 files

0.1.10

2 files

This release

0.1.9 This release

2 files

0.1.8

2 files

0.1.7

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page