Skip to main content

vdi2770

Read a VDI 2770 handover-documentation container and get back a typed model — without extracting anything to disk, without opening a socket, and without importing anything outside the standard library.

That last part is a property of the package, not a promise: a test reads every import in it and fails on anything the standard library does not provide. It is what makes this worth installing on its own — the rule set that judges what this reads is a separate distribution, released beside it under the same version, and nothing here reaches it.

pip install vdi2770
import vdi2770

box = vdi2770.read_container_file("handover.zip")
for c in box.walk():
    if c.metadata_bytes is None:
        continue
    doc = vdi2770.build_document(vdi2770.parse_xml(c.metadata_bytes), c.where)
    print(c.path, [(i.domain_id, i.id) for i in doc.identifiers],
          [k.class_id for k in doc.classifications])
handover.zip [('SUPPLIER', 'DOC-2024-0001')] ['03-01']
handover.zip!/pumps.zip [('SUPPLIER', 'DOC-2024-0002')] ['02-04']

It decides nothing

There is no is_valid() here, on purpose. Whether a container is correct is a question about VDI 2770, and the answer depends on which supplement your customer sent you. This library tells you what is in the file and where it is written; the opinion is yours to supply.

If you want an opinion supplied for you, the rules live here too, behind an extra: pip install "vdi2770[validate]" adds the schema parser and the rule set, and you run it as python -m vdi2770.validate check YOUR-CONTAINER.zip (that install carries no command). vdi2770-validate is the older name for the same thing and installs the vdi2770-validate command; it asks for this package at exactly its own version, so the two are always the release they say they are. A floor was used for four releases and stopped only the direction that never went wrong: it let a newer engine install beside an older command, and halves that disagree about which release they are refuse to judge rather than guess.

Three properties, each tested rather than promised

Nothing is extracted to disk. Members are decompressed into memory under a budget and dropped. There is no temporary directory to clean up and no path traversal to get wrong, because no path is ever joined.

Nothing is fetched. No socket is opened for any input, ever — including XML that asks for one. An entity declaration is refused outright rather than resolved-but-locally, so there is no parser setting to get wrong later.

A refusal is reported, not raised. A member that blows a budget becomes a Defect on the container and the read continues, so one hostile file inside a supplier archive does not cost you the other four hundred.

What comes back

read_container(data, name) returns a Container:

path handover.zip!/pumps.zip — the JAR convention, so it stays greppable
kind DOCUMENTATION, DOCUMENT, UNKNOWN, or UNREADABLE
members, file_names what the reader can open — the budget filter and the readability sweep have both run
present every file name the archive declares, including members that were refused. Whether a name is there is a fact about the directory; being unable to inflate the bytes behind it does not unsay it
metadata_bytes, metadata_name the metadata that was found, if any
children, walk() inner containers, opened to three levels
defects what the reader could not do, and why
rejected members present in the archive but refused, and why
near_misses reserved name → (kind, the member name that nearly matched, as the archive spells it), kind being in-a-subfolder, path-prefixed, case-differs or case-differs-elsewhere. vdi2770_metadata.xml in an archive with no metadata is worth saying; how to say it is yours, not ours
duplicate_names a ZIP may carry the same name twice; readers disagree about which one wins

build_document(node, where) returns a Document whose every node carries a Location with the line and column it was written at, which is the reason this package parses XML itself instead of handing you an ElementTree.

read_pdf(data) returns four facts and no verdict: is_pdf, header, encrypted, and pdfa_claim — the last being what the file's own metadata claims, such as "2b". Nothing here verifies that claim. Verifying PDF/A takes a PDF/A validator, and this is not one.

is_pdf has three values, not two. True found an indirect object, False looked at the whole file and found none, and None gave up at MAX_OBJ_PROBES without an answer. A conforming file can reach that bound — a comment is legal between any two tokens — so None is not "no", and code that needs the difference compares with is True / is False rather than testing truthiness. Folding the third value into either of the other two decides, in a layer that knows nothing about what the file is for, a question only the caller holding the obligation can answer.

Defect kinds

not-a-zip, too-many-members, unsafe-member-name, member-too-large, suspicious-compression, archive-too-large, metadata-too-large, metadata-unreadable, member-unreadable, nesting-too-deep, container-budget-exhausted, decompression-budget-exhausted, member-budget-exhausted, ambiguous-name, nameless-member.

These strings are part of the public surface; a test in this package fails if the code grows a kind that this list does not name.

vdi2770.REFUSAL_KINDS is the subset of those that can name a member in Container.rejected — what a caller needs a sentence for. Working that subset out by reading this module's source is how two of them came to be missed.

The last three are the budgets that span the whole read rather than one archive: a documentation container may legitimately hold hundreds of inner containers, and their metadata is held for as long as you walk the tree. Ten thousand of them, each with sixteen megabytes of metadata, is a permitted input under every per-archive limit and about 156 GiB of memory — and the same tree can ask the readability sweep to inflate terabytes while no single member is over its cap. MAX_CONTAINERS and MAX_TOTAL_METADATA_BYTES bound the first, MAX_TOTAL_DECOMPRESSED the second. MAX_TOTAL_MEMBERS bounds a third thing the other two do not: this package keeps one record per entry named anywhere in the tree, and ten thousand entries in each of a thousand archives is ten million of them whatever their bytes weigh. Hitting any of them is reported rather than silently truncating the tree.

Supported

Python 3.9 and up. The budgets are module constants in vdi2770.zipread — per archive: MAX_MEMBERS, MAX_MEMBER_BYTES, MAX_TOTAL_BYTES, MAX_RATIO with its MIN_SUSPICIOUS_BYTES floor, MAX_METADATA_BYTES, MAX_CONTAINER_LEVELS; across one read: MAX_CONTAINERS, MAX_TOTAL_METADATA_BYTES, MAX_TOTAL_DECOMPRESSED, MAX_TOTAL_MEMBERS. vdi2770.xmlread adds MAX_ELEMENTS, MAX_TEXT_PIECES, MAX_ATTRIBUTES_PER_ELEMENT and MAX_ATTRIBUTES. The first bounds the tree built out of one metadata file — the bytes were bounded and that tree was not, and the expansion between them is the sender's to choose. The second bounds the text hung off it, which the element count does not see: a document of three elements whose text is 450,000 character references — a 4.1 KiB archive — held 48 MB before this bound existed and holds 23 MB now. The last two bound the attributes hung off it, which neither of the others sees: attributes are cheap to write and the schema check downstream is quadratic in how many sit on one element, so 12,000 of them in a 27 KiB archive cost 13.6 seconds. vdi2770.pdfread has thirteen of its own for the PDF scan: MAX_STREAMS with the MAX_STREAM_MARKERS that bounds how many places are looked at to find that many — the second exists because a marker the scan rejects still costs it something — MAX_STREAM_SCAN, MAX_INFLATED_PER_STREAM, MAX_INFLATED_TOTAL, MAX_INFLATED_PER_READ, MAX_OBJ_PROBES, MAX_XMP_PACKETS, MAX_PDFA_PREFIXES, MAX_TRAILER_SCAN with MAX_TRAILER_BYTES and the MAX_TRAILERS the second is derived from — one bounds how much of a single trailer dictionary is read and the other how much all of them together may cost — and MAX_LINE_LOOKBACK, how far back a token looks for the start of its line, which is what examining one costs. That last is small on purpose: it divides into MAX_TRAILER_BYTES to bound how many tokens are examined at all, and the two attacks it sits between pull in opposite directions. Every trailer in the file is a candidate, newest first, because a bound on how many to read is a bound on where to look, and whoever appends to a file can push the real trailer past one. A test fails if either module grows one this list does not name. You can read them all, and they are deliberately not arguments, so a caller cannot turn them off by accident.

Unofficial

Not affiliated with, endorsed by, or connected to VDI, the Digital Data Chain Consortium, or IDTA. VDI 2770 is a guideline published by the Verein Deutscher Ingenieure; this is an independent reader for the container format it describes, written without access to the guideline text, which is sold rather than published. What that means for what this library can and cannot claim is spelled out in the validator's scope note.

Apache-2.0.

Release files for vdi2770 0.9.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for vdi2770 0.9.3
File Size Uploaded
vdi2770-0.9.3.tar.gz 178.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for vdi2770 0.9.3
File Interpreter ABI Platform
vdi2770-0.9.3-py3-none-any.whl Python 3 none any Details

Total release size: 330.0 kB

Release files / vdi2770-0.9.3.tar.gz

Download URL vdi2770-0.9.3.tar.gz
Size 178.0 kB
Tags Source
SHA-256 checksum
How to use checksums
e58e75b5ad9f3b7dd18c3040debd3c11ab1d5abc5c12351c04de87dab6feccef
BLAKE2b-256 checksum
How to use checksums
f61bf8ffa7bcba61829aad588f31cca6ecbd11033f2948cbb5db6391a26982c7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release files / vdi2770-0.9.3-py3-none-any.whl

Download URL vdi2770-0.9.3-py3-none-any.whl
Size 151.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7d2335e22941d6d2d3d78cec8cfee67e46e88230f81cf2dc3d9d3d0c36305c66
BLAKE2b-256 checksum
How to use checksums
ccd39ec06d8a75e04015109a9afc128d2c51f3b6517f58290f544b04c99ab7cb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release history Release notifications | RSS feed

0.10.0

2 release files

0.9.7

2 release files

0.9.6

2 release files

0.9.5

2 release files

0.9.4

2 release files

This release

0.9.3 This release

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.4.0

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page