A robust, secure, and maintainable ONIX 3.x Python library using xsdata.
Project description
pyonix-core
A high-performance, type-safe, and secure Python library for processing ONIX 3.0 XML feeds. Built on top of xsdata and lxml, pyonix-core provides a robust foundation for bibliographic data exchange in the publishing industry.
Overview
ONIX for Books is the international standard for representing and communicating book industry product information in electronic form. pyonix-core simplifies the complexity of the ONIX 3.0 standard by providing:
- Strict Type Safety: Fully typed Python dataclasses generated directly from EDItEUR's official XSD schemas.
- Memory Efficiency: Iterative, streaming parsing capabilities designed to handle multi-gigabyte ONIX files with minimal memory footprint.
- Security First: Hardened XML parsing configuration to prevent XXE (XML External Entity) attacks and other common XML vulnerabilities.
- Developer Friendly: A high-level "Facade" pattern to abstract away the deeply nested structure of raw ONIX messages, providing simple access to common fields like ISBNs, titles, and prices.
Installation
Requires Python 3.11 or higher.
pip install pyonix-core
Quick Start
Parsing an ONIX File
The core entry point is parse_onix_stream, which yields product records one by one. This allows you to process massive files without loading the entire document into memory.
from pyonix_core.parsing.parser import parse_onix_stream
from pyonix_core.facade.product import ProductFacade
file_path = "path/to/onix_feed.xml"
# parse_onix_stream automatically detects Reference vs Short tags
for product in parse_onix_stream(file_path):
# Wrap the raw data model in a Facade for easier access
facade = ProductFacade(product)
print(f"Title: {facade.title}")
print(f"ISBN-13: {facade.isbn13}")
print(f"Price: {facade.price_amount} {facade.price_currency}")
print("-" * 20)
Working with the Facade
The ProductFacade simplifies data extraction. Instead of navigating complex nested objects, you can access properties directly.
# Example of accessing data via Facade
print(f"Record Reference: {facade.record_reference}")
print(f"Contributors: {', '.join(facade.contributors)}")
Architecture
Data Models
The data models in pyonix_core.models are auto-generated using xsdata from the official ONIX 3.0 schemas. This ensures 100% compliance with the standard and provides excellent IDE support (autocompletion, type checking).
Security
XML parsing is handled by lxml with strict security settings:
resolve_entities=False: Prevents XXE attacks.no_network=True: Blocks remote resource fetching.load_dtd=False: Disables DTD processing.
Development
Running Tests
python -m unittest discover tests
Regenerating Models
If the schemas change, you can regenerate the data models:
python scripts/generate_models.py
License
This project is licensed under the MIT License.
Note: This project is currently in active development.
Recent Features
The following high-value features were recently added to pyonix-core. Each one is implemented to integrate cleanly with the existing facade and parsing architecture.
- Short Tag to Reference Tag Normalizer
- Module:
pyonix_core.parsing.normalization - Description: Uses an XSLT transform (bundleable EDItEUR XSLT) to convert ONIX Short Tags into Reference Tags before parsing. This lets the parser and generated dataclasses work consistently regardless of input tag style. For very large files, a streaming/SAX-based fallback is planned.
- Quick usage:
from pyonix_core.parsing.normalization import TagNormalizer
from pyonix_core.parsing.parser import parse_onix_stream
norm = TagNormalizer()
with open('short_tag_onix.xml', 'rb') as fh:
normalized_stream = norm.normalize_stream(fh)
for product in parse_onix_stream(normalized_stream):
...
- Data Flattener (Pandas/CSV ready)
- Module:
pyonix_core.utils.flatten - Description:
ProductFlattenerconverts aProductFacadeinto a flat dictionary using a configurableSerializationProfile. Useful for creating CSV rows or constructing Pandas DataFrames without manually traversing the ONIX model. - Quick usage:
from pyonix_core.utils.flatten import ProductFlattener
from pyonix_core.facade.product import ProductFacade
f = ProductFlattener()
row = f.flatten(ProductFacade(product_model))
- HTML Sanitizer & Extractor
- Module:
pyonix_core.utils.text - Description: Safe extraction and sanitization of HTML-like content found in ONIX
TextContentcomposites. Usesbleachto sanitize andhtml2textto produce Markdown. These are optional dependencies under thetextextra inpyproject.toml. - Quick usage:
from pyonix_core.utils.text import clean_html, to_markdown
safe_html = clean_html(raw_html)
md = to_markdown(raw_html)
- ISBN Tools (10/13 converter)
- Module:
pyonix_core.utils.identifiers - Description: Pure-Python
ISBNutility for cleaning, validating, and converting ISBN-10 to ISBN-13.ProductFacade.isbn13now prefers an explicit ISBN-13 but will auto-convert valid ISBN-10 values when necessary. - Quick usage:
from pyonix_core.utils.identifiers import ISBN
ISBN.clean('0-306-40615-2')
ISBN.validate('9780306406157')
ISBN.to_13('0-306-40615-2')
- Media Asset Helper
- Module:
pyonix_core.facade.assets - Description:
AssetHelpersimplifies locating front cover images and other collateral assets within theCollateralDetailcomposite and exposes ahelperproperty onProductFacade. - Quick usage:
from pyonix_core.facade.product import ProductFacade
facade = ProductFacade(product_model)
cover_url = facade.helper.get_cover_image()
All of these features are exercised in the test-suite (excluding extremely large-file performance tests). Optional dependencies are declared under [project.optional-dependencies] in pyproject.toml (see the text extra for HTML utilities).
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pyonix_core-0.2.1.tar.gz.
File metadata
- Download URL: pyonix_core-0.2.1.tar.gz
- Upload date:
- Size: 1.3 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ad223cc90bd803019a55ad69277494a1082a3bed2d265bd063b5e089c57cad34
|
|
| MD5 |
a4c8dddb0e1478d37c9692adfd1a7b5a
|
|
| BLAKE2b-256 |
a22d6c938692868c7f411ee8da4da4c7e92705210419fcc592f1cd57dde53319
|
File details
Details for the file pyonix_core-0.2.1-py3-none-any.whl.
File metadata
- Download URL: pyonix_core-0.2.1-py3-none-any.whl
- Upload date:
- Size: 2.1 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8b45c798b7c55d6467480f38733d0418dd1bc338718891a3c0bb60b4573dce75
|
|
| MD5 |
df1524c39cb1449652632eb335305c43
|
|
| BLAKE2b-256 |
7181b8c66835b60f36801c5a0e0d36c85d5d154eda7feb15a0a7598753d6d2f9
|