Skip to main content

(pyscgen) python schema generator

A Python Package to analyze JSON Documents, build a merged JSON-Document out of multiple provided JSON Documents and create an AVRO-Schemas based on multiple given JSON Documents.

Installation

pip

pip install pyscgen

poetry

poetry add pyscgen

JSON Analyzer

receives a list of json documents, analyzes the structure and outputs the following infos:

-Output:

  • collection_data:
    • dict which stores infos about the columns of each document with attributes like name, path, data type and if its null
  • column_infos:
    • condensed/merged column infos of all given documents with attributes like name, path, nullability, density, unique values, data types found, parent column config etc.
    • This is the "real" result of the analyzer and the building plan for the JSON Merger and AVRO Schema generator.
  • df_flattened
    • A pandas DataFrame which stores the json documents flattened and contains every found column with data.
    • One document is represented by one row, index starts at 0 and matches to the order in the given list of documents.
  • df_dtypes
    • A pandas DataFrame which stores the python data type for each column of all given documents.
    • One document is represented by one row, index starts at 0 and matches to the order in the given list of documents.
  • df_unique
    • A pandas DataFrame which stores unique values found for each column of all given documents.
    • This data frame is pivoted
      • The column "0" stores all found column. One column is represented by one row.
      • The columns 1 - n contain one distinct value each. If you analyze 10 documents and one field has a distinct value in each document, you´ll produce 10 value columns.

JSON Merger

receives a list of json documents, analyzes the structure with the JSON Analyzer and outputs one merged dictionary/json document with all found columns and dummy values according to the found data types:

  • Output: merged_doc:
    • dictionary merged together from the all given JSON documents.

AVRO Schema generator

receives a list of json documents, analyzes the structure with the JSON Analyzer and outputs a "Schema"-object which can be converted to a dict and stored in an avsc-file

  • Flow:
    • JSON Analyzer -> JSON Merger -> AVRO Schema generator
  • Output: avro_schema:
    • AVRO Schema object can be converted to a dict and stored to an avsc file, see examples.
  • Limitations/Currently not supported:
    • Resolution of duplicated names in the AVRO-Output
    • empty dicts/maps:
      • If one of your JSON input documents has an empty dictionary, e. g. {"field1": "value1", "field2": {}} an empty AVRO record will be generated, which is in valid in AVRO.
      • You can resolve this by deleting the empty record or filling it with live if you know that it´s still needed in the future.
    • Mixed types for one field in the the JSON Input documents currently results in the most generic AVRO type, which ist a String (Union types, as an option, is planned)
      • This leads to the case that not all input documents automaticall validate positively against the AVRO schema, because you need to convert types beforehand.
      • You can use the pydantic model generator to create a model, which automatically converts other data types to string, to solve this issue. Just let all your JSON flow through the model bevore validation.

Pydantic Model generator

receives a list of json documents analyzes the structure with the JSON Analyzer, creates an AVRO Schema which is then used to create a pydantic Model with pydantic-avro

  • Flow
    • JSON Analyzer -> JSON Merger -> AVRO Schema generator -> Pydantic Model generator
  • Output: pydantic model:
    • Pydantic Model as a string which can be written to a .py-File.
  • Limitations/Currently not supported:
    • Python/AVRO bytes is currently not supported in pydantic model generation because there is no support in pydantic for this datatype right now.

Why should you use it?

Considering the fact that there are already other AVRO Schema generators, why should you use this one?

The simple answer is "null" or "nullability".

All other solutions I tried simply let you pass one JSON Document as a base for the AVRO-Schema generation. This does obviously not allow to infer nullabilty, because only what's present in this one message can be observed and therefore used for schema generation.

This library lets you pass a list of JSON Documents and therefore can gather the infos which fields are always present - aka. mandatory - and which fields are only found in a couple documents - and therefore nullable.

Handling nullability in AVRO-Schemas for records and arrays is quite painful in my opinion, this library gets rid of this chore for you.

Of course, it works just fine with only one JSON document like any other AVRO generator, just without the nullability part then.

Principles

  • CI-CO (Crap in - Crap out):
    • The JSON Documents are not validated, therefore, if you pass in garbage, crap in - crap out.
    • The goal is to creata an AVRO-Schema no matter what. In my opinion, an half ready AVRO-Schema which needs to be edited by hand is better than no AVRO-Schema at all!

Examples:

Other useful resources regarding AVRO:

License:

BSD-3-Clause

Metadata

Release files for pyscgen 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyscgen 0.3.0
File Size Uploaded
pyscgen-0.3.0.tar.gz 52.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyscgen 0.3.0
File Interpreter ABI Platform
pyscgen-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 74.3 kB

Release files / pyscgen-0.3.0.tar.gz

Download URL pyscgen-0.3.0.tar.gz
Size 52.9 kB
Tags Source
SHA-256 checksum
How to use checksums
799f72a816b06993d544fec503054a7522a082572bab0c6571b552c39aa0add5
BLAKE2b-256 checksum
How to use checksums
8b8e1661c434f071408b1cf5b3483494bd09620a21400e46a855c6733c2f6452
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.2.1 CPython/3.13.9 Windows/11

Release files / pyscgen-0.3.0-py3-none-any.whl

Download URL pyscgen-0.3.0-py3-none-any.whl
Size 21.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
fc6ddd5a10352e9bb30f4355ccf35d32f09259f239fb3ee65d73bce7652b5e3e
BLAKE2b-256 checksum
How to use checksums
80ca1d5a144b82efac1b73388d3f8050c375501533f854ed75de4f7a7c6ae1a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.2.1 CPython/3.13.9 Windows/11

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page