dv_schema_models
Pydantic models for Dataverse metadata — parse the schema, load dataset exports, and validate field values against the schema.
[!CAUTION] This library is under active development and the API is not yet stable. Breaking changes may occur between releases. Please pin to a specific version in your
pyproject.tomlorrequirements.txtif you want to avoid surprises.
Pre-requisites
- Python 3.10+
Installation
- With
uv(recommended):
uv add dv_schema_models
- With
pip:
pip install dv_schema_models
To export schemas to Excel (see usage #5), install with the spreadsheet extra:
uv add "dv_schema_models[spreadsheet]" # or: pip install "dv_schema_models[spreadsheet]"
Concepts
| Thing | What it is |
|---|---|
| Schema | /api/metadatablocks response — defines what fields can exist, their types, and rules |
| Dataset instance | GET /api/datasets/:id response — the actual metadata values for one dataset |
| Record model | A Pydantic model generated from the schema, used to validate instance values |
Usage
1. Load and query the schema
import json
from dv_schema_models.dataverse_schema import load_schema
schema = load_schema(json.load(open("dv_schema.json")))
schema.block_names() # ['citation', 'geospatial', ...]
block = schema.get_block("citation")
block.fields.keys() # top-level field names
block.required_fields() # leaf fields where isRequired=True
block.all_leaf_fields() # flattened, including nested compound fields
field = block.get_field("keyword")
field.is_compound() # True — has childFields
field.iter_leaf_fields() # [keywordValue, keywordVocabulary, ...]
2. Load a dataset and read values
import json
from dv_schema_models.dataset_instance import load_dataset
dataset = load_dataset(json.load(open("ds_metadata.json")))
# Load the possible typeNames for a given block
dataset.field_names("citation") # ['title', 'author', 'keyword', ...]
dataset.data.latestVersion.metadataBlocks.get("citation").field_names() # same
# Shortcut from the top level
dataset.get_value("citation", "title") # plain string
# Or drill down
block = dataset.data.latestVersion.metadataBlocks.get("citation")
block.get_value("keyword") # unwrapped Python value (str / list / dict)
block.get_field("author").simple_value() # [{'authorName': 'Author1', 'authorAffiliation': 'Author1Aff'...} ... {'authorName': 'Author2', 'authorAffiliation': 'Author2Aff'...}]
# Pull one subfield out of a compound field
block.get_subfield_values("author", "authorName") # ['Author1', 'Author2']
3. Work with files
from dv_schema_models.dataset_instance import load_dataset
from dv_schema_models.file_instance import FileInstance
dataset = load_dataset(json.load(open("ds_metadata.json")))
files = dataset.data.latestVersion.files or []
FileInstance.sum_field(files, "filesize") # sum a DataFile field across files, e.g. total filesize
FileInstance.list_field(files, "dataFile.filename") # list a field's values across files, dotted path for nested fields
sum_field skips files with no dataFile or a None value for the field. Returns None (and logs a warning) if any present value isn't numeric.
list_field takes a dotted path (e.g. "restricted" for a top-level field, "dataFile.checksum.type" for a nested one) and skips entries where the path is missing or None.
4. Validate instance values against the schema
import json
from dv_schema_models.dataverse_schema import load_schema
from dv_schema_models.dataset_instance import load_dataset
from dv_schema_models.schema_driven_records import build_record_model, flatten_instance
schema = load_schema(json.load(open("dv_schema.json")))
dataset = load_dataset(json.load(open("ds_metadata.json")))
citation_schema = schema.get_block("citation")
CitationRecord = build_record_model(citation_schema) # dynamic Pydantic model
block = dataset.data.latestVersion.metadataBlocks.get("citation")
raw = flatten_instance(block) # {typeName: value, ...}
record = CitationRecord.model_validate(raw)
The generated model enforces field names, required/optional status, list wrapping for multiple=True fields, and int/float types where declared by the schema.
5. Discover available fields
# Fields actually present in this dataset instance
block = dataset.data.latestVersion.metadataBlocks.get("citation")
block.field_names() # e.g. ['title', 'author', 'keyword', ...]
# All fields the schema defines (including absent/optional ones)
schema.get_block("citation").all_leaf_fields().keys()
# After validation, access as typed attributes
record = CitationRecord.model_validate(flatten_instance(block))
record.title # str
record.author # list[...] for multiple=True compound fields
record.keyword # None if not present in this dataset (optional fields default to None)
# Note: field names with dots become underscores — e.g. 'resolution.Spatial' → record.resolution_Spatial
6. Export the schema to a spreadsheet
Requires the spreadsheet extra (see Installation).
import json
from dv_schema_models.dataverse_schema import load_schema
from dv_schema_models.schema_spreadsheet import SchemaSpreadsheet
schema = load_schema(json.load(open("dv_schema.json")))
SchemaSpreadsheet(schema).write("dv_schema.xlsx")
Writes an .xlsx workbook with one formatted worksheet per metadata block plus a combined All sheet. See docs/schema_spreadsheet/README.md for the output layout, column mapping, and architecture.
Input file shapes
Schema — output of Dataverse /api/metadatablocks:
{"status": "OK", "data": [{"id": 10, "name": "citation", "fields": {...}}]}
Dataset — output of Dataverse GET /api/datasets/:id:
{"status": "OK", "data": {"latestVersion": {"metadataBlocks": {"citation": {"fields": [...]}}}}}
Citation
If you use this library in your work, please cite according to CITATION
License
Metadata
Release files for dv_schema_models 0.8.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dv_schema_models-0.8.0.tar.gz | 13.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dv_schema_models-0.8.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 28.6 kB
Release files / dv_schema_models-0.8.0.tar.gz
| Download URL | dv_schema_models-0.8.0.tar.gz |
|---|---|
| Size | 13.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d6312f2c59e9ec2a0c901ac519f92b945c3b93c5cfe9e74dbfed4a2e2f390e9f
|
|
BLAKE2b-256 checksum How to use checksums |
4e7ec4dbad35a18eaf93c0b589023075d3742e14ed11559f4e82046b5b560d26
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / dv_schema_models-0.8.0-py3-none-any.whl
| Download URL | dv_schema_models-0.8.0-py3-none-any.whl |
|---|---|
| Size | 14.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f008c5e6390bac68fde4be1f6a6a2e031585f79fb9113fa599153a39556b7c5a
|
|
BLAKE2b-256 checksum How to use checksums |
b3e7092083b7143d1b60c6979ee73f720e6a8516550bfa13135facb925607a50
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|