dv_schema_models
Pydantic models for Dataverse metadata — parse the schema, load dataset exports, and validate field values against the schema.
[!CAUTION] This library is under active development and the API is not yet stable. Breaking changes may occur between releases. Please pin to a specific version in your
pyproject.tomlorrequirements.txtif you want to avoid surprises.
Pre-requisites
- Python 3.10+
Installation
- With
uv(recommended):
uv add dv_schema_models
- With
pip:
pip install dv_schema_models
To export schemas to Excel (see usage #5), install with the spreadsheet extra:
uv add "dv_schema_models[spreadsheet]" # or: pip install "dv_schema_models[spreadsheet]"
Concepts
| Thing | What it is |
|---|---|
| Schema | /api/metadatablocks response — defines what fields can exist, their types, and rules |
| Dataset instance | GET /api/datasets/:id response — the actual metadata values for one dataset |
| Record model | A Pydantic model generated from the schema, used to validate instance values |
Usage
1. Load and query the schema
import json
from dv_schema_models.dataverse_schema import load_schema
schema = load_schema(json.load(open("dv_schema.json")))
schema.block_names() # ['citation', 'geospatial', ...]
block = schema.get_block("citation")
block.fields.keys() # top-level field names
block.required_fields() # leaf fields where isRequired=True
block.all_leaf_fields() # flattened, including nested compound fields
field = block.get_field("keyword")
field.is_compound() # True — has childFields
field.iter_leaf_fields() # [keywordValue, keywordVocabulary, ...]
2. Load a dataset and read values
import json
from dv_schema_models.dataset_instance import load_dataset
dataset = load_dataset(json.load(open("ds_metadata.json")))
# Load the possible typeNames for a given block
dataset.field_names("citation") # ['title', 'author', 'keyword', ...]
dataset.data.latestVersion.metadataBlocks.get("citation").field_names() # same
# Shortcut from the top level
dataset.get_value("citation", "title") # plain string
# Or drill down
block = dataset.data.latestVersion.metadataBlocks.get("citation")
block.get_value("keyword") # unwrapped Python value (str / list / dict)
block.get_field("author").simple_value() # [{'authorName': 'Author1', 'authorAffiliation': 'Author1Aff'...} ... {'authorName': 'Author2', 'authorAffiliation': 'Author2Aff'...}]
# Pull one subfield out of a compound field
block.get_subfield_values("author", "authorName") # ['Author1', 'Author2']
3. Work with files
from dv_schema_models.dataset_instance import load_dataset
from dv_schema_models.file_instance import FileInstance
dataset = load_dataset(json.load(open("ds_metadata.json")))
files = dataset.data.latestVersion.files or []
FileInstance.sum_field(files, "filesize") # sum a DataFile field across files, e.g. total filesize
FileInstance.list_field(files, "dataFile.filename") # list a field's values across files, dotted path for nested fields
sum_field skips files with no dataFile or a None value for the field. Returns None (and logs a warning) if any present value isn't numeric.
list_field takes a dotted path (e.g. "restricted" for a top-level field, "dataFile.checksum.type" for a nested one) and skips entries where the path is missing or None.
4. Work with role assignments
import json
from dv_schema_models.role_assignments import load_role_assignments
role_assignments = load_role_assignments(json.load(open("ds_role_assignments.json")))
role_assignments.count_field("assignee") # number of assignments with an "assignee" field
role_assignments.count_field("roleName", "Curator") # number of assignments where roleName == "Curator"
role_assignments.get_value("assignee") # ['@personA', '@personB', ...]
# Fields not on the schema (e.g. the `_roleAlias` Dataverse sends) are still reachable
role_assignments.data[0].get_raw("_roleAlias") # 'curator'
load_role_assignments also accepts the error envelope Dataverse returns when the request isn't permitted ({"status": "ERROR", "message": "..."}) — data is None, and message is reachable via role_assignments.model_extra.
5. Validate instance values against the schema
import json
from dv_schema_models.dataverse_schema import load_schema
from dv_schema_models.dataset_instance import load_dataset
from dv_schema_models.schema_driven_records import build_record_model, flatten_instance
schema = load_schema(json.load(open("dv_schema.json")))
dataset = load_dataset(json.load(open("ds_metadata.json")))
citation_schema = schema.get_block("citation")
CitationRecord = build_record_model(citation_schema) # dynamic Pydantic model
block = dataset.data.latestVersion.metadataBlocks.get("citation")
raw = flatten_instance(block) # {typeName: value, ...}
record = CitationRecord.model_validate(raw)
The generated model enforces field names, required/optional status, list wrapping for multiple=True fields, and int/float types where declared by the schema.
6. Discover available fields
# Fields actually present in this dataset instance
block = dataset.data.latestVersion.metadataBlocks.get("citation")
block.field_names() # e.g. ['title', 'author', 'keyword', ...]
# All fields the schema defines (including absent/optional ones)
schema.get_block("citation").all_leaf_fields().keys()
# After validation, access as typed attributes
record = CitationRecord.model_validate(flatten_instance(block))
record.title # str
record.author # list[...] for multiple=True compound fields
record.keyword # None if not present in this dataset (optional fields default to None)
# Note: field names with dots become underscores — e.g. 'resolution.Spatial' → record.resolution_Spatial
7. Export the schema to a spreadsheet
Requires the spreadsheet extra (see Installation).
import json
from dv_schema_models.dataverse_schema import load_schema
from dv_schema_models.schema_spreadsheet import SchemaSpreadsheet
schema = load_schema(json.load(open("dv_schema.json")))
SchemaSpreadsheet(schema).write("dv_schema.xlsx")
Writes an .xlsx workbook with one formatted worksheet per metadata block plus a combined All sheet. See docs/schema_spreadsheet/README.md for the output layout, column mapping, and architecture.
Input file shapes
Schema — output of Dataverse /api/metadatablocks:
{"status": "OK", "data": [{"id": 10, "name": "citation", "fields": {...}}]}
Dataset — output of Dataverse GET /api/datasets/:id:
{"status": "OK", "data": {"latestVersion": {"metadataBlocks": {"citation": {"fields": [...]}}}}}
Role assignments — output of Dataverse GET /api/datasets/:id/assignments:
{"status": "OK", "data": [{"id": 1, "assignee": "@user", "roleId": 7, "roleName": "Curator", "definitionPointId": 34847}]}
Error responses (e.g. {"status": "ERROR", "message": "..."}, no data key) are also accepted — see usage #4.
Citation
If you use this library in your work, please cite according to CITATION
License
Metadata
Release files for dv_schema_models 0.8.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dv_schema_models-0.8.1.tar.gz | 14.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dv_schema_models-0.8.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 29.9 kB
Release files / dv_schema_models-0.8.1.tar.gz
| Download URL | dv_schema_models-0.8.1.tar.gz |
|---|---|
| Size | 14.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
03c178e8c842f07464785327f76b202bf5a252824adb17f89501c89c71ebe4f9
|
|
BLAKE2b-256 checksum How to use checksums |
8136df36eae9be73050d9bbf27c8ca0f71a91554d6998fceedfdd271c6a9eaad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / dv_schema_models-0.8.1-py3-none-any.whl
| Download URL | dv_schema_models-0.8.1-py3-none-any.whl |
|---|---|
| Size | 15.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e7176cf75c48896cac83a1205fb5a848eb033e9b14c0a7a91c56f2ebe01b346a
|
|
BLAKE2b-256 checksum How to use checksums |
e3a0b394113ddeb50ceeeac1fe8be46c466bb209d24de96d7811e6ddb9233045
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|