This release is a pre-release and may not be stable for production use.
dv_schema_models
Pydantic models for Dataverse metadata — parse the schema, load dataset exports, and validate field values against the schema.
[!CAUTION] This library is under active development and the API is not yet stable. Breaking changes may occur between releases. Please pin to a specific version in your
pyproject.tomlorrequirements.txtif you want to avoid surprises.
Pre-requisites
- Python 3.13+
Installation
- With
uv(recommended):
uv add dv_schema_models
- With
pip:
pip install dv_schema_models
Concepts
| Thing | What it is |
|---|---|
| Schema | /api/metadatablocks response — defines what fields can exist, their types, and rules |
| Dataset instance | GET /api/datasets/:id response — the actual metadata values for one dataset |
| Record model | A Pydantic model generated from the schema, used to validate instance values |
Usage
1. Load and query the schema
import json
from dv_schema_models.dataverse_schema import load_schema
schema = load_schema(json.load(open("dv_schema.json")))
schema.block_names() # ['citation', 'geospatial', ...]
block = schema.get_block("citation")
block.fields.keys() # top-level field names
block.required_fields() # leaf fields where isRequired=True
block.all_leaf_fields() # flattened, including nested compound fields
field = block.get_field("keyword")
field.is_compound() # True — has childFields
field.iter_leaf_fields() # [keywordValue, keywordVocabulary, ...]
2. Load a dataset and read values
import json
from dv_schema_models.dataset_instance import load_dataset
dataset = load_dataset(json.load(open("ds_metadata.json")))
# Shortcut from the top level
dataset.get_value("citation", "title") # plain string
# Or drill down
block = dataset.data.latestVersion.metadataBlocks.get("citation")
block.get_value("keyword") # unwrapped Python value (str / list / dict)
block.get_field("author").simple_value() # same, from the DatasetFieldValue directly
3. Validate instance values against the schema
import json
from dv_schema_models.dataverse_schema import load_schema
from dv_schema_models.dataset_instance import load_dataset
from dv_schema_models.schema_driven_records import build_record_model, flatten_instance
schema = load_schema(json.load(open("dv_schema.json")))
dataset = load_dataset(json.load(open("ds_metadata.json")))
citation_schema = schema.get_block("citation")
CitationRecord = build_record_model(citation_schema) # dynamic Pydantic model
block = dataset.data.latestVersion.metadataBlocks.get("citation")
raw = flatten_instance(block) # {typeName: value, ...}
record = CitationRecord.model_validate(raw)
The generated model enforces field names, required/optional status, list wrapping for multiple=True fields, and int/float types where declared by the schema.
4. Discover available fields
# Fields actually present in this dataset instance
block = dataset.data.latestVersion.metadataBlocks.get("citation")
block.field_names() # e.g. ['title', 'author', 'keyword', ...]
# All fields the schema defines (including absent/optional ones)
schema.get_block("citation").all_leaf_fields().keys()
# After validation, access as typed attributes
record = CitationRecord.model_validate(flatten_instance(block))
record.title # str
record.author # list[...] for multiple=True compound fields
record.keyword # None if not present in this dataset (optional fields default to None)
# Note: field names with dots become underscores — e.g. 'resolution.Spatial' → record.resolution_Spatial
Input file shapes
Schema — output of Dataverse /api/metadatablocks:
{"status": "OK", "data": [{"id": 10, "name": "citation", "fields": {...}}]}
Dataset — output of Dataverse GET /api/datasets/:id:
{"status": "OK", "data": {"latestVersion": {"metadataBlocks": {"citation": {"fields": [...]}}}}}
Citation
If you use this library in your work, please cite according to CITATION
License
Metadata
Release files for dv_schema_models 0.2.0a0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dv_schema_models-0.2.0a0.tar.gz | 9.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dv_schema_models-0.2.0a0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 18.6 kB
Release files / dv_schema_models-0.2.0a0.tar.gz
| Download URL | dv_schema_models-0.2.0a0.tar.gz |
|---|---|
| Size | 9.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
236dce39298bf2707b97d1254801026260dded7e87d669f7b5ccd1a98784f629
|
|
BLAKE2b-256 checksum How to use checksums |
6fbc1fb0d683477260a7c4dd2e644a3a29758bbb3769116570c5b45569bd1007
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.28 {"installer":{"name":"uv","version":"0.11.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / dv_schema_models-0.2.0a0-py3-none-any.whl
| Download URL | dv_schema_models-0.2.0a0-py3-none-any.whl |
|---|---|
| Size | 9.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
de93c62620b7ff5827425ede7cfa7806286d646967a4d2b122a49060b527ad3b
|
|
BLAKE2b-256 checksum How to use checksums |
76fb4b8797390cf34798fe1dbdc9e9f4e51f8e17a8c9a02938197ab69efdaf3e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.28 {"installer":{"name":"uv","version":"0.11.28","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|