Skip to main content

dv_schema_models

Pydantic models for Dataverse metadata — parse the schema, load dataset exports, and validate field values against the schema.

[!CAUTION] This library is under active development and the API is not yet stable. Breaking changes may occur between releases. Please pin to a specific version in your pyproject.toml or requirements.txt if you want to avoid surprises.

Pre-requisites

  1. Python 3.10+

Installation

  1. With uv (recommended):
uv add dv_schema_models
  1. With pip:
pip install dv_schema_models

To export schemas to Excel (see usage #5), install with the spreadsheet extra:

uv add "dv_schema_models[spreadsheet]"   # or: pip install "dv_schema_models[spreadsheet]"

Concepts

Thing What it is
Schema /api/metadatablocks response — defines what fields can exist, their types, and rules
Dataset instance GET /api/datasets/:id response — the actual metadata values for one dataset
Record model A Pydantic model generated from the schema, used to validate instance values

Usage

1. Load and query the schema

import json
from dv_schema_models.dataverse_schema import load_schema

schema = load_schema(json.load(open("dv_schema.json")))

schema.block_names()                        # ['citation', 'geospatial', ...]
block = schema.get_block("citation")
block.fields.keys()                         # top-level field names
block.required_fields()                     # leaf fields where isRequired=True
block.all_leaf_fields()                     # flattened, including nested compound fields

field = block.get_field("keyword")
field.is_compound()                         # True — has childFields
field.iter_leaf_fields()                    # [keywordValue, keywordVocabulary, ...]

2. Load a dataset and read values

import json
from dv_schema_models.dataset_instance import load_dataset

dataset = load_dataset(json.load(open("ds_metadata.json")))

# Load the possible typeNames for a given block
dataset.field_names("citation")  # ['title', 'author', 'keyword', ...] 
dataset.data.latestVersion.metadataBlocks.get("citation").field_names() # same


# Shortcut from the top level
dataset.get_value("citation", "title")      # plain string

# Or drill down
block = dataset.data.latestVersion.metadataBlocks.get("citation")
block.get_value("keyword")                  # unwrapped Python value (str / list / dict)
block.get_field("author").simple_value()    # [{'authorName': 'Author1', 'authorAffiliation': 'Author1Aff'...} ... {'authorName': 'Author2', 'authorAffiliation': 'Author2Aff'...}]

# Pull one subfield out of a compound field
block.get_subfield_values("author", "authorName")  # ['Author1', 'Author2']

3. Work with files

from dv_schema_models.dataset_instance import load_dataset
from dv_schema_models.file_instance import FileInstance

dataset = load_dataset(json.load(open("ds_metadata.json")))

files = dataset.data.latestVersion.files or []
FileInstance.sum_field(files, "filesize")   # sum a DataFile field across files, e.g. total filesize
FileInstance.list_field(files, "dataFile.filename")   # list a field's values across files, dotted path for nested fields

sum_field skips files with no dataFile or a None value for the field. Returns None (and logs a warning) if any present value isn't numeric.

list_field takes a dotted path (e.g. "restricted" for a top-level field, "dataFile.checksum.type" for a nested one) and skips entries where the path is missing or None.

4. Work with role assignments

import json
from dv_schema_models.role_assignments import load_role_assignments

role_assignments = load_role_assignments(json.load(open("ds_role_assignments.json")))

role_assignments.count_field("assignee")            # number of assignments with an "assignee" field
role_assignments.count_field("roleName", "Curator") # number of assignments where roleName == "Curator"
role_assignments.get_value("assignee")              # ['@personA', '@personB', ...]

# Fields not on the schema (e.g. the `_roleAlias` Dataverse sends) are still reachable
role_assignments.data[0].get_raw("_roleAlias")      # 'curator'

load_role_assignments also accepts the error envelope Dataverse returns when the request isn't permitted ({"status": "ERROR", "message": "..."}) — data is None, and message is reachable via role_assignments.model_extra.

5. Validate instance values against the schema

import json
from dv_schema_models.dataverse_schema import load_schema
from dv_schema_models.dataset_instance import load_dataset
from dv_schema_models.schema_driven_records import build_record_model, flatten_instance


schema = load_schema(json.load(open("dv_schema.json")))
dataset = load_dataset(json.load(open("ds_metadata.json")))

citation_schema = schema.get_block("citation")
CitationRecord = build_record_model(citation_schema)   # dynamic Pydantic model

block = dataset.data.latestVersion.metadataBlocks.get("citation")
raw = flatten_instance(block)              # {typeName: value, ...}
record = CitationRecord.model_validate(raw)

The generated model enforces field names, required/optional status, list wrapping for multiple=True fields, and int/float types where declared by the schema.

6. Discover available fields

# Fields actually present in this dataset instance
block = dataset.data.latestVersion.metadataBlocks.get("citation")
block.field_names()                            # e.g. ['title', 'author', 'keyword', ...]

# All fields the schema defines (including absent/optional ones)
schema.get_block("citation").all_leaf_fields().keys()

# After validation, access as typed attributes
record = CitationRecord.model_validate(flatten_instance(block))
record.title          # str
record.author         # list[...] for multiple=True compound fields
record.keyword        # None if not present in this dataset (optional fields default to None)
# Note: field names with dots become underscores — e.g. 'resolution.Spatial' → record.resolution_Spatial

7. Export the schema to a spreadsheet

Requires the spreadsheet extra (see Installation).

import json
from dv_schema_models.dataverse_schema import load_schema
from dv_schema_models.schema_spreadsheet import SchemaSpreadsheet

schema = load_schema(json.load(open("dv_schema.json")))
SchemaSpreadsheet(schema).write("dv_schema.xlsx")

Writes an .xlsx workbook with one formatted worksheet per metadata block plus a combined All sheet. See docs/schema_spreadsheet/README.md for the output layout, column mapping, and architecture.

Input file shapes

Schema — output of Dataverse /api/metadatablocks:

{"status": "OK", "data": [{"id": 10, "name": "citation", "fields": {...}}]}

Dataset — output of Dataverse GET /api/datasets/:id:

{"status": "OK", "data": {"latestVersion": {"metadataBlocks": {"citation": {"fields": [...]}}}}}

Role assignments — output of Dataverse GET /api/datasets/:id/assignments:

{"status": "OK", "data": [{"id": 1, "assignee": "@user", "roleId": 7, "roleName": "Curator", "definitionPointId": 34847}]}

Error responses (e.g. {"status": "ERROR", "message": "..."}, no data key) are also accepted — see usage #4.

Citation

If you use this library in your work, please cite according to CITATION

License

MIT

Metadata

Release files for dv_schema_models 0.8.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dv_schema_models 0.8.1
File Size Uploaded
dv_schema_models-0.8.1.tar.gz 14.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dv_schema_models 0.8.1
File Interpreter ABI Platform
dv_schema_models-0.8.1-py3-none-any.whl Python 3 none any Details

Total release size: 29.9 kB

Release files / dv_schema_models-0.8.1.tar.gz

Download URL dv_schema_models-0.8.1.tar.gz
Size 14.7 kB
Tags Source
SHA-256 checksum
How to use checksums
03c178e8c842f07464785327f76b202bf5a252824adb17f89501c89c71ebe4f9
BLAKE2b-256 checksum
How to use checksums
8136df36eae9be73050d9bbf27c8ca0f71a91554d6998fceedfdd271c6a9eaad
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / dv_schema_models-0.8.1-py3-none-any.whl

Download URL dv_schema_models-0.8.1-py3-none-any.whl
Size 15.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e7176cf75c48896cac83a1205fb5a848eb033e9b14c0a7a91c56f2ebe01b346a
BLAKE2b-256 checksum
How to use checksums
e3a0b394113ddeb50ceeeac1fe8be46c466bb209d24de96d7811e6ddb9233045
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.30 {"installer":{"name":"uv","version":"0.11.30","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page