microdata-tools
Tools for the microdata.no platform
Installation
microdata-tools can be installed from PyPI using pip:
pip install microdata-tools
Usage
Once you have your metadata and data files ready to go, they should be named and stored like this:
my-input-directory/
MY_DATASET_NAME/
MY_DATASET_NAME.csv
MY_DATASET_NAME.json
The CSV file is optional in some cases.
Package dataset
The package_dataset() function will encrypt and package your dataset as a tar archive. The process is as follows:
- Generate the symmetric key for a dataset.
- Encrypt the dataset data (CSV) using an AES-256-GCM symmetric key and store the encrypted file as
<DATASET_NAME>.csv.encr - Encrypt the symmetric key using HPKE with the combined ML-KEM-768/X25519 public key from
microdata_public_key.pemand store the resulting HPKE ciphertext as <DATASET_NAME>.kem.encr - Gather the encrypted CSV, ciphertext file and metadata (JSON) file in one tar file.
Unpackage dataset
The unpackage_dataset() function will untar and your dataset and use the combined ML-KEM-768/X25519 private key from microdata_private_key.pem to recover the symmetric key, which is then used to decrypt the dataset.
The packaged file has to have the <DATASET_NAME>.tar extension. Its contents should be as follows:
<DATASET_NAME>.json : Required medata file.
<DATASET_NAME>.csv.encr : Optional encrypted dataset file.
<DATASET_NAME>.kem.encr : Optional HPKE ciphertext file containing the encrypted symmetric key required to decrypt the dataset file. Required if the .csv.encr file is present.
Decryption uses the combined ML-KEM-768/X25519 private key located at PRIVATE_KEY_DIR to recover the symmetric decryption key.
The packaged file is then stored in output_dir/archive/unpackaged after a successful run or output_dir/archive/failed after an unsuccessful run.
Example
Store your metadata and data files according to the structure described above, and put the provided public key in a directory of your choice. Then:
from pathlib import Path
from microdata_tools import package_dataset
package_dataset(
public_key_dir=Path("path/to/key_directory"),
dataset_dir=Path("path/to/MY_DATASET_NAME"),
output_dir=Path("path/to/output"),
)
This produces path/to/output/MY_DATASET_NAME.tar, which can be uploaded to microdata.
Validation
Once you have your metadata and data files ready to go, they should be named and stored like this:
my-input-directory/
MY_DATASET_NAME/
MY_DATASET_NAME.csv
MY_DATASET_NAME.json
Note that the filename only allows upper case letters A-Z, number 0-9 and underscores.
Import microdata-tools in your script and validate your files:
from microdata_tools import validate_dataset
validation_errors = validate_dataset(
"MY_DATASET_NAME",
input_directory="path/to/my-input-directory"
)
if not validation_errors:
print("My dataset is valid")
else:
print("Dataset is invalid :(")
# You can print your errors like this:
for error in validation_errors:
print(error)
For a more in-depth explanation of usage visit the usage documentation.
Data format description
A dataset as defined in microdata consists of one data file, and one metadata file.
The data file is a csv file seperated by semicolons. A valid example would be:
000000000000001;123;2020-01-01;2020-12-31;
000000000000002;123;2020-01-01;2020-12-31;
000000000000003;123;2020-01-01;2020-12-31;
000000000000004;123;2020-01-01;2020-12-31;
For datasets with temporalityType FIXED the value does not change over time, so there is no time period to describe — both the start and stop columns should be left empty. On microdata.no such variables are shown with an infinite validity period ("∞"), see BEFOLKNING_KJOENN for an example.
000000000000001;123;;;
000000000000002;123;;;
000000000000003;123;;;
000000000000004;123;;;
Read more about the data format and columns in the documentation.
The metadata files should be in json format. The requirements for the metadata is best described through the Pydantic model, the examples, and the metadata model.
Contribute
Set up
To work on this repository you need to install uv:
# macOS / linux / BashOnWindows
curl -LsSf https://astral.sh/uv/install.sh | sh
# Windows powershell
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
Then install the virtual environment from the root directory:
uv sync
Running unit tests
Open terminal and go to root directory of the project and run:
uv run pytest
Pre-commit
There are currently 3 active rules: Ruff-format, Ruff-lint and sync lock file. Install pre-commit
pip install pre-commit
If you've made changes to the pre-commit-config.yaml or its a new project install the hooks with:
pre-commit install
Now it should run when you do:
git commit
By default it only runs against changed files. To force the hooks to run against all files:
pre-commit run --all-files
if you dont have it installed on your system you can use: (but then it won't run when you use the git-cli)
uv run pre-commit
Read more about pre-commit
Metadata
Release files for microdata-tools 2.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| microdata_tools-2.2.1.tar.gz | 44.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| microdata_tools-2.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 106.4 kB
Release files / microdata_tools-2.2.1.tar.gz
| Download URL | microdata_tools-2.2.1.tar.gz |
|---|---|
| Size | 44.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7f77ee16cbde7be924102f8564f1092f7c2cc43a3ceb68ccb16245bd13c00a8c
|
|
BLAKE2b-256 checksum How to use checksums |
6524d091c0c719404e7c028c7b6d0a12ba81995e52c06d67c3efe5bcce4f8cdc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.
Transparency logRelease files / microdata_tools-2.2.1-py3-none-any.whl
| Download URL | microdata_tools-2.2.1-py3-none-any.whl |
|---|---|
| Size | 62.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
937b4aec4e8a727e4b86d71463f8f841c5ebf4ac8b0ea603ac0e8763d18815b0
|
|
BLAKE2b-256 checksum How to use checksums |
afbb4931ecc1851be4d83b0a00cb2b20e0620f8ee3aa598586dfb8aeee4bb38d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 22, 2026.
Transparency log