Skip to main content

🧨 Dynaword tooling

Version 0.1.0 (Changelog)
Python 3.12 or newer
License CC0-1.0, the same as the Dynaword datasets
Used by The Dynaword dataset repositories
Contact If you have questions about the tooling please create an issue in the discussions of the Dynaword repository you are contributing to

Table of Contents

Description

Summary

The Dynaword tooling is the code shared by all Dynaword dataset repositories. It computes descriptive statistics, fills in the datasheets and the README, draws the plots, provides the helpers that create.py scripts use to clean and tokenize a dataset, and runs the tests every repository has to pass.

Each repository used to keep its own copy of this code under src/dynaword. With the package installed, a repository only contains its data, dataset cards and create.py scripts, and every repository runs the same version of the tooling.

What a dataset repository contains

README.md              # main dataset card; its front matter lists every dataset as a config
CHANGELOG.md
CONTRIBUTING.md
pyproject.toml         # dataset version and the dependency on this package
makefile               # the commands below
dynaword.toml          # corpus-specific settings, see Settings
data/
  dataset_name/
    dataset_name.md    # dataset card
    data.parquet       # the documents
    metadata.parquet   # Propella annotations, one row per document
    create.py          # recreates data.parquet from the source
    descriptive_stats.json   # generated
    images/                  # generated

The pyproject.toml, makefile and dynaword.toml a repository needs are in examples/.

Installing dependencies

The repositories use uv for environment and dependency management. After installing uv, run:

make install

Adding a new dataset

To add a new dataset create a folder under data/{dataset_name}/ with a dataset card, data.parquet, metadata.parquet and a create.py, then register the dataset in the configs list of the README front matter. Guidance on writing the dataset card is in the CONTRIBUTING.md of the repository.

Writing create.py

create.py recreates the dataset from its source. Use the shared helpers for the final processing steps:

from dynaword.process_dataset import (
    add_token_count,        # token_count column, Llama 3 tokenizer
    ensure_column_order,    # id, text, source, added, created, token_count
    remove_duplicate_text,
    remove_empty_texts,
)

ds = remove_empty_texts(ds)
ds = remove_duplicate_text(ds)
ds = add_token_count(ds)
ds = ensure_column_order(ds)
ds.to_parquet("data.parquet")

Declare the package in the script header so uv run data/{dataset_name}/create.py works on its own:

# /// script
# requires-python = ">=3.12"
# dependencies = [
#     "datasets>=3.0.0",
#     "dynaword>=0.1",
# ]
# ///

Until the package is published, add the [tool.uv.sources] entry from examples/pyproject.toml to the header as well.

Updating descriptive statistics

make update-descriptive-statistics

This computes descriptive_stats.json and the plots for every dataset that does not have them yet, fills in the marked sections of its dataset card, and then regenerates the tables, plots and statistics of the main README. To recompute one dataset after changing it:

uv run python -m dynaword.update_descriptive_statistics --dataset {dataset_name} --force

Running tests

make test

The tests check that every dataset listed in the README loads, has unique ids, matches the schema, has a complete dataset card, has matching annotations, and contains no duplicate or one-token documents. The output is written to test_results.log.

Bumping the version

make bump-version

Bumps the patch version of the dataset in pyproject.toml and in the README table.

Checklist

  • I have run make test.
  • I have updated descriptive statistics with make update-descriptive-statistics after adding or changing a dataset.
  • I have bumped the version with make bump-version if the repo contents changed materially.
  • I have updated CHANGELOG.md when appropriate.
  • I have declared any extra create.py dependencies in the script header.
  • I have reviewed the generated dataset card, README and plot changes before staging them.

Settings

Values that differ between corpora live in dynaword.toml in the repository root. The file is optional; the corpus name and its languages are read from the README front matter.

Key Used for Default
hub_id Links from dataset cards to the annotations section of the main README Relative link to README.md
reference_corpora Dashed reference lines in the tokens-over-time plot (name and tokens each) No lines
license_names Licence names grouped under a generic label in the licence table Built-in list
removed_sources Folders under data/ removed on purpose and therefore not listed in the README None

See examples/dynaword.toml.

Frequently asked questions

Do I need src/ in the repository?

No. Delete it and add the package as a dependency in pyproject.toml, as in examples/pyproject.toml. The make targets keep their names; only the commands behind them change, see examples/makefile.

Can I run the tooling on a repository from another directory?

Yes. The commands work on the current directory; set DYNAWORD_REPO to the path of the repository to work on it from elsewhere.

How do I develop the tooling itself?

Clone this repository, run make install, and try your change against a dataset repository with the commands above. make lint formats and checks the code.

Metadata

Release files for dynaword 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for dynaword 0.1.0
File Size Uploaded
dynaword-0.1.0.tar.gz 266.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for dynaword 0.1.0
File Interpreter ABI Platform
dynaword-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 335.7 kB

Release files / dynaword-0.1.0.tar.gz

Download URL dynaword-0.1.0.tar.gz
Size 266.8 kB
Tags Source
SHA-256 checksum
How to use checksums
474e2b8a29d4c75c903d91c9501cb1402a3d05562491a4e50ce6fe31bab36282
BLAKE2b-256 checksum
How to use checksums
d2c428bae13fbe7abf403773ef98286369f284c1389cfea276dcaf0afc905bdd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / dynaword-0.1.0-py3-none-any.whl

Download URL dynaword-0.1.0-py3-none-any.whl
Size 69.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6d8647ececa87da58d6fcf02a2542ad89b75eee71bb0b6078c2bc91e5411a8a0
BLAKE2b-256 checksum
How to use checksums
86783b11b3261c0c5a9ae1b63c2e51584d651189993a068f9c83a6dc41f69c68
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page