🧨 Dynaword tooling
| Version | 0.1.0 (Changelog) |
| Python | 3.12 or newer |
| License | CC0-1.0, the same as the Dynaword datasets |
| Used by | The Dynaword dataset repositories |
| Contact | If you have questions about the tooling please create an issue in the discussions of the Dynaword repository you are contributing to |
Table of Contents
Description
Summary
The Dynaword tooling is the code shared by all Dynaword dataset repositories. It computes descriptive statistics,
fills in the datasheets and the README, draws the plots, provides the helpers that create.py scripts use to
clean and tokenize a dataset, and runs the tests every repository has to pass.
Each repository used to keep its own copy of this code under src/dynaword. With the package installed, a
repository only contains its data, dataset cards and create.py scripts, and every repository runs the same
version of the tooling.
What a dataset repository contains
README.md # main dataset card; its front matter lists every dataset as a config
CHANGELOG.md
CONTRIBUTING.md
pyproject.toml # dataset version and the dependency on this package
makefile # the commands below
dynaword.toml # corpus-specific settings, see Settings
data/
dataset_name/
dataset_name.md # dataset card
data.parquet # the documents
metadata.parquet # Propella annotations, one row per document
create.py # recreates data.parquet from the source
descriptive_stats.json # generated
images/ # generated
The pyproject.toml, makefile and dynaword.toml a repository needs are in examples/.
Installing dependencies
The repositories use uv for environment and dependency management. After installing uv, run:
make install
Adding a new dataset
To add a new dataset create a folder under data/{dataset_name}/ with a dataset card, data.parquet,
metadata.parquet and a create.py, then register the dataset in the configs list of the README front matter.
Guidance on writing the dataset card is in the CONTRIBUTING.md of the repository.
Writing create.py
create.py recreates the dataset from its source. Use the shared helpers for the final processing steps:
from dynaword.process_dataset import (
add_token_count, # token_count column, Llama 3 tokenizer
ensure_column_order, # id, text, source, added, created, token_count
remove_duplicate_text,
remove_empty_texts,
)
ds = remove_empty_texts(ds)
ds = remove_duplicate_text(ds)
ds = add_token_count(ds)
ds = ensure_column_order(ds)
ds.to_parquet("data.parquet")
Declare the package in the script header so uv run data/{dataset_name}/create.py works on its own:
# /// script
# requires-python = ">=3.12"
# dependencies = [
# "datasets>=3.0.0",
# "dynaword>=0.1",
# ]
# ///
Until the package is published, add the [tool.uv.sources] entry from examples/pyproject.toml
to the header as well.
Updating descriptive statistics
make update-descriptive-statistics
This computes descriptive_stats.json and the plots for every dataset that does not have them yet, fills in the
marked sections of its dataset card, and then regenerates the tables, plots and statistics of the main README.
To recompute one dataset after changing it:
uv run python -m dynaword.update_descriptive_statistics --dataset {dataset_name} --force
Running tests
make test
The tests check that every dataset listed in the README loads, has unique ids, matches the schema, has a
complete dataset card, has matching annotations, and contains no duplicate or one-token documents.
The output is written to test_results.log.
Bumping the version
make bump-version
Bumps the patch version of the dataset in pyproject.toml and in the README table.
Checklist
- I have run
make test. - I have updated descriptive statistics with
make update-descriptive-statisticsafter adding or changing a dataset. - I have bumped the version with
make bump-versionif the repo contents changed materially. - I have updated
CHANGELOG.mdwhen appropriate. - I have declared any extra
create.pydependencies in the script header. - I have reviewed the generated dataset card, README and plot changes before staging them.
Settings
Values that differ between corpora live in dynaword.toml in the repository root. The file is optional; the
corpus name and its languages are read from the README front matter.
| Key | Used for | Default |
|---|---|---|
hub_id |
Links from dataset cards to the annotations section of the main README | Relative link to README.md |
reference_corpora |
Dashed reference lines in the tokens-over-time plot (name and tokens each) |
No lines |
license_names |
Licence names grouped under a generic label in the licence table | Built-in list |
removed_sources |
Folders under data/ removed on purpose and therefore not listed in the README |
None |
Frequently asked questions
Do I need src/ in the repository?
No. Delete it and add the package as a dependency in pyproject.toml, as in examples/pyproject.toml.
The make targets keep their names; only the commands behind them change, see examples/makefile.
Can I run the tooling on a repository from another directory?
Yes. The commands work on the current directory; set DYNAWORD_REPO to the path of the repository to work on it from elsewhere.
How do I develop the tooling itself?
Clone this repository, run make install, and try your change against a dataset repository with the commands above.
make lint formats and checks the code.
Metadata
Release files for dynaword 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dynaword-0.1.0.tar.gz | 266.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dynaword-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 335.7 kB
Release files / dynaword-0.1.0.tar.gz
| Download URL | dynaword-0.1.0.tar.gz |
|---|---|
| Size | 266.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
474e2b8a29d4c75c903d91c9501cb1402a3d05562491a4e50ce6fe31bab36282
|
|
BLAKE2b-256 checksum How to use checksums |
d2c428bae13fbe7abf403773ef98286369f284c1389cfea276dcaf0afc905bdd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / dynaword-0.1.0-py3-none-any.whl
| Download URL | dynaword-0.1.0-py3-none-any.whl |
|---|---|
| Size | 69.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6d8647ececa87da58d6fcf02a2542ad89b75eee71bb0b6078c2bc91e5411a8a0
|
|
BLAKE2b-256 checksum How to use checksums |
86783b11b3261c0c5a9ae1b63c2e51584d651189993a068f9c83a6dc41f69c68
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.15 {"installer":{"name":"uv","version":"0.12.15","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|