Skip to main content

The Nordic Pile

Project description

The Nordic Pile replication code

The Nordic Pile is a repository with the aim of providing tools and code to download and replicate a large nordic language dataset. The dataset consists of many smaller datasets combined together. With the objective to cover a broad set of language modalities.

Workflow

To propose a new dataset be added to the Nordic Pile, open an issue. Your issue should include a description of the dataset, its size, what language(s) it is in, a link to the data, and any other relevant information. If a project manger approves your proposal, they will change its label to Datasets and add it to Project: Datasets. Datasets that we elect to not include in the current version of the Pile will receive a Deferred or Declined label. We will now focus on datasets in the languages of the nordics: Swedish, Danish, Norwegian and Finnish.

To claim responsibility for implementing an unclaimed dataset, leave a comment on one of our unassigned issues. Once a dataset has been assigned to you, make the necessary changes to datasets.py and pile.py in a fork and submit a pull request. If you require, you can also submit a script for processing the data as shown here.

To raise an issue that is not proposing a new dataset, open an issue with the tag Feature Request or Bug as appropriate.

Data ready for final implementation should meet the following criteria:

  • The data must be in lm_dataformat format.
  • The data must be shuffled.

Attribution

This initiative is heavily inspired by Eleuther AIs The Pile project.
https://www.eleuther.ai/
https://pile.eleuther.ai/

Datasets

Dataset Status
Wikipedia-Swedish 🙋‍♀️ Waiting for contributor
Wikipedia-Danish 🙋‍♀️ Waiting for contributor
Wikipedia-Norwegian 🙋‍♀️ Waiting for contributor
Wikipedia-Finnish 🙋‍♀️ Waiting for contributor
Swedish Parliament 🙋‍♀️ Waiting for contributor

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

the-nordic-pile-0.0.2.tar.gz (5.9 kB view details)

Uploaded Source

Built Distribution

the_nordic_pile-0.0.2-py3-none-any.whl (6.3 kB view details)

Uploaded Python 3

File details

Details for the file the-nordic-pile-0.0.2.tar.gz.

File metadata

  • Download URL: the-nordic-pile-0.0.2.tar.gz
  • Upload date:
  • Size: 5.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.7.0 importlib_metadata/4.8.2 pkginfo/1.8.2 requests/2.26.0 requests-toolbelt/0.9.1 tqdm/4.62.3 CPython/3.7.12

File hashes

Hashes for the-nordic-pile-0.0.2.tar.gz
Algorithm Hash digest
SHA256 06b1bc5cccb58e4ea77a8e6928ebed3867e8f202af1efb4dda6be2a1c292df9a
MD5 175d98c1674347abe8389182ab019d4c
BLAKE2b-256 b094779fbec2c77aa48347e060d11b97c103596e302090740942bcde31552daf

See more details on using hashes here.

File details

Details for the file the_nordic_pile-0.0.2-py3-none-any.whl.

File metadata

  • Download URL: the_nordic_pile-0.0.2-py3-none-any.whl
  • Upload date:
  • Size: 6.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.7.0 importlib_metadata/4.8.2 pkginfo/1.8.2 requests/2.26.0 requests-toolbelt/0.9.1 tqdm/4.62.3 CPython/3.7.12

File hashes

Hashes for the_nordic_pile-0.0.2-py3-none-any.whl
Algorithm Hash digest
SHA256 16b4506c51f69a9e8220753ec325bcee312fe4be9750e4d1ba2811c5e7e24845
MD5 fd461bbe83b62967265ab2a5304758a8
BLAKE2b-256 4a37ef14f7915026c6c22691fc80dfb8a2cad4b2c705b6b673cc3fa8c921a0b6

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page