Skip to main content

Wiki-Dump Reader

Travis Coverage

Extract corpora from wiki-dump.

Install

pip install wiki-dump-reader

Usage

The dump file *wiki-*-pages-articles.xml should be downloaded first. Then you can iterate and get cleaned text from the text:

from wiki_dump_reader import Cleaner, iterate

cleaner = Cleaner()
for title, text in iterate('*wiki-*-pages-articles.xml'):
    text = cleaner.clean_text(text)
    cleaned_text, links = cleaner.build_links(text)

Just ignore links if you don't need them:

cleaned_text, _ = cleaner.build_links(text)

See examples for an intuitive feeling.

Metadata

Release files for wiki-dump-reader 0.0.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for wiki-dump-reader 0.0.4
File Size Uploaded
wiki-dump-reader-0.0.4.tar.gz 3.4 kB Details

Release files / wiki-dump-reader-0.0.4.tar.gz

Download URL wiki-dump-reader-0.0.4.tar.gz
Size 3.4 kB
Tags Source
SHA-256 checksum
How to use checksums
86532997c6870b46182eed6c461049dfebaea37d6f59f90d7ff5bcd4d85db04d
BLAKE2b-256 checksum
How to use checksums
ab79e70b9c27a3038bad28448e8183ee59a248d968aeb942dff94d61dcf10c45
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/1.12.1 pkginfo/1.4.2 requests/2.20.1 setuptools/40.7.1 requests-toolbelt/0.8.0 tqdm/4.24.0 CPython/3.6.4

Release history Release notifications | RSS feed

This release

0.0.4 This release

1 release file

0.0.3

1 release file

0.0.2

1 release file

0.0.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page