Word embedding: generic iterative stemmer
A generic helper for training gensim and fasttext word embedding models.
Specifically, this repository was created in order to
implement stemming
on a Wikipedia-based corpus in Hebrew, but it will probably also work for other
corpus sources and languages as well.
Important to note that while there are sophisticated and efficient approaches to the stemming task, this repository implements a naive approach with no strict time or memory considerations (more about that in the explanation section).
Based on https://github.com/liorshk/wordembedding-hebrew.
Setup
- Create a
python3virtual environment. - Install dependencies using
make install(this will run tests too).
Usage
The general flow is as follows:
- Get a text corpus (for example, from Wikipedia).
- Create a training program.
- Run a
StemmingTrainer.
The output of the training process is a generic_iterative_stemmer.models.StemmedKeyedVectors object
(in the form of a .kv file). It has the same interface as the standard gensim.models.KeyedVectors,
so the 2 can be used interchangeably.
0. (Optional) Set up a language data cache
generic_iterative_stemmer uses a language data cache to store its output and intermediate results.
The language data directory is useful if you want to train multiple models on the same corpus, or if you want to
train a model on a corpus that you've already trained on in the past, with different parameters.
To set up the language data cache, run mkdir -p ~/.cache/language_data.
Tip: soft-link the language data cache to your project's root directory,
e.g. ln -s ~/.cache/language_data language_data.
1. Get a text corpus
If you don't a specific corpus in mind, you can use Wikipedia. Here's how:
- Under
~/.cache/language_datafolder, create a directory for your corpus (for example,wiki-he). - Download Hebrew (or any other language) dataset from Wikipedia:
- Go to wikimedia dumps (in the URL, replace
hewith your language code). - Download the matching
wiki-latest-pages-articles.xml.bz2file, and place it in your corpus directory.
- Go to wikimedia dumps (in the URL, replace
- Create initial text corpus: run the script inside
notebooks/create_corpus.py(change parameters as needed).
This will create acorpus.txtfile in your corpus directory. It takes roughly 15 minutes to run (depending on the corpus size and your computer).
2. Create a training program
TODO
3. Run a StemmingTrainer
TODO
4. Play with your trained model
Play with your trained model using playground.ipynb.
Generic iterative stemming
TODO: Explain the algorithm.
Release files for generic-iterative-stemmer 1.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| generic_iterative_stemmer-1.2.0.tar.gz | 16.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| generic_iterative_stemmer-1.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 36.3 kB
Release files / generic_iterative_stemmer-1.2.0.tar.gz
| Download URL | generic_iterative_stemmer-1.2.0.tar.gz |
|---|---|
| Size | 16.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ce9c39bc73d1206037699afbb36fa00f66419936cafbde53f62f75cf3b29cfe8
|
|
BLAKE2b-256 checksum How to use checksums |
ff5fecbb3d56f7f78f6eb2b4e9d0b9bcdf9beb8dacfcc14e110fac6f9aa28848
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/5.1.1 CPython/3.12.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Nov 12, 2024.
Transparency logRelease files / generic_iterative_stemmer-1.2.0-py3-none-any.whl
| Download URL | generic_iterative_stemmer-1.2.0-py3-none-any.whl |
|---|---|
| Size | 20.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7325fd1eb2d03f6e4f6729316b253a72e1a508429cbacb4a78aceeb4852d20d3
|
|
BLAKE2b-256 checksum How to use checksums |
05d0422e766c324126a3996db70d82704aa9e9e66dd264a3c65d3d37ed80e635
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/5.1.1 CPython/3.12.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Nov 12, 2024.
Transparency log