The streaming access to the Google ngram data.
The data is so big, that storing it is almost impossible. However, sometimes you need an aggregate data over the dataset. For example to build a co-occurrence matrix.
This package provides an iterator over the dataset stored at Google. It decompresses the data on the fly and provides you the access to the underlying data.
>>> from google_ngram_downloader import readline_google_store >>> >>> fname, url, records = next(readline_google_store(ngram_len=5)) >>> fname 'googlebooks-eng-all-5gram-20120701-0.gz' >>> url 'http://storage.googleapis.com/books/ngrams/books/googlebooks-eng-all-5gram-20120701-0.gz' >>> next(records) Record(ngram=u'0 " A most useful', year=1860, match_count=1, volume_count=1)
pip install google-ngram-downloader
The command line tool
It also provides a simple command line tool to download the ngrams called google-ngram-downloader.
- Non-unique contexts are taken into account inside of an ngram.
- The cooccurrence command does not perform any ngram modification.
- download, readile and cooccurrence subcommands.
- readline_google_store transforms lines to Record in several processes.
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Hashes for google-ngram-downloader-3.1.1.tar.gz