Skip to main content

fastcat

Tests Publish PyPI Python Versions Downloads License Pending Pull-Requests Github Issues Commits Since Release

Fastcat is a little Python library for quickly looking up broader/narrower relations in Wikipedia categories locally. The idea is that fastcat can be useful in situations where you need to rapidly lookup category relations, but don't want to hammer on the Wikipedia API. Fastcat relies on Redis and the SKOS files that DBpedia makes available based on the Wikipedia MySQL dumps.

fastcat logo

Attribution

This software is a fork of fastcat tool created by Ed Summers. Some changes were made under the Creative Commons Attribution-ShareAlike 3.0 license, and they are described in commit messages. Major changes are porting the code to Python 3 as well as adding support for more than one language.

Usage

Basic usage

The first time you import fastcat you'll need to populate your Redis database with the category data from DBpedia. To do that instantiate a FastCat object and call the load method. After that you can use it to do lookups.

>>> import fastcat
>>> f = fastcat.FastCat()
>>> f.load()  # downloads the dump and loads it into redis (about a minute for English)
...
>>> print(f.broader("Computer programming"))
['Software engineering', 'Software development']
>>> print(f.narrower("Computer programming"))
['Programming languages', 'Algorithms', 'Data structures', 'Computer programming tools', 'Programming games', 'Programming paradigms', 'Anti-patterns', 'Software design patterns', 'Programming constructs', 'Programming contests', 'Concurrent computing', 'Source code', 'Debugging', 'Computer programmers', 'Programming idioms', 'Computer libraries', 'Self-hosting software', 'Programming principles', 'Software optimization', 'Computer programming books', 'Code refactoring', 'Live coding', 'Source code generation', 'Program derivation', 'Visual programming', 'Computer programming folklore']

Non-english categories

Just fill-in the language argument in the FastCat() constructor with a language code listed below.

>>> import fastcat
>>> f = fastcat.FastCat(language='de')
>>> f.load()  # downloads the dump and loads it into redis (about a minute for English)
...
>>> print(f.broader("Berlin"))
['Europa nach Ort', 'Deutschland nach Gemeinde', 'Deutschland nach Bundesland']
>>> print(f.narrower("Berlin"))
['Umwelt- und Naturschutz (Berlin)', 'Veranstaltung (Berlin)', 'Stadtplanung (Berlin)', 'Verwaltung (Berlin)', 'Urbaner Freiraum in Berlin als Thema']
Currently supported languages (and their codes)
  1. English (en)
  2. Estonian (et)
  3. German (de)
  4. Japanese (ja)
  5. Polish (pl)
  6. Portuguese (pt)
  7. Russian (ru)
  8. Ukrainian (ua)
  9. Czech (cs)

How long loading takes

load() pipelines its writes to Redis and streams the dump rather than holding it in memory, so populating a language is quick and cheap:

Language Keys written Load time Peak memory
Czech 151,020 4.1s 52 MB
English 1,838,205 58s 55 MB

(Measured on a laptop against a local Redis, dump already downloaded. Before 0.2.4 the same Czech load took 91.5s and English held 1.26 GB in memory.)

load(batch_size=...) controls how many Redis commands are buffered per pipeline flush; the default of 10000 is a reasonable trade between speed and memory.

Where the data comes from

The engine argument selects the download source:

>>> import fastcat
>>> f = fastcat.FastCat(engine='wiki-archive')  # the default
>>> f.load()
Engine Status What it downloads
wiki-archive default DBpedia's archived 2016-10 release -- fixed URLs, data frozen in 2016
databus available DBpedia's current Databus release (2022.12.01) -- current data, resolved through metadata

For up-to-date categories, ask for the Databus:

>>> f = fastcat.FastCat(engine='databus')
>>> f.load(language='et')

The Databus has no fixed download URLs, so fastcat resolves the newest version of the generic/categories artifact and picks the SKOS part for your language. It also publishes a sha256 per file, which fastcat checks after downloading; a mismatch discards the file rather than loading it.

The two engines are not interchangeable

They are different snapshots of Wikipedia, six years apart, and the categories really do differ:

>>> fastcat.FastCat(engine='wiki-archive').broader("Tallinn")   # 2016
['Eesti linnad', 'Harju maakonna omavalitsused', 'Hansalinnad']
>>> fastcat.FastCat(engine='databus').broader("Tallinn")        # 2022
['Eesti asulate kategooriad', 'Harju maakonna omavalitsusüksuste kategooriad']

Estonian is 31,678 keys from the archive and 41,254 from the Databus. Because mixing them would produce a set of relations belonging to neither snapshot, fastcat records which engine filled a Redis db and refuses to load the other on top:

>>> f = fastcat.FastCat(engine='databus')
>>> f.load(language='et')
>>> f.load(language='et', engine='wiki-archive')
RuntimeError: Language 'et' is already loaded in this redis db from the 'databus'
engine, so loading 'wiki-archive' data would merge two different snapshots.
Flush this db first (e.g. FastCat.db.flushdb()) or point this language at another db.

wiki-archive stays the default: it needs no metadata lookup, is unaffected by the Databus being down, and leaves existing installs on the data they already have.

You can list the engines from code:

>>> fastcat.FastCat.get_supported_engines()
('wiki-archive', 'databus')
>>> fastcat.FastCat.get_implemented_engines()
('wiki-archive', 'databus')

Install

Redis installation

You first need to setup Redis server on your machine as follows.

On Mac:

$ brew install redis

On Linux:

$ sudo apt-get install redis-server

On Windows:

Please refer to instruction on installing Vagrant Redis. You will need an Ubuntu installation on your Windows, more information can be found here: Install your Linux Distribution of Choice

With Docker (any platform):

If you would rather not install Redis at all, the bundled compose file spins one up on localhost:6379, with the loaded categories kept in a named volume so they survive a restart:

$ docker compose up -d redis

The same file also defines a fastcat container with the package and its dev dependencies installed, which is handy for running the suite in a clean environment:

$ docker compose run --rm fastcat pytest

Inside a container Redis is not on localhost, so fastcat reads the FASTCAT_REDIS_HOST and FASTCAT_REDIS_PORT environment variables (already set for the fastcat service) to find it.

Installing the module

If you are ready, installing Fastcat is pretty straightforward:

$ pip install fastcat

Or if you wish to get the newest dev code:

$ pip install git+https://github.com/oskar-j/fastcat.git

That's it!

Contributing to the project

Guidelines

See CONTRIBUTING.md for more details

Running unit tests

With uv, which installs the exact versions from the committed uv.lock:

$ uv sync
$ uv run pytest

Or with pip:

$ pip install -e '.[dev]'
$ pytest

That runs the fast, offline tests. The end-to-end tests need a Redis server and download a SKOS dump per language from DBpedia, so they are opt-in:

$ docker compose up -d redis      # or your own local Redis
$ pytest --run-integration

Q&A

How much is is tested?

It's still in early stage of development, please share some feedback with me (under the ticket #7).

What are biggest drawbacks of Fastcat?

DBpedia SKOS files move around, and that has already bitten this project once: the rolling current tree fastcat downloaded from was withdrawn, and downloading Wikipedia data stopped working entirely until 0.2.3 repointed it at the archived 2016-10 release. Pinning to an archive buys stable URLs at the cost of data that stops in 2016 -- so pass engine='databus' for current categories. Moreover, due to the infrastructure of Redis, you can have a maximum number of 16 languages (1 slot for a language). Last but not least, it takes around 40 MB of your web transfer (size depends on the selected language) to download a single SKOS file.

Which Python versions are supported?

Python 3.10 and above (tested on GitHub Actions against 3.10 through 3.14). Releases up to 0.1.2 supported Python 3.5+; if you are stuck on an older interpreter, pin fastcat==0.1.2.

Which languages are supported?

There are two ways to check the list of available languages.

First, is a manual inspection of the lang.py file.

Second way is to call the get_supported_languages() method on the FastCat object.

What's coming next?

Support for the rest of european languages. Exporting n-size tree of categories to a CSV or GraphML file. Moving the downloaded dumps and the language mapping out of the package directory into a proper user cache directory.

License

Creative Commons Attribution-ShareAlike 3.0

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

fastcat-0.3.0.tar.gz (33.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

fastcat-0.3.0-py3-none-any.whl (23.6 kB view details)

Uploaded Python 3

File details

Details for the file fastcat-0.3.0.tar.gz.

File metadata

  • Download URL: fastcat-0.3.0.tar.gz
  • Upload date:
  • Size: 33.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for fastcat-0.3.0.tar.gz
Algorithm Hash digest
SHA256 696ab6d5924aa200df3c73ad2d90e31fecdb7cc54dc78c14bb163c90b0bcd33c
MD5 4a0da41280270ed9f456ab4a942f3865
BLAKE2b-256 84f6a0cd4d7969294cc9b19ae143a98e16990f2a7f69afb4dc4b21b27ec58033

See more details on using hashes here.

Provenance

The following attestation bundles were made for fastcat-0.3.0.tar.gz:

Publisher: publish.yml on oskar-j/fastcat

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file fastcat-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: fastcat-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 23.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for fastcat-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a49612a056b1f8fd529828cc8ad05508c1379855314009e4a3431e2c3c4fa94c
MD5 b37e1390c01f273b898153e562b8d44c
BLAKE2b-256 a015b75a4cd16cebf5b7502fda312f222a99348d27f18f8315f9a0257b261159

See more details on using hashes here.

Provenance

The following attestation bundles were made for fastcat-0.3.0-py3-none-any.whl:

Publisher: publish.yml on oskar-j/fastcat

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 files

0.2.5

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

0.2.1

2 files

0.1.2

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page