fastcat
Fastcat is a little Python library for quickly looking up broader/narrower relations in Wikipedia categories locally. The idea is that fastcat can be useful in situations where you need to rapidly lookup category relations, but don't want to hammer on the Wikipedia API. Fastcat relies on Redis and the SKOS files that DBpedia makes available based on the Wikipedia MySQL dumps.
Attribution
This software is a fork of fastcat tool created by Ed Summers. Some changes were made under the Creative Commons Attribution-ShareAlike 3.0 license, and they are described in commit messages. Major changes are porting the code to Python 3 as well as adding support for more than one language.
Usage
Basic usage
The first time you import fastcat you'll need to populate your Redis database
with the category data from DBpedia. To do that instantiate a FastCat object
and call the load method. After that you can use it to do lookups.
>>> import fastcat
>>> f = fastcat.FastCat()
>>> f.load() # downloads the dump and loads it into redis (about a minute for English)
...
>>> print(f.broader("Computer programming"))
['Software engineering', 'Software development']
>>> print(f.narrower("Computer programming"))
['Programming languages', 'Algorithms', 'Data structures', 'Computer programming tools', 'Programming games', 'Programming paradigms', 'Anti-patterns', 'Software design patterns', 'Programming constructs', 'Programming contests', 'Concurrent computing', 'Source code', 'Debugging', 'Computer programmers', 'Programming idioms', 'Computer libraries', 'Self-hosting software', 'Programming principles', 'Software optimization', 'Computer programming books', 'Code refactoring', 'Live coding', 'Source code generation', 'Program derivation', 'Visual programming', 'Computer programming folklore']
Non-english categories
Just fill-in the language argument in the FastCat() constructor with a language code listed below.
>>> import fastcat
>>> f = fastcat.FastCat(language='de')
>>> f.load() # downloads the dump and loads it into redis (about a minute for English)
...
>>> print(f.broader("Berlin"))
['Europa nach Ort', 'Deutschland nach Gemeinde', 'Deutschland nach Bundesland']
>>> print(f.narrower("Berlin"))
['Umwelt- und Naturschutz (Berlin)', 'Veranstaltung (Berlin)', 'Stadtplanung (Berlin)', 'Verwaltung (Berlin)', 'Urbaner Freiraum in Berlin als Thema']
Currently supported languages (and their codes)
- English (
en) - Estonian (
et) - German (
de) - Japanese (
ja) - Polish (
pl) - Portuguese (
pt) - Russian (
ru) - Ukrainian (
ua) - Czech (
cs)
How long loading takes
load() pipelines its writes to Redis and streams the dump rather than holding
it in memory, so populating a language is quick and cheap:
| Language | Keys written | Load time | Peak memory |
|---|---|---|---|
| Czech | 151,020 | 4.1s | 52 MB |
| English | 1,838,205 | 58s | 55 MB |
(Measured on a laptop against a local Redis, dump already downloaded. Before
0.2.4 the same Czech load took 91.5s and English held 1.26 GB in memory.)
load(batch_size=...) controls how many Redis commands are buffered per
pipeline flush; the default of 10000 is a reasonable trade between speed and
memory.
Where the data comes from
The engine argument selects the download source:
>>> import fastcat
>>> f = fastcat.FastCat(engine='wiki-archive') # the default
>>> f.load()
| Engine | Status | What it downloads |
|---|---|---|
wiki-archive |
default | DBpedia's archived 2016-10 release |
databus |
not implemented yet | DBpedia's current Databus distribution |
Asking for databus raises NotImplementedError:
>>> fastcat.FastCat(engine='databus')
NotImplementedError: The 'databus' engine is not implemented yet. Use 'wiki-archive' instead (the default).
Note that the archived release is a snapshot of Wikipedia as it stood in
2016, so categories added since then are missing. That is the price of stable
URLs: the rolling current tree fastcat used before was withdrawn by DBpedia
and every URL under it now returns 404. Fetching today's data is what the
databus engine is for.
You can list the engines from code:
>>> fastcat.FastCat.get_supported_engines()
('wiki-archive', 'databus')
>>> fastcat.FastCat.get_implemented_engines()
('wiki-archive',)
Install
Redis installation
You first need to setup Redis server on your machine as follows.
On Mac:
$ brew install redis
On Linux:
$ sudo apt-get install redis-server
On Windows:
Please refer to instruction on installing Vagrant Redis. You will need an Ubuntu installation on your Windows, more information can be found here: Install your Linux Distribution of Choice
With Docker (any platform):
If you would rather not install Redis at all, the bundled compose file spins one
up on localhost:6379, with the loaded categories kept in a named volume so
they survive a restart:
$ docker compose up -d redis
The same file also defines a fastcat container with the package and its dev
dependencies installed, which is handy for running the suite in a clean
environment:
$ docker compose run --rm fastcat pytest
Inside a container Redis is not on localhost, so fastcat reads the
FASTCAT_REDIS_HOST and FASTCAT_REDIS_PORT environment variables (already set
for the fastcat service) to find it.
Installing the module
If you are ready, installing Fastcat is pretty straightforward:
$ pip install fastcat
Or if you wish to get the newest dev code:
$ pip install git+https://github.com/oskar-j/fastcat.git
That's it!
Contributing to the project
Guidelines
See CONTRIBUTING.md for more details
Running unit tests
With uv, which installs the exact versions from
the committed uv.lock:
$ uv sync
$ uv run pytest
Or with pip:
$ pip install -e '.[dev]'
$ pytest
That runs the fast, offline tests. The end-to-end tests need a Redis server and download a SKOS dump per language from DBpedia, so they are opt-in:
$ docker compose up -d redis # or your own local Redis
$ pytest --run-integration
Q&A
How much is is tested?
It's still in early stage of development, please share some feedback with me (under the ticket #7).
What are biggest drawbacks of Fastcat?
DBpedia SKOS files move around, and that has already bitten this project once: the rolling current tree fastcat
downloaded from was withdrawn, and downloading Wikipedia data stopped working entirely until 0.2.3 repointed it at
the archived 2016-10 release. Pinning to an archive buys stable URLs at the cost of data that stops in 2016 --
until the databus engine lands, categories created after that are simply not there. Moreover, due to the
infrastructure of Redis, you can have a maximum number
of 16 languages (1 slot for a language). Last but not least, it takes around 40 MB of your web transfer (size depends
on the selected language) to download a single SKOS file.
Which Python versions are supported?
Python 3.10 and above (tested on GitHub Actions against 3.10 through 3.14).
Releases up to 0.1.2 supported Python 3.5+; if you are stuck on an older
interpreter, pin fastcat==0.1.2.
Which languages are supported?
There are two ways to check the list of available languages.
First, is a manual inspection of the lang.py file.
Second way is to call the get_supported_languages() method on the FastCat object.
What's coming next?
Implementing the databus engine, so fastcat can pull current categories instead of a 2016 snapshot.
Support for the rest of european languages. Exporting n-size tree of categories to a CSV or GraphML file.
Moving the downloaded dumps and the language mapping out of the package directory into a proper user cache
directory.
License
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file fastcat-0.2.4.tar.gz.
File metadata
- Download URL: fastcat-0.2.4.tar.gz
- Upload date:
- Size: 28.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
765b23cd9c276a57dfdc5b964c7d6dd1fb45261dae585ce7ac4d2d6628be1953
|
|
| MD5 |
17cf70bb94241a6bac12ffc937d603e3
|
|
| BLAKE2b-256 |
b8af27c5fc223887ec88361f4cbc87ed9e997466b6426128f7284a1dc9ad6940
|
Provenance
The following attestation bundles were made for fastcat-0.2.4.tar.gz:
Publisher:
publish.yml on oskar-j/fastcat
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
fastcat-0.2.4.tar.gz -
Subject digest:
765b23cd9c276a57dfdc5b964c7d6dd1fb45261dae585ce7ac4d2d6628be1953 - Sigstore transparency entry: 2809417166
- Sigstore integration time:
-
Permalink:
oskar-j/fastcat@281fdaaadcc40ccd48ab66146d1a06ae461de88f -
Branch / Tag:
refs/heads/master - Owner: https://github.com/oskar-j
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@281fdaaadcc40ccd48ab66146d1a06ae461de88f -
Trigger Event:
push
-
Statement type:
File details
Details for the file fastcat-0.2.4-py3-none-any.whl.
File metadata
- Download URL: fastcat-0.2.4-py3-none-any.whl
- Upload date:
- Size: 20.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
94af4e98644be9d39d2bd14efb7196c7def3c349ab1431c88d9e0f6cbac9c2fe
|
|
| MD5 |
e91f69567be7ebf472eccc7c98317923
|
|
| BLAKE2b-256 |
2c465a67eb44c7fd12a8e4a13a42090e509205bb3ca8bc59d21e783433897c2a
|
Provenance
The following attestation bundles were made for fastcat-0.2.4-py3-none-any.whl:
Publisher:
publish.yml on oskar-j/fastcat
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
fastcat-0.2.4-py3-none-any.whl -
Subject digest:
94af4e98644be9d39d2bd14efb7196c7def3c349ab1431c88d9e0f6cbac9c2fe - Sigstore transparency entry: 2809417218
- Sigstore integration time:
-
Permalink:
oskar-j/fastcat@281fdaaadcc40ccd48ab66146d1a06ae461de88f -
Branch / Tag:
refs/heads/master - Owner: https://github.com/oskar-j
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@281fdaaadcc40ccd48ab66146d1a06ae461de88f -
Trigger Event:
push
-
Statement type: