A CLI for language-learning vocabulary extraction
Project description
Blitzer CLI
A minimalist command-line tool for language text processing
Author: Blitzer Team
Overview
Blitzer CLI is a command-line tool for language-learners that produces a word frequency list from input text. In addition to the simple word frequency list, the tool supports "lemmatization" for different languages. In simple terms, this means that, taking English as an example, the output can optionally treat "running" "ran" and "runs" as three instances of the same word as opposed to three different words.
Users can create and include their own word exclusion lists which will bar words from the result. The upfront cost of making and maintaining this exclusion list comes with the benefit of instant custom vocabulary for texts you intent to read, listen to, or study, as well as a way to track how many words you know in your language. Pretty cool!
The tool was designed for language learners and supports completely downloadable language pack plugins. There are no built-in lemmatization processors - all language processing capabilities are acquired using blitzer languages install <lang-code> command. Currently 2 languages are supported. Contributions for new languages are warmly accepted, and should be fairly straightforward to implement.
Installation
Using pip (recommended)
pip install -e .
Prerequisites
- Python 3.7+
clicklibrarytomlilibrary (for Python < 3.11)
Usage
Basic Usage
The tool reads text from stdin and outputs processed word lists to stdout using a subcommand architecture with flags:
# With stdin
echo "Your text here" | blitzer blitz -l [language_code] [flags]
# With direct text input
blitzer blitz -t "Your text here" -l [language_code] [flags]
Available Languages
base:: Basic processor for space-separated languages (no lemmatization support)- Downloadable language packs available via plugins (e.g.,
blitzer languages install plifor Pali,blitzer languages install slvfor Slovenian)
Available Commands
blitz:: Main command for processing text with configurable flagslist:: Lists supported languages for lemmatizationlanguages:: Manage language packs (install, uninstall, list)
Flags
-L,--lemmatize:: Treats different declensions/forms of the same word as one word-f,--freq:: Include frequency counts in the output-c,--context:: Include sample context for each word in output-p,--prompt:: Include custom prompt for LLM at the top of output-s,--src:: Include the full source text at the top of output-l,--language_code:: ISO 639 three-character language code-t,--text:: Direct text input (overrides stdin)-h,--help:: Show help message
Examples
# Basic word list
echo "This is a test." | blitzer blitz -l pli
# With frequency counts
echo "This is a test. This is only a test." | blitzer blitz -l pli -f
# With context sentences
echo "This is a test." | blitzer blitz -l pli -f -c
# Lemmatized output
echo "To je test." | blitzer blitz -l slv -L
# Using multiple flags
echo "This is a test." | blitzer blitz -l pli -L -f -c -p
# List available languages (plugins only)
blitzer languages list
# Manage language packs
blitzer languages list
blitzer languages install [lang-code]
blitzer languages uninstall [lang-code]
Configuration
The tool uses XDG specifications for configuration management:
Configuration Location
- Config file:
~/.config/blitzer/config.toml - Exclusion files:
~/.config/blitzer/[lang_code]_exclusion.txt(only location looked up)
Default Configuration
When the config file doesn't exist, it will be automatically created with these defaults:
# Blitzer CLI Configuration
# This file uses TOML format
# Default flag values
default_lemmatize = false # Default value for --lemmatize/-L flag
default_freq = false # Default value for --freq/-f flag
default_context = false # Default value for --context/-c flag
default_prompt = false # Default value for --prompt/-p flag
default_src = false # Default value for --src/-s flag
# Language-specific prompts
# Each key in the prompts table represents a language code with its custom prompt
[prompts]
"base" = "Convert the following wordlist into tab separated anki cards."
"en" = "Convert the following wordlist into tab separated anki cards."
Configuration defaults can be overridden with negative flags: --no-lemmatize, --no-freq, --no-context, --no-prompt, --no-src.
Exclusion Lists
Exclusion lists prevent known words from appearing in output. Language-specific exclusion lists are automatically created in the config directory when first accessing a language.
Language Support
Extending Language Support
The architecture uses a plugin system with entry points to make adding new languages much more dynamic and extensible. There is now only one way to add support for new languages:
For Plugin Languages (downloadable packages):
- Create a separate Python package with name format
blitzer-language-[lang_code] - Include your processor configuration in the package's
__init__.pyfile with aregister()function - Add an entry point in your package's
setup.pyorpyproject.tomlunderblitzer.languages - Bundle any required data files (like SQLite databases) with the package
- Distribute as a pip-installable package
Language Pack Management
The languages command allows for easy management of downloadable language packs:
# List all available languages (built-in and installed plugins)
blitzer languages list
# Install a language pack
blitzer languages install [lang-code]
# Uninstall a language pack
blitzer languages uninstall [lang-code]
Language Dictionaries
The tool supports language-specific dictionaries that enable lemmatization when using the -L flag. Lemmatization is the process of grouping together the different inflected forms of a word so they can be analyzed as a single item. For example, in Pali, both "deva" and "devo" would be mapped to the same root form "deva". Language dictionaries are stored in SQLite databases bundled with language plugins:
- Database location for plugins: Bundled with the plugin package
The tool looks for these databases in the installed plugins but does not create them automatically. When a language dictionary is not available, the tool falls back to basic word processing without lemmatization capabilities.
Current Language Support
base:: Basic processor for space-separated languages without lemmatization (always available)- Downloadable processors: Available as separate plugins (install with
blitzer languages install [lang-code])
Technical Architecture
Core Components
cli.py:: Command-line interface using Clickconfig.py:: XDG configuration managementprocessor.py:: Core text processing logicdata_manager.py:: Language data management utilities
Processing Pipeline
- Read text from stdin or from text argument
- Load appropriate language specification via entry points and register function
- Apply text normalization (if language-specific)
- Tokenize text into words or lemmatize as needed
- Apply exclusion filtering
- Format output according to flags
- Write results to stdout
Dependencies
click:: Command-line interface frameworktomli(ortomllibfor Python 3.11+) :: TOML configuration parsing
Development
Adding New Features
The architecture is designed for extensibility:
- Add new language processors in
languages/directory with get_processor function - Extend processing capabilities in
processor.py - Modify configuration schema in
config.py
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file blitzer_cli-0.1.0.tar.gz.
File metadata
- Download URL: blitzer_cli-0.1.0.tar.gz
- Upload date:
- Size: 26.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7d7adab8a03c6c9b76a4605bb2905b13d764c47ef468970c9a0a2e380aafed67
|
|
| MD5 |
cd2f3a061ba0907975a12b17fd2c33e9
|
|
| BLAKE2b-256 |
6c151e7b1699f8acb45c45cd44de800fcccec0af5efaeea9a1acd8f34b6d0833
|
File details
Details for the file blitzer_cli-0.1.0-py3-none-any.whl.
File metadata
- Download URL: blitzer_cli-0.1.0-py3-none-any.whl
- Upload date:
- Size: 27.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cc1a868958efe7a6f8d3691311871cfb25fa214548f76c43cfc7b4049168a7c8
|
|
| MD5 |
309386a7c2aa284511300dccfe0a84e7
|
|
| BLAKE2b-256 |
7cc352ebbdd7d5a5cbfca5c65a099b17e77d1bcd4a7c142a8981322d8cb6f643
|