lindera-python
Python binding for Lindera, a Japanese morphological analysis engine.
Overview
lindera-python provides a comprehensive Python interface to the Lindera 3.0.0 morphological analysis engine, supporting Japanese, Korean, and Chinese text analysis. This implementation includes all major features:
- Multi-language Support: Japanese (IPADIC, IPADIC-NEologd, UniDic), Korean (ko-dic), Chinese (CC-CEDICT, Jieba)
- Character Filters: Text preprocessing with mapping, regex, Unicode normalization, and Japanese iteration mark handling
- Token Filters: Post-processing filters including lowercase, length filtering, stop words, and Japanese-specific filters
- Flexible Configuration: Configurable tokenization modes and penalty settings
- Metadata Support: Complete dictionary schema and metadata management
Features
Core Components
- TokenizerBuilder: Fluent API for building customized tokenizers
- Tokenizer: High-performance text tokenization with integrated filtering
- CharacterFilter: Pre-processing filters for text normalization
- TokenFilter: Post-processing filters for token refinement
- Metadata & Schema: Dictionary structure and configuration management
- Training & Export (optional): Train custom morphological analysis models from corpus data
Supported Dictionaries
- Japanese: IPADIC, IPADIC-NEologd, UniDic
- Korean: ko-dic
- Chinese: CC-CEDICT, Jieba
- Custom: User dictionary support
Pre-built dictionaries are available from GitHub Releases.
Download a dictionary archive (e.g. lindera-ipadic-*.zip) and specify the extracted path when loading.
Filter Types
Character Filters:
- Mapping filter (character replacement)
- Regex filter (pattern-based replacement)
- Unicode normalization (NFKC, etc.)
- Japanese iteration mark normalization
Token Filters:
- Text case transformation (lowercase, uppercase)
- Length filtering (min/max character length)
- Stop words filtering
- Japanese-specific filters (base form, reading form, etc.)
- Korean-specific filters
Install project dependencies
- pyenv : https://github.com/pyenv/pyenv?tab=readme-ov-file#installation
- Poetry : https://python-poetry.org/docs/#installation
- Rust : https://www.rust-lang.org/tools/install
Install Python
# Install Python
% pyenv install 3.13.5
Setup repository and activate virtual environment
# Clone lindera project repository
% git clone git@github.com:lindera/lindera.git
% cd lindera
# Create Python virtual environment and initialize
% make init
# Activate Python virtual environment
% source .venv/bin/activate
Install lindera-python in the virtual environment
This command builds the library with development settings (debug build).
(.venv) % make python-develop
Quick Start
Basic Tokenization
from lindera.dictionary import load_dictionary
from lindera.tokenizer import Tokenizer
# Load dictionary from a local path (download from GitHub Releases)
dictionary = load_dictionary("/path/to/ipadic")
# Create a tokenizer
tokenizer = Tokenizer(dictionary, mode="normal")
# Tokenize Japanese text
text = "すもももももももものうち"
tokens = tokenizer.tokenize(text)
for token in tokens:
print(f"Text: {token.surface}, Position: {token.byte_start}-{token.byte_end}")
Converting Tokens to Plain Data
to_dict() returns the token as a plain dict, keeping each field's natural
Python type, so it serializes without a custom encoder:
import json
data = tokens[0].to_dict()
# {'surface': ..., 'byte_start': 0, 'byte_end': 9, 'position': 0,
# 'word_id': 12345, 'is_unknown': False, 'details': [...]}
json.dumps([token.to_dict() for token in tokens])
Using Character Filters
from lindera import TokenizerBuilder
# Create tokenizer builder
builder = TokenizerBuilder()
builder.set_mode("normal")
builder.set_dictionary("/path/to/ipadic")
# Add character filters
builder.append_character_filter("mapping", {"mapping": {"ー": "-"}})
builder.append_character_filter("unicode_normalize", {"kind": "nfkc"})
# Build tokenizer with filters
tokenizer = builder.build()
text = "テストー123"
tokens = tokenizer.tokenize(text) # Will apply filters automatically
Using Token Filters
from lindera import TokenizerBuilder
# Create tokenizer builder
builder = TokenizerBuilder()
builder.set_mode("normal")
builder.set_dictionary("/path/to/ipadic")
# Add token filters
builder.append_token_filter("lowercase")
builder.append_token_filter("length", {"min": 2, "max": 10})
builder.append_token_filter("japanese_stop_tags", {
"tags": ["助詞,格助詞,一般", "助詞,係助詞", "助詞,連体化", "助動詞"]
})
# Build tokenizer with filters
tokenizer = builder.build()
tokens = tokenizer.tokenize("テキストの解析")
Integrated Pipeline
from lindera import TokenizerBuilder
# Build tokenizer with integrated filters
builder = TokenizerBuilder()
builder.set_mode("normal")
builder.set_dictionary("/path/to/ipadic")
# Add character filters
builder.append_character_filter("mapping", {"mapping": {"ー": "-"}})
builder.append_character_filter("unicode_normalize", {"kind": "nfkc"})
# Add token filters
builder.append_token_filter("lowercase")
builder.append_token_filter("japanese_base_form")
# Build and use
tokenizer = builder.build()
tokens = tokenizer.tokenize("コーヒーショップ")
Working with Metadata
from lindera import Metadata
# Get metadata for a specific dictionary
metadata = Metadata.load("/path/to/ipadic")
print(f"Dictionary: {metadata.dictionary_name}")
print(f"Version: {metadata.dictionary_version}")
# Access schema information
schema = metadata.dictionary_schema
print(f"Schema has {len(schema.fields)} fields")
print(f"Fields: {schema.fields[:5]}") # First 5 fields
Advanced Usage
Filter Configuration Examples
Character filters and token filters accept configuration as dictionary arguments:
from lindera import TokenizerBuilder
builder = TokenizerBuilder()
builder.set_dictionary("/path/to/ipadic")
# Character filters with dict configuration
builder.append_character_filter("unicode_normalize", {"kind": "nfkc"})
builder.append_character_filter("japanese_iteration_mark", {
"normalize_kanji": "true",
"normalize_kana": "true"
})
builder.append_character_filter("mapping", {
"mapping": {"リンデラ": "lindera", "トウキョウ": "東京"}
})
# Token filters with dict configuration
builder.append_token_filter("japanese_katakana_stem", {"min": 3})
builder.append_token_filter("length", {"min": 2, "max": 10})
builder.append_token_filter("japanese_stop_tags", {
"tags": ["助詞,格助詞,一般", "助詞,係助詞", "助詞,連体化", "助動詞", "記号,句点", "記号,読点"]
})
# Filters without configuration can omit the dict
builder.append_token_filter("lowercase")
builder.append_token_filter("japanese_base_form")
tokenizer = builder.build()
See examples/ directory for comprehensive examples including:
tokenize.py: Basic tokenizationtokenize_with_filters.py: Using character and token filterstokenize_with_userdict.py: Custom user dictionarytrain_and_export.py: Train and export custom dictionaries (requirestrainfeature)- Multi-language tokenization
- Advanced configuration options
Dictionary Support
Japanese
- IPADIC: Default Japanese dictionary, good for general text
- UniDic: Academic dictionary with detailed morphological information
Korean
- ko-dic: Standard Korean dictionary for morphological analysis
Chinese
- CC-CEDICT: Community-maintained Chinese-English dictionary
Custom Dictionaries
- User dictionary support for domain-specific terms
- CSV format for easy customization
Dictionary Training (Experimental)
lindera-python supports training custom morphological analysis models from annotated corpus data when built with the train feature.
Building with Training Support
# Install with training support
(.venv) % maturin develop --features train
Training a Model
import lindera.trainer
# Train a model from corpus
lindera.trainer.train(
seed="path/to/seed.csv", # Seed lexicon
corpus="path/to/corpus.txt", # Training corpus
char_def="path/to/char.def", # Character definitions
unk_def="path/to/unk.def", # Unknown word definitions
feature_def="path/to/feature.def", # Feature templates
rewrite_def="path/to/rewrite.def", # Rewrite rules
output="model.dat", # Output model file
lambda_=0.01, # L1 regularization
max_iter=100, # Max iterations
max_threads=None # Auto-detect CPU cores
)
Exporting Dictionary Files
# Export trained model to dictionary files
lindera.trainer.export(
model="model.dat", # Trained model
output="exported_dict/", # Output directory
metadata="metadata.json" # Optional metadata file
)
This will create:
lex.csv: Lexicon filematrix.def: Connection cost matrixunk.def: Unknown word definitionschar.def: Character definitionsmetadata.json: Dictionary metadata (if provided)
See examples/train_and_export.py for a complete example.
API Reference
Core Classes
TokenizerBuilder: Fluent builder for tokenizer configurationTokenizer: Main tokenization engineToken: Individual token with text, position, and linguistic featuresCharacterFilter: Text preprocessing filtersTokenFilter: Token post-processing filtersMetadata: Dictionary metadata and configurationSchema: Dictionary schema definition
Training Functions (requires train feature)
train(): Train a morphological analysis model from corpusexport(): Export trained model to dictionary files
See the test_basic.py file for comprehensive API usage examples.
Release files for lindera 6.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| lindera-6.2.0.tar.gz | 501.0 kB | Details |
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| lindera-6.2.0-cp310-abi3-win_arm64.whl | CPython 3.10 | abi3 | Windows ARM64 | Details |
| lindera-6.2.0-cp310-abi3-win_amd64.whl | CPython 3.10 | abi3 | Windows x86-64 | Details |
| lindera-6.2.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl | CPython 3.10 | abi3 | Linux glibc 2.17+ x86-64 | Details |
| lindera-6.2.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl | CPython 3.10 | abi3 | Linux glibc 2.17+ ARM64 | Details |
| lindera-6.2.0-cp310-abi3-macosx_11_0_arm64.whl | CPython 3.10 | abi3 | macOS 11.0+ ARM64 | Details |
| lindera-6.2.0-cp310-abi3-macosx_10_12_x86_64.whl | CPython 3.10 | abi3 | macOS 10.12+ x86-64 | Details |
Total release size: 13.9 MB
Release files / lindera-6.2.0.tar.gz
| Download URL | lindera-6.2.0.tar.gz |
|---|---|
| Size | 501.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
70c52d1192412bbe00c1b271fb9d23086c1c0b0637a22d8506e62973c780f981
|
|
BLAKE2b-256 checksum How to use checksums |
668db1b8cb97691c4a050879a4e90a50ff91687b94c220266e1d311a01cc10c4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.2.0-cp310-abi3-win_arm64.whl
| Download URL | lindera-6.2.0-cp310-abi3-win_arm64.whl |
|---|---|
| Size | 2.0 MB |
| Tags | CPython 3.10 Windows ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
c57669d2f1f256e6812b4441eb4da014cba9a3784791952ce5fa0a29fc01b6fe
|
|
BLAKE2b-256 checksum How to use checksums |
13e4d706ada7e24053e9ec447ba52a9061af2b269a6fc1f33cb5aad460db8446
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.2.0-cp310-abi3-win_amd64.whl
| Download URL | lindera-6.2.0-cp310-abi3-win_amd64.whl |
|---|---|
| Size | 2.2 MB |
| Tags | CPython 3.10 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
d39584d33a9aee00d9ea70521b4f8635798e2740c2943ec2c4aa6600a1836b7c
|
|
BLAKE2b-256 checksum How to use checksums |
b1cac17288a0b67d5614d4b2b85ae5dcd7d05d15ed09e1ef4261a84a68c3d754
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.2.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
| Download URL | lindera-6.2.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl |
|---|---|
| Size | 2.4 MB |
| Tags | CPython 3.10 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
7a7de3c2cd9b96cc1d8685fb8eac7fea83e09b78d36ab2c3b4fe59c70e9bce92
|
|
BLAKE2b-256 checksum How to use checksums |
e11fd17e23ebd6b84642fb0a3ec3d25a61e9cd2fa24e60871d36b384d8f75767
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.2.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
| Download URL | lindera-6.2.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl |
|---|---|
| Size | 2.3 MB |
| Tags | CPython 3.10 Linux glibc 2.17+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
be2555f78559e4e2c6a9634853a8e9a857f9487fa1578579655ebe43b4152b68
|
|
BLAKE2b-256 checksum How to use checksums |
a641a8df1d450eac6fe54e56157a62c6add2ba801329a2dff8df8c8697f4e886
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.2.0-cp310-abi3-macosx_11_0_arm64.whl
| Download URL | lindera-6.2.0-cp310-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 2.2 MB |
| Tags | CPython 3.10 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
e4e7635563ef816852219906545076f2b465555553111f7c74ffba42bf361afb
|
|
BLAKE2b-256 checksum How to use checksums |
ee56378c0f5595673564c602917c12ab0cb466b368607c3cf36e63bd4651b636
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.2.0-cp310-abi3-macosx_10_12_x86_64.whl
| Download URL | lindera-6.2.0-cp310-abi3-macosx_10_12_x86_64.whl |
|---|---|
| Size | 2.3 MB |
| Tags | CPython 3.10 abi3 macOS 10.12+ x86-64 |
|
SHA-256 checksum How to use checksums |
7d60ccc7e052458b208168b0458f9842a89fafdf420a72511330fa51c4cc58a8
|
|
BLAKE2b-256 checksum How to use checksums |
3550759a510bbee31226cf0bf05ab97f2f9aa0224e7b5dee10b1326a930a83d7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|