lindera-python
Python binding for Lindera, a Japanese morphological analysis engine.
Overview
lindera-python provides a comprehensive Python interface to the Lindera 3.0.0 morphological analysis engine, supporting Japanese, Korean, and Chinese text analysis. This implementation includes all major features:
- Multi-language Support: Japanese (IPADIC, IPADIC-NEologd, UniDic), Korean (ko-dic), Chinese (CC-CEDICT, Jieba)
- Character Filters: Text preprocessing with mapping, regex, Unicode normalization, and Japanese iteration mark handling
- Token Filters: Post-processing filters including lowercase, length filtering, stop words, and Japanese-specific filters
- Flexible Configuration: Configurable tokenization modes and penalty settings
- Metadata Support: Complete dictionary schema and metadata management
Features
Core Components
- TokenizerBuilder: Fluent API for building customized tokenizers
- Tokenizer: High-performance text tokenization with integrated filtering
- CharacterFilter: Pre-processing filters for text normalization
- TokenFilter: Post-processing filters for token refinement
- Metadata & Schema: Dictionary structure and configuration management
- Training & Export (optional): Train custom morphological analysis models from corpus data
Supported Dictionaries
- Japanese: IPADIC, IPADIC-NEologd, UniDic
- Korean: ko-dic
- Chinese: CC-CEDICT, Jieba
- Custom: User dictionary support
Pre-built dictionaries are available from GitHub Releases.
Download a dictionary archive (e.g. lindera-ipadic-*.zip) and specify the extracted path when loading.
Filter Types
Character Filters:
- Mapping filter (character replacement)
- Regex filter (pattern-based replacement)
- Unicode normalization (NFKC, etc.)
- Japanese iteration mark normalization
Token Filters:
- Text case transformation (lowercase, uppercase)
- Length filtering (min/max character length)
- Stop words filtering
- Japanese-specific filters (base form, reading form, etc.)
- Korean-specific filters
Install project dependencies
- pyenv : https://github.com/pyenv/pyenv?tab=readme-ov-file#installation
- Poetry : https://python-poetry.org/docs/#installation
- Rust : https://www.rust-lang.org/tools/install
Install Python
# Install Python
% pyenv install 3.13.5
Setup repository and activate virtual environment
# Clone lindera project repository
% git clone git@github.com:lindera/lindera.git
% cd lindera
# Create Python virtual environment and initialize
% make init
# Activate Python virtual environment
% source .venv/bin/activate
Install lindera-python in the virtual environment
This command builds the library with development settings (debug build).
(.venv) % make python-develop
Quick Start
Basic Tokenization
from lindera.dictionary import load_dictionary
from lindera.tokenizer import Tokenizer
# Load dictionary from a local path (download from GitHub Releases)
dictionary = load_dictionary("/path/to/ipadic")
# Create a tokenizer
tokenizer = Tokenizer(dictionary, mode="normal")
# Tokenize Japanese text
text = "すもももももももものうち"
tokens = tokenizer.tokenize(text)
for token in tokens:
print(f"Text: {token.surface}, Position: {token.byte_start}-{token.byte_end}")
Converting Tokens to Plain Data
to_dict() returns the token as a plain dict, keeping each field's natural
Python type, so it serializes without a custom encoder:
import json
data = tokens[0].to_dict()
# {'surface': ..., 'byte_start': 0, 'byte_end': 9, 'position': 0,
# 'word_id': 12345, 'is_unknown': False, 'details': [...]}
json.dumps([token.to_dict() for token in tokens])
Using Character Filters
from lindera import TokenizerBuilder
# Create tokenizer builder
builder = TokenizerBuilder()
builder.set_mode("normal")
builder.set_dictionary("/path/to/ipadic")
# Add character filters
builder.append_character_filter("mapping", {"mapping": {"ー": "-"}})
builder.append_character_filter("unicode_normalize", {"kind": "nfkc"})
# Build tokenizer with filters
tokenizer = builder.build()
text = "テストー123"
tokens = tokenizer.tokenize(text) # Will apply filters automatically
Using Token Filters
from lindera import TokenizerBuilder
# Create tokenizer builder
builder = TokenizerBuilder()
builder.set_mode("normal")
builder.set_dictionary("/path/to/ipadic")
# Add token filters
builder.append_token_filter("lowercase")
builder.append_token_filter("length", {"min": 2, "max": 10})
builder.append_token_filter("japanese_stop_tags", {
"tags": ["助詞,格助詞,一般", "助詞,係助詞", "助詞,連体化", "助動詞"]
})
# Build tokenizer with filters
tokenizer = builder.build()
tokens = tokenizer.tokenize("テキストの解析")
Integrated Pipeline
from lindera import TokenizerBuilder
# Build tokenizer with integrated filters
builder = TokenizerBuilder()
builder.set_mode("normal")
builder.set_dictionary("/path/to/ipadic")
# Add character filters
builder.append_character_filter("mapping", {"mapping": {"ー": "-"}})
builder.append_character_filter("unicode_normalize", {"kind": "nfkc"})
# Add token filters
builder.append_token_filter("lowercase")
builder.append_token_filter("japanese_base_form")
# Build and use
tokenizer = builder.build()
tokens = tokenizer.tokenize("コーヒーショップ")
Working with Metadata
from lindera import Metadata
# Get metadata for a specific dictionary
metadata = Metadata.load("/path/to/ipadic")
print(f"Dictionary: {metadata.dictionary_name}")
print(f"Version: {metadata.dictionary_version}")
# Access schema information
schema = metadata.dictionary_schema
print(f"Schema has {len(schema.fields)} fields")
print(f"Fields: {schema.fields[:5]}") # First 5 fields
Advanced Usage
Filter Configuration Examples
Character filters and token filters accept configuration as dictionary arguments:
from lindera import TokenizerBuilder
builder = TokenizerBuilder()
builder.set_dictionary("/path/to/ipadic")
# Character filters with dict configuration
builder.append_character_filter("unicode_normalize", {"kind": "nfkc"})
builder.append_character_filter("japanese_iteration_mark", {
"normalize_kanji": "true",
"normalize_kana": "true"
})
builder.append_character_filter("mapping", {
"mapping": {"リンデラ": "lindera", "トウキョウ": "東京"}
})
# Token filters with dict configuration
builder.append_token_filter("japanese_katakana_stem", {"min": 3})
builder.append_token_filter("length", {"min": 2, "max": 10})
builder.append_token_filter("japanese_stop_tags", {
"tags": ["助詞,格助詞,一般", "助詞,係助詞", "助詞,連体化", "助動詞", "記号,句点", "記号,読点"]
})
# Filters without configuration can omit the dict
builder.append_token_filter("lowercase")
builder.append_token_filter("japanese_base_form")
tokenizer = builder.build()
See examples/ directory for comprehensive examples including:
tokenize.py: Basic tokenizationtokenize_with_filters.py: Using character and token filterstokenize_with_userdict.py: Custom user dictionarytrain_and_export.py: Train and export custom dictionaries (requirestrainfeature)- Multi-language tokenization
- Advanced configuration options
Dictionary Support
Japanese
- IPADIC: Default Japanese dictionary, good for general text
- UniDic: Academic dictionary with detailed morphological information
Korean
- ko-dic: Standard Korean dictionary for morphological analysis
Chinese
- CC-CEDICT: Community-maintained Chinese-English dictionary
Custom Dictionaries
- User dictionary support for domain-specific terms
- CSV format for easy customization
Dictionary Training (Experimental)
lindera-python supports training custom morphological analysis models from annotated corpus data when built with the train feature.
Building with Training Support
# Install with training support
(.venv) % maturin develop --features train
Training a Model
import lindera.trainer
# Train a model from corpus
lindera.trainer.train(
seed="path/to/seed.csv", # Seed lexicon
corpus="path/to/corpus.txt", # Training corpus
char_def="path/to/char.def", # Character definitions
unk_def="path/to/unk.def", # Unknown word definitions
feature_def="path/to/feature.def", # Feature templates
rewrite_def="path/to/rewrite.def", # Rewrite rules
output="model.dat", # Output model file
lambda_=0.01, # L1 regularization
max_iter=100, # Max iterations
max_threads=None # Auto-detect CPU cores
)
Exporting Dictionary Files
# Export trained model to dictionary files
lindera.trainer.export(
model="model.dat", # Trained model
output="exported_dict/", # Output directory
metadata="metadata.json" # Optional metadata file
)
This will create:
lex.csv: Lexicon filematrix.def: Connection cost matrixunk.def: Unknown word definitionschar.def: Character definitionsmetadata.json: Dictionary metadata (if provided)
See examples/train_and_export.py for a complete example.
API Reference
Core Classes
TokenizerBuilder: Fluent builder for tokenizer configurationTokenizer: Main tokenization engineToken: Individual token with text, position, and linguistic featuresCharacterFilter: Text preprocessing filtersTokenFilter: Token post-processing filtersMetadata: Dictionary metadata and configurationSchema: Dictionary schema definition
Training Functions (requires train feature)
train(): Train a morphological analysis model from corpusexport(): Export trained model to dictionary files
See the test_basic.py file for comprehensive API usage examples.
Release files for lindera 6.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| lindera-6.1.0.tar.gz | 497.9 kB | Details |
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| lindera-6.1.0-cp310-abi3-win_arm64.whl | CPython 3.10 | abi3 | Windows ARM64 | Details |
| lindera-6.1.0-cp310-abi3-win_amd64.whl | CPython 3.10 | abi3 | Windows x86-64 | Details |
| lindera-6.1.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl | CPython 3.10 | abi3 | Linux glibc 2.17+ x86-64 | Details |
| lindera-6.1.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl | CPython 3.10 | abi3 | Linux glibc 2.17+ ARM64 | Details |
| lindera-6.1.0-cp310-abi3-macosx_11_0_arm64.whl | CPython 3.10 | abi3 | macOS 11.0+ ARM64 | Details |
| lindera-6.1.0-cp310-abi3-macosx_10_12_x86_64.whl | CPython 3.10 | abi3 | macOS 10.12+ x86-64 | Details |
Total release size: 13.8 MB
Release files / lindera-6.1.0.tar.gz
| Download URL | lindera-6.1.0.tar.gz |
|---|---|
| Size | 497.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
758c43bec1eb3dc8cde7fc3d219b03d8108e71ed4f256333fac3b69c63a6a328
|
|
BLAKE2b-256 checksum How to use checksums |
ba5edc77b2c7732d41bebd286224cc5111955d4d6c888f082f0932ed46e3985d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.1.0-cp310-abi3-win_arm64.whl
| Download URL | lindera-6.1.0-cp310-abi3-win_arm64.whl |
|---|---|
| Size | 2.0 MB |
| Tags | CPython 3.10 Windows ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
4d71c35a2c8a8bd3b63127d16ac45c40102877312b99ef5eaac074dbb5ed24aa
|
|
BLAKE2b-256 checksum How to use checksums |
e34b8785502c0cdb0427e78c17088555ebe677ad08e56178f593824080fdee20
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.1.0-cp310-abi3-win_amd64.whl
| Download URL | lindera-6.1.0-cp310-abi3-win_amd64.whl |
|---|---|
| Size | 2.2 MB |
| Tags | CPython 3.10 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
5e41211a2cdaf7d79e0567816bd242ff73091bb8bfff2c2d0a5268573ce06ffc
|
|
BLAKE2b-256 checksum How to use checksums |
5ef7d1e2a5a6886959bb2a1acf9c22e07d77be12915cbab8ac31cd486aab91f4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.1.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
| Download URL | lindera-6.1.0-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl |
|---|---|
| Size | 2.4 MB |
| Tags | CPython 3.10 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
2918fd83ceafcdf8d8e5b234a4e80fdf9c5c63915f222038c0039e7dbd9c495b
|
|
BLAKE2b-256 checksum How to use checksums |
53387560638203b15fbc42554b8b525504a43dd321872e37a776b21588beb11a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.1.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
| Download URL | lindera-6.1.0-cp310-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl |
|---|---|
| Size | 2.3 MB |
| Tags | CPython 3.10 Linux glibc 2.17+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
67ebe384633261fb0318b7c16fbcde0b5e880271a7db09490a9930e4b075cd99
|
|
BLAKE2b-256 checksum How to use checksums |
7641726972219a4ee25035f97c48c72750ba83d86e8d6a6ab0ff5fccb79cb3fe
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.1.0-cp310-abi3-macosx_11_0_arm64.whl
| Download URL | lindera-6.1.0-cp310-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 2.2 MB |
| Tags | CPython 3.10 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
f83a8e59aa0e32857da92ef5f53f54ea790635e613aabce24d1930c03fa065a5
|
|
BLAKE2b-256 checksum How to use checksums |
cb3c7d726dd154ab1552d34e22204290ff7a657ac93e5c4c057a3d957881d703
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|
Release files / lindera-6.1.0-cp310-abi3-macosx_10_12_x86_64.whl
| Download URL | lindera-6.1.0-cp310-abi3-macosx_10_12_x86_64.whl |
|---|---|
| Size | 2.3 MB |
| Tags | CPython 3.10 abi3 macOS 10.12+ x86-64 |
|
SHA-256 checksum How to use checksums |
85c08f7d080ca19bc362c4003ea9933a2b402b7c949525743a1ba0519771d781
|
|
BLAKE2b-256 checksum How to use checksums |
a8a7d74410bc30f2cfdbb73bf269c850a8e1facacc8dd0ade1723babb96b235b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
maturin/1.15.0
|