Skip to main content

Lucytok

Lucene's boring English tokenizers recreated for Python. Compatible with SearchArray.

Lets you configure a handful of normal tokenization rules like ascii folding, posessive removal, both types of porter stemming, English stopwords, etc.

Usage

Creating a tokenizer close to Elasticsearch's default english analyzer

from lucytok import english
es_english = english("Nsp->NNN->l->NsNN->1")
tokenized = es_english("The quick brown fox jumps over the lazy døg")
print(tokenized)

Outputs

['_', 'quick', 'brown', 'fox', 'jump', 'over', '_', 'lazi', 'døg']

Make a tokenizer with ASCII folding...

from lucytok import english
es_english_folded = english("asp->NNN->l->NsNN->1")
print(es_english_folded("The quick brown fox jumps over the lazy døg"))
['_', 'quick', 'brown', 'fox', 'jump', 'over', '_', 'lazi', 'dog']

Split compounds and convert British to American spelling...

from lucytok import english
es_british = english("asp->NNN->l->cbsN->1")
print(es_british("The watercolour fox jumps over the lazy døg"))
['_', 'water', 'color', 'fox', 'jump', 'over', '_', 'lazi', 'dog']

Spec

Create a tokenizer using the following settings (these concepts correspond to their Elasticsearch counterparts):


#  |- ASCII fold (a) or not (N)
#  ||- Standard (s) or WS tokenizer (w)
#  ||- Remove possessive suffixes (p) or not (N)
#  |||
# "NsN->NNN->N->NNNN->N"
#       |||  |  ||||  |
#       |||  |  ||||  |- Porter stem version (1) or version (2) vs N/0 for none
#       |||  |  ||||- Manually convert irregular plurals (p) or not (N)
#       |||  |  |||- Blank out stopwords (s) or not (N)
#       |||  |  ||- Convert british to american spelling (b) or not (N)
#       |||  |  |- Split Compounds (c) or not (N)
#       |||  |- Lowercase (l) or not (N)
#       |||- Split on letter/number transitions (n) or not (N)
#       ||- Split on case changes (c) or not (N)
#       |- Split on punctuation (p) or not (N)


# "NsN->NNN->N->NNN->N"
#  ---
#  (tokenization)

# "NsN->NNN->N->NNNN->N"
#       ---
#       (word splitting on rules, like WordDelimeterFilter in Lucene)

# "NsN->NNN->N->NNNN->N"
#            -
#            (lowercasing or not)

# "NsN->NNN->N->NNNN->N"
#               ----
#               (dictionary based edits: synonyms, stopwords, etc)

# "NsN->NNN->N->NNNN->N"
#                     - stemming (porter)

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

lucytok-0.1.5.tar.gz (33.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

lucytok-0.1.5-py3-none-any.whl (32.9 kB view details)

Uploaded Python 3

File details

Details for the file lucytok-0.1.5.tar.gz.

File metadata

  • Download URL: lucytok-0.1.5.tar.gz
  • Upload date:
  • Size: 33.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.8.3 CPython/3.12.6 Darwin/24.1.0

File hashes

Hashes for lucytok-0.1.5.tar.gz
Algorithm Hash digest
SHA256 7671716d30bf39acc2b22aaf27a23d24e2d65ce67bfa5231cff9600286253a26
MD5 c919b902f41646efbee406be6c3bf1f7
BLAKE2b-256 f19cfa4388d31bd312239f7b9de2392d2bd3ab9c0d2a3b6ca974652149b6214e

See more details on using hashes here.

File details

Details for the file lucytok-0.1.5-py3-none-any.whl.

File metadata

  • Download URL: lucytok-0.1.5-py3-none-any.whl
  • Upload date:
  • Size: 32.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.8.3 CPython/3.12.6 Darwin/24.1.0

File hashes

Hashes for lucytok-0.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 dd0c91ce8b9662f0b463a21c92fb0f156cae69cdb192aa10a401d5cb02044bdf
MD5 a057c0af803993c690106950ffc54e35
BLAKE2b-256 b872422e8e71ad3af6d547da0eeb1e35d1ed5c797cb930e9a43bccab0b1783e2

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page