This is a pre-production deployment of Warehouse. Changes made here affect the production instance of PyPI (pypi.python.org).
Help us improve Python packaging - Donate today!

tokenizers for Whoosh designed for Japanese language

Project Description

About

Tokenizers for Whoosh full text search library designed for Japanese language. This package conteins two Tokenizers.

  • IgoTokenizer
  • TinySegmenterTokenizer
  • MeCabTokenizer

How To Use

IgoTokenizer:

import igo.Tagger
import whooshjp
from whooshjp.IgoTokenizer import IgoTokenizer

tk = IgoTokenizer(igo.Tagger.Tagger('ipadic'))
scm = Schema(title=TEXT(stored=True, analyzer=tk), path=ID(unique=True,stored=True), content=TEXT(analyzer=tk))

TinySegmenterTokenizer:

import tinysegmenter
import whooshjp
from whooshjp.TinySegmenterTokenizer import TinySegmenterTokenizer

tk = TinySegmenterTokenizer(tinysegmenter.TinySegmenter())
scm = Schema(title=TEXT(stored=True, analyzer=tk), path=ID(unique=True,stored=True), content=TEXT(analyzer=tk))

Changelog for Japanese Tokenizers for Whoosh

2011-02-19 – 0.1
  • first release.
2011-02-21 – 0.2
  • add TinySegmenterTokenizer
  • change module name
2011-02-24 – 0.3
  • add FeatureFilter
2011-02-27 – 0.4
  • add MeCabTokenizer
  • add a mode for don’t pickle igo tagger to minimize index.
2011-04-17 – 0.5
  • correct char offsets
2011-04-17 – 0.6
  • correct char offsets(TinySegmenterTokenizer)
2012-04-14 – 0.7
  • rename package(WhooshJapaneseTokenizer to whooshjp)
  • no longer import sub modules automatically
  • Python3 compatibility(3.2, 3.3)
  • Drop Python2.5 support
Release History

Release History

This version
History Node

0.7

History Node

0.6

History Node

0.5

History Node

0.4

History Node

0.2

History Node

0.1

Download Files

Download Files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

File Name & Checksum SHA256 Checksum Help Version File Type Upload Date
whoosh-igo-0.7.tar.gz (8.0 kB) Copy SHA256 Checksum SHA256 Source Jul 16, 2012

Supported By

WebFaction WebFaction Technical Writing Elastic Elastic Search Pingdom Pingdom Monitoring Dyn Dyn DNS Sentry Sentry Error Logging CloudAMQP CloudAMQP RabbitMQ Heroku Heroku PaaS Kabu Creative Kabu Creative UX & Design Fastly Fastly CDN DigiCert DigiCert EV Certificate Rackspace Rackspace Cloud Servers DreamHost DreamHost Log Hosting