Chinese text preprocess
You can extract numbers, email, website, emoji, tex, and delete spaces, punctuations.
Install
>> pip install cnprep
Usage
from cnprep import Extractor ext = Extractor(args=['email', 'number'], limit=5) ext.extract(message)
args: option
e.g. ['email', 'telephone'] or 'email, telephone'
email
telephone
web
QQ
tex
wechat
message (without punctuation)
blur (Ⅰ①壹...)
limit: parameter for get_number (blur)
Also, you can use ‘’ext.reset_param()’’ to reset the parameters.
Attention
The URL extractor only support ASCII
Metadata
Release files for cnprep 0.1.12
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cnprep-0.1.12-py2.py3-none-any.whl | Python 2, Python 3 | none | any | Details |
Release files / cnprep-0.1.12-py2.py3-none-any.whl
| Download URL | cnprep-0.1.12-py2.py3-none-any.whl |
|---|---|
| Size | 6.6 kB |
| Tags | Python 2 Python 3 |
|
SHA-256 checksum How to use checksums |
483ae57599402fd4421ed1b436d3fe7e7a29c813bb54586d49b48b98f6ce5323
|
|
BLAKE2b-256 checksum How to use checksums |
a8f57940a7544d04cfbaf3f892e2e7b9019a60d54f5f1120916feae94bcebfd7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |