Cleantextclip
Library to prepare text for machine learning and NLP tasks. Originated from CLIP model preparation, but a few more rules were added.
Installation
pip install -U ternaus_cleantext
Cleans text similar, but stricter than in the CLIP model:
- Escapes HTML characters
- Removes html tags
- Removes URLs
- Removes extra white spaces
- Text to lower case
from ternaus_cleantext.ternaus_cleantext import clean_text
print(clean_text("This is a test https://ternaus.com <b>bold</b>"))
returns
this is a test bold
Metadata
Release files for ternaus-cleantext 0.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ternaus_cleantext-0.0.1.tar.gz | 5.3 kB | Details |
Release files / ternaus_cleantext-0.0.1.tar.gz
| Download URL | ternaus_cleantext-0.0.1.tar.gz |
|---|---|
| Size | 5.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
29dbf62943c1717b65c108ce62a824eb00757544f19e8148ceb6a0590b3a5a1b
|
|
BLAKE2b-256 checksum How to use checksums |
0e71fdcf492b444a973555001a9b73173215902e4cd7ea49ee4cf8666d8b70d1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.10.8
|