An NLP text preprocessing function package
Project description
This package is all in one text preprocessing pipeline and takes text input in the form of pandas series and outputs as pandas series object.
All the text_processing sub-functions may be viewed here https://github.com/tariqjamil-bwp/pypi_projects/tree/main/tj_nlp
Installing: pip install .....
Usage: (by example) df2 = df1.copy() df2['input_data'] = text_preprocessing_ds(df2['Phrase'])
Note: It will not work with .apply(text_....) method, as .apply is item by item processing which may not be optimized at dataframe level.
Details: It has a variety of input flags, controlling the appropritae text preprocessing activity, and are by default set to True. The details of flags is as under:
** The function is optimized to take advantage of the multi threaded processors **
Arguments:
- pdSeries: A positional argument, of pandas series object type, and should contain text phrases to be preprocessed
** All following arguments are key-word type **
-
NewLine removal: Parameter: 'newlines_tabs' This setting when enablesd causes the function to remove all the occurrences of newlines, tabs, and combinations like: \n, \.
Example: Input : This is her \ first day at this place.\n Please,\t Be nice to her.\n Output : This is her first day at this place. Please, Be nice to her.
-
HTML Tags Parameter: 'remove_html' This setting when enablesd causes the function to remove all the occurrences of html tags from the text.
Example: Input : This is a nice place to live. Output : This is a nice place to live.
-
Hyper Links removal This setting when enablesd causes the function to remove all the occurrences of links. Parameter: 'links'
Example: Input : To know more about cats and food & website: catster.com visit: https://catster.com//how-to-feed-cats Output : To know more about cats and food & website: visit:
-
Removing extra white spaces Parameter: 'extra_whitespace' This setting when enablesd causes the function to remove extra whitespaces from the text
Example: Input : How are you doing ? Output : How are you doing ?
-
Accented Characters Removal: Parameter: 'accented_chars' This setting when enablesd causes the function to remove accented characters from the text contained within the Dataset.
Example: Input : Málaga, àéêöhello Output : Malaga, aeeohello
-
Lower Casing Parameter: 'lowercase' This setting when enablesd causes the function to convert text into lower case.
Example: Input : The World is Full of Surprises! Output : the world is full of surprises!
-
Removing Character Repetitions: Parameter: 'repeatition' This setting when enablesd causes the function to reduce repeatition to two characters for alphabets and to one character for punctuations.
Example: Input : Realllllllllyyyyy, Greeeeaaaatttt !!!!?....;;;;:) Output : Reallyy, Greeaatt !?.;:)
-
Expand Contractions: Parameter: 'contractions' This setting when enablesd causes the function to expand shortened words to the actual form. e.g. don't to do not
Example: Input : ain't, aren't, can't, cause, can't've Output : is not, are not, cannot, because, cannot have -
Removing Special Characters Parameter: 'special_chars' This setting when enablesd causes the function to remove all the special characters except a preset list of characters as they have imp meaning in the text provided.
Example: Input : Hello, T-a-r-i-q. Thi*s is $100.05 : the payment that you will recieve! (Is this okay?) Output : Hello, Tariq. This is $100.05 : the payment that you will recieve! Is this okay?
-
Removong Stop Words Parameter: 'stop_words' This setting when enablesd causes the function to remove the stopwords which doesn't add much meaning to a sentence & they can be remove safely without comprimising meaning of the sentence.
Example: Input : This is Tariq from karachi who came here to study. Output : ["'This", 'Tariq', 'karachi', 'came', 'study', '.', "'"]
-
Mis-spelled words Correction: Parameter: 'mis_spell' This setting when enablesd causes the function to correct spellings. Currently only English language is supported.
Example: Input : This is OAkbar from Londn who came heree to studdy. Output : This is Akbar from London who came here to study.
-
Lamentization: Parameters: 'lemmatization_allow' This setting when enablesd causes the function to converts word to their root words without explicitely cut down as done in stemming.
Example: Input : text reduced Output : text reduce
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file tjtext_preproc_ds-0.1.0.tar.gz.
File metadata
- Download URL: tjtext_preproc_ds-0.1.0.tar.gz
- Upload date:
- Size: 11.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/4.0.2 CPython/3.8.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c93a192dbadbcf63a7801747bc61e34c0bf8a35cad308ba24ef4a21c67ccef61
|
|
| MD5 |
ca4d33ac79ed3d9a8a5387e2d6dc13ee
|
|
| BLAKE2b-256 |
37bd92fc9944dbb4347cd27683a7db379090bb03b6f17c123a1de608b561b20c
|
File details
Details for the file tjtext_preproc_ds-0.1.0-py3-none-any.whl.
File metadata
- Download URL: tjtext_preproc_ds-0.1.0-py3-none-any.whl
- Upload date:
- Size: 9.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/4.0.2 CPython/3.8.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1ce640d46e8cddcf23d910bce290e8072fe4a8b6fe0d753d8e0ea0f85e65ddad
|
|
| MD5 |
babade4938d34e11013ee10d305b2da7
|
|
| BLAKE2b-256 |
d27e84eb6dd67561b6131461ab2427d867fc0457ea3bd0cea87d7e5ca424e396
|