NLP Classification Python Package For Profile Category and NER
Functionality of the Package
Performs Classification using Regex for social media user's profile categorization and NER tasks
- If profile:
- Categories: Politician, Information Vehicule, Health Professional, Science (those with scientific/academic background), Education Professional, Artist, Organization and Journalist;
- NER:
- Vaccines, Products, Drugs, Diseases, Symptoms, Science, Part of the body; returns a dataframe that has information of the entity id, name and it's occurence frequency on the input text.
The package takes the following parameters as input:
- Profile:
- Text containing informations about the user's biography or channel description in the context of communication reaseach porpuses
- NER:
- Text
Usage
pip install ProfileNER-classifier
Example
from nlpclassifier_profilener import NLPClassifier
#Instatiate the classifier
#If profile classification is the task of choice
nlpc = NLPClassifier('profile')
#Preprocess the text. Its's important to lowercase and remove accents
yt['channelDesc'] = yt['channelDesc'].apply(lambda x: pipeline.preprocess(x, lower = True)).apply(lambda x: pipeline.strip_accents(x))
#Classify channelCategory
yt['channelCategory'] = yt['channelDesc'].apply(lambda x: tpc.classifier(str(x)))
#See the distribution of profile categories if Profile task
yt['channelCategory'].value_counts()
#If NER classification, there's two options
#The first is to get the entities name as you would get with Spacy, for example. For this purpose, use the classifier function with 'ner' use.
nlpc = NLPClassifier('ner')
yt['ner'] = yt['videoTranscription'].apply(lambda x: nlpc.classifier(str(x))) #already preprocessed
#The second is to get the entities occurence frequency in the text. For this purpose, use the ner_classifier function with 'ner' use. Remember that it returns the input dataframe updated with the ner entities as new columns and their frequency as their rows.
nlpc = NLPClassifier('ner')
yt = yt['videoTranscription'].apply(lambda x: nlpc.ner_classifier(str(x))) #already preprocessed
Note
The package is still a work is in progress. In case of error, feel free to contact me.
Change Log
0.0.7.6 (04/09/2022)
- First Release
Release files for ProfileNER-classifier 0.7.7
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ProfileNER-classifier-0.7.7.tar.gz | 19.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ProfileNER_classifier-0.7.7-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 40.4 kB
Release files / ProfileNER-classifier-0.7.7.tar.gz
| Download URL | ProfileNER-classifier-0.7.7.tar.gz |
|---|---|
| Size | 19.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c492a81fcb08a88faae56324b6e00b19681817d4a44b97916f46d9ee622affe7
|
|
BLAKE2b-256 checksum How to use checksums |
84dd159f5dafd849cfc9f66585338687061414edc38d371ab0d5bbb9dddaf39b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.1 CPython/3.8.5
|
Release files / ProfileNER_classifier-0.7.7-py3-none-any.whl
| Download URL | ProfileNER_classifier-0.7.7-py3-none-any.whl |
|---|---|
| Size | 21.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f8b2d164d67a4fd5b0138bd046843341bcea055310510abee1725c0ab5584700
|
|
BLAKE2b-256 checksum How to use checksums |
759135ee9f75c101e6c6d98887a94b3b6a36c12d3ff42b1b4ddb5bd8409dfdf4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.1 CPython/3.8.5
|