A Package for text preprocessing

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Project description

text-preprocessing

1. Text Cleaning

from nlp_preprocessing import clean

texts = ["Hi I am's nakdur"]
cleaned_texts = clean.clean_v1(texts)

2. Dataset Prepration

from nlp_preprocessing import dataset as ds
import pandas as pd

text = ['I am Test 1','I am Test 2']
label = ['A','B']
aspect = ['C','D']
data = pd.DataFrame({'text':text*5,'label':label*5,'aspect':aspect*5})
data

data_config = {
            'data_class':'multi-label',
            'x_columns':['text'],
            'y_columns':['label','aspect'],
            'one_hot_encoded_columns':[],
            'label_encoded_columns':['label','aspect'],
            'data':data,
            'split_ratio':0.1
          }

dataset = ds.Dataset(data_config)
train, test = dataset.get_train_test_data()

print(train['Y_train'],train['X_train'])
print(test['Y_test'],test['X_test'])
print(dataset.data_config)

3. Seq token generator

texts = ['I am Test 2', 'I am Test 1', 'I am Test 1', 'I am Test 1','I am Test 1', 'I am Test 2', 'I am Test 1', 'I am Test 2','I am Test 2']

tokens = seq_gen.get_word_sequences(texts)
print(tokens)

Project details

These details have not been verified by PyPI

Project links

Homepage

GitHub Statistics

View statistics for this project via Libraries.io, or by using our public dataset on Google BigQuery

Release history Release notifications | RSS feed

0.2.0

Aug 15, 2020

0.1.13

May 31, 2020

0.1.12

May 27, 2020

0.1.11

May 18, 2020

0.1.10

May 18, 2020

0.1.9

May 17, 2020

0.1.8

May 17, 2020

0.1.7

May 17, 2020

0.1.6

Apr 21, 2020

This version

0.1.5

Apr 16, 2020

0.1.4

Apr 16, 2020

0.1.3

Apr 16, 2020

0.1.2

Apr 6, 2020

0.1.1

Apr 4, 2020

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nlp_preprocessing-0.1.5.tar.gz (8.1 kB view hashes)

Uploaded Apr 16, 2020 Source

Built Distribution

nlp_preprocessing-0.1.5-py3-none-any.whl (8.5 kB view hashes)

Uploaded Apr 16, 2020 Python 3

Hashes for nlp_preprocessing-0.1.5.tar.gz

Hashes for nlp_preprocessing-0.1.5.tar.gz
Algorithm	Hash digest
SHA256	`dcd5424696afb65868897b7dfe40c26ddbea4df9f5c2aa84f94b6ecb7c691e91`
MD5	`1b0cef7cfad9bba5aac73dd19c65653c`
BLAKE2b-256	`da0580a7e87715b211d50d286b53a41f8eeb4503aea48a50fd4cb4bb6daf8014`

Hashes for nlp_preprocessing-0.1.5-py3-none-any.whl

Hashes for nlp_preprocessing-0.1.5-py3-none-any.whl
Algorithm	Hash digest
SHA256	`37e88a1aaee0ecd08a9d36b4d2353659ee82dc6b802ba7c57d388e64e942ce32`
MD5	`f1442b9aaa2779194a5001ce4c693d2a`
BLAKE2b-256	`65ea5b97243f3c0c9ce78549767e67d8da72a71b0cb59a02f36cd44d22f1fa14`