Skip to main content
# preprocessingtext

A tool short, but very usefull to help in pre-processing data from texts.



## How to Install


>> pip install --user preprocessingtext



## Usage

#### Using stem_sentence()

>> from preprocessingtext import CleanSentence

>> cleaner = CleanSentence(idiom='portuguese')

>> cleaner.stem_sentence(sentence="String", remove_stop_words=True, remove_punctuation=True, normalize_text=True, replace_garbage=True)

To init a a class, you need to pass the idiom that you want to work. The custom value, is "portuguese".

Before, you can instance a new object from CleanSentence, and call the method stem_sentence. You can choose in use
"remove_stop_words" from string (pass True or False) and "remove_punctuation" from string (pass True or False),
"replace_garbage" (True or False) removing values from data, and "normalize_text" (True or False) to normalize text.

#### Usage of list_to_replace
You can improve what you need to replace (clean) in your data. You can use "cleaner.list_to_replace.append('what_you_need_to_add')",
or you can pass a new list of values: cleaner.list_to_replace = ['item1', 'item2', 'item3']

# Custom value of list_to_replace
>> list_to_replace = ['https://', 'http://', '$']

# Adding new values
list_to_replace.append('item1')
['https://', 'http://', 'R$', '$', 'item1']

# Replacing values
>> list_to_replace = ['item1', 'item2', 'item3']


#### Using tokenizer()

>> cleaner.tokenizer('Um exemplo de tokens.')

>> ['Um', 'exemplo', 'de', 'tokens']

## Example

## Using all parameters of stem_sentence()
>> string = "Eu sou uma sentença comum. Serei pré-processada com este modulo, veremos a serguir usando os métodos disponiveis"
>> cleaner.stem_sentence(sentence=string,
remove_stop_words=True,
remove_punctuation=True,
normalize_text=True,
replace_garbage=True
)
>> sentenc comum pre-process modul ver segu us metod disponi

## Don't using remove_stop_words
>> print(cleaner.stem_sentence(sentence=string,
remove_stop_words=False,
remove_punctuation=True,
normalize_text=True,
replace_garbage=True
)
)
>> eu sou uma sentenc comum ser pre-process com est modul ver a segu us os metod disponi

## Tokenizer
>> print(cleaner.tokenizer('Um exemplo de tokens.'))
>> ['Um', 'exemplo', 'de', 'tokens']

## Cleaning garbage words
>> string_web = 'Acesse esses links para ganhar dinheiro: https://easymoney.com.net and http://falselink.com'
>> cleaner.stem_sentence(sentence=string_web,
remove_stop_words=False,
remove_punctuation=True,
replace_garbage=True
)
>> acess ess link par ganh dinh easymoney.com.net and falselink.com

## English example
>> en_cleaner = CleanSentences(idiom='english')

>> string_web = 'Access these links to gain money: https://easymoney.com.net and http://falselink.com'
>> print(en_cleaner.stem_sentence(sentence=string_web,
remove_stop_words=True,
remove_punctuation=True,
replace_garbage=True
)
)
>> acc link gain money easymoney.com.net falselink.com


# Author
{
'name': Everton Tomalok,
'email': evertontomalok123@gmail.com
}

Metadata

Release files for preprocessingtext 0.0.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for preprocessingtext 0.0.4
File Size Uploaded
preprocessingtext-0.0.4.tar.gz 3.9 kB Details

Release files / preprocessingtext-0.0.4.tar.gz

Download URL preprocessingtext-0.0.4.tar.gz
Size 3.9 kB
Tags Source
SHA-256 checksum
How to use checksums
81a3365bf106b901bf3bd968d75e9ba0190b363fd943d6e140e76d84f7ff2e75
BLAKE2b-256 checksum
How to use checksums
353301f321ffff1e01f90fd97974fc6bd1e0eb1ad35d8dd35b795d54eeda3c69
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via Python-urllib/3.6

Release history Release notifications | RSS feed

This release

0.0.4 This release

1 release file

0.0.3

1 release file

0.0.2

1 release file

0.0.1

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page