Skip to main content

TextPreprocessor

Overview

The TextPreprocessor is a Python class designed for comprehensive text preprocessing. This library can perform the following tasks:

1. Text lematization    --> text_lematizacion()
2. Text stemming        --> text_stemming()
3. Remove links         --> remove_links()
4. Remove hastags       --> remove_hastags()
5. Remove characters    --> remove_characters()
6. Convert to lowercase --> convert_to_lowercase()
7. Remove emojis        --> remove_emojis()
8. Remove numbers       --> remove_numbers()
9. Remove stopwords     --> remove_stopwords()
10. Cloud words         --> cloudwords()
11. Frequency words     --> frecwords()

Installation

Ensure you have NLTK installed. You can install NLTK via pip:

pip install nltk

Usage

### Import the TextPreprocessor class
from text_preprocessor import TextPreprocessor
# Initialize the preprocessor with default settings
preprocessor = TextPreprocessor()

# Customize the preprocessor by setting flags
preprocessor = TextPreprocessor(
    text_lematizacion_flag =True, 
    text_stemming_flag=True, 
    remove_links=True, 
    remove_hastags=True,
    remove_characters=True,
    convert_to_lowercase=True, 
    remove_emojis=True, 
    remove_numbers=True,
    remove_stopwords_flag= True, 
    cloudwords =True,
    frecwords = True
)


text = "Your text goes here..."
processed_text = preprocessor.preprocess_text(text)

Available Methods

  1. preprocess_text(text): Preprocesses the input text based on the initialized flags.
  2. Other methods in the class can be used individually for specific preprocessing steps (e.g., remove_links, remove_stopwords, etc.).

Examples

text = "Hello! This is an example text with #hashtags and links: https://example.com"
# Initialize preprocessor
preprocessor = TextPreprocessor(remove_links=True, remove_hashtags=True)

# Preprocess text
processed_text = preprocessor.preprocess_text(text)
print(processed_text)
Output: "Hello This is an example text with and links"
texto = "If you need to tokenize the text before removing stopwords, you might want to incorporate a tokenization function such as word_tokenize from NLTK."
# lematizacion
p1 = pt.TextPreprocessor()
doc = p1.lematizacion_text(texto)
Output: "if you need to tokenize the text before remove stopword , you might want to incorporate a tokenization function such as word_tokenize from NLTK ."
texto = "If you need to tokenize the text before removing stopwords, you might want to incorporate a tokenization function such as word_tokenize from NLTK."
# stemming
p1 = pt.TextPreprocessor()
doc = p1.stemming_text(texto)
Output: "if you need to token the text befor remov stopword , you might want to incorpor a token function such as word_token from nltk ."

License

This project is licensed under the MIT License - see the LICENSE file for details.

Release files for procesadortext 1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for procesadortext 1.0
File Size Uploaded
procesadortext-1.0.tar.gz 4.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for procesadortext 1.0
File Interpreter ABI Platform
procesadortext-1.0-py3-none-any.whl Python 3 none any Details

Total release size: 9.4 kB

Release files / procesadortext-1.0.tar.gz

Download URL procesadortext-1.0.tar.gz
Size 4.6 kB
Tags Source
SHA-256 checksum
How to use checksums
33bc3aed22af327c42a50fbb4d909537654fb9e868a14b5621978017278a54de
BLAKE2b-256 checksum
How to use checksums
1ca8acae0bb0cb8adddb247b0c62fcca045fedc08a0ae8134279b289fe4630f0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.12.0

Release files / procesadortext-1.0-py3-none-any.whl

Download URL procesadortext-1.0-py3-none-any.whl
Size 4.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5b4ffa59b1181b675c2179b3979ad650180a3a76659c0d10effc4d7e650d78fe
BLAKE2b-256 checksum
How to use checksums
0e6319b6639be90feffa454afd594f534d26c5463242072c85f350b2479a956f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.12.0

Release history Release notifications | RSS feed

This release

1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page