Skip to main content

tidyX

GitHub stars

Downloads

Before and After tidyX

tidyX is a Python package designed for cleaning and preprocessing text for machine learning applications, especially for text written in Spanish and originating from social networks. This library provides a complete pipeline to remove unwanted characters, normalize text, group similar terms, etc. to facilitate NLP applications.

To deep dive in the package visit our website

Installation

Install the package using pip:

pip install tidyX

Make sure you have the necessary dependencies installed. If you plan on lemmatizing, you'll need spaCy along with the appropriate language models. For Spanish lemmatization, we recommend downloading the es_core_news_sm model:

python -m spacy download es_core_news_sm 

For English lemmatization, we suggest the en_core_web_sm model:

python -m spacy download en_core_web_sm 

To see a full list of available models for different languages, visit Spacy's documentation.

Features

  • Standardize Text Pipeline: The preprocess() method provides an all-encompassing solution for quickly and effectively standardizing input strings, with a particular focus on tweets. It transforms the input to lowercase, strips accents (and emojis, if specified), and removes URLs, hashtags, and certain special characters. Additionally, it offers the option to delete stopwords in a specified language, trims extra spaces, extracts mentions, and removes 'RT' prefixes from retweets.
from tidyX import TextPreprocessor as tp



# Raw tweet example

raw_tweet = "RT @user: Check out this link: https://example.com 🌍 #example 😃"



# Applying the preprocess method

cleaned_text = tp.preprocess(raw_tweet)



# Printing the cleaned text

print("Cleaned Text:", cleaned_text)

Output:


Cleaned Text: check out this link

To remove English stopwords, simply add the parameters remove_stopwords=True and language_stopwords="english":

from tidyX import TextPreprocessor as tp



# Raw tweet example

raw_tweet = "RT @user: Check out this link: https://example.com 🌍 #example 😃"



# Applying the preprocess method with additional parameters

cleaned_text = tp.preprocess(raw_tweet, remove_stopwords=True, language_stopwords="english")



# Printing the cleaned text

print("Cleaned Text:", cleaned_text)

Output:


Cleaned Text: check link

For a more detailed explanation of the customizable steps of the function, visit the official preprocess() documentation.

  • Stemming and Lemmatizing: One of the foundational steps in preparing text for NLP applications is bringing words to a common base or root. This library provides both stemmer() and lemmatizer() functions to perform this task across various languages.

  • Group similar terms: When working with a corpus sourced from social networks, it's common to encounter texts with grammatical errors or words that aren't formally included in dictionaries. These irregularities can pose challenges when creating Term Frequency matrices for NLP algorithms. To address this, we developed the create_bol() function, which allows you to create specific bags of terms to cluster related terms.

  • Remove unwanted elements: such as special characters, extra spaces, accents, emojis, urls, tweeter mentions, among others.

  • Dependency Parsing Visualization: Incorporates visualization tools that enable the display of dependency parses, facilitating linguistic analysis and feature engineering.

  • Much more!

Tutorials

Contributing

Contributions to enhance tidyX are welcome! Feel free to open issues for bug reports, feature requests, or submit pull requests. If this package has been helpful, please give us a star :D

Metadata

Release files for tidyX 1.7.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tidyX 1.7.1
File Size Uploaded
tidyX-1.7.1.tar.gz 5.6 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for tidyX 1.7.1
File Interpreter ABI Platform
tidyX-1.7.1-py3-none-any.whl Python 3 none any Details

Total release size: 5.7 MB

Release files / tidyX-1.7.1.tar.gz

Download URL tidyX-1.7.1.tar.gz
Size 5.6 MB
Tags Source
SHA-256 checksum
How to use checksums
7b20e88e7873ad36872704b7ce1cf6e8d3b0b75d9474ede873eb281bf99e76e2
BLAKE2b-256 checksum
How to use checksums
451d764c0aeca0a7b4d0b6207140d9af6d839b8cbb277c9e58e1f9112b15ac3c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.18

Release files / tidyX-1.7.1-py3-none-any.whl

Download URL tidyX-1.7.1-py3-none-any.whl
Size 149.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c8612ac2892cde0aa38b3abb5fa582f5a051f0ac889798d71ebc3a949306eb23
BLAKE2b-256 checksum
How to use checksums
ba5e45fcb042d0eb7deb3d1f0d7a6db3217638f2c236db3e8631c920396a25fd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.9.18

Release history Release notifications | RSS feed

This release

1.7.1 This release

2 release files

1.7.0

2 release files

1.6.7

2 release files

1.6.6

2 release files

1.6.5

2 release files

1.6.4

2 release files

1.6.3

2 release files

1.6.2

2 release files

1.6.1

2 release files

1.5.4

2 release files

1.5.3

2 release files

1.5.2

2 release files

1.5.1

2 release files

1.5

2 release files

1.4.3

2 release files

1.4.2

2 release files

1.4.1

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.1.1

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page