A powerful CLI-based text corpus analyser for extracting palindromes, anagrams, word frequencies, pattern matches, emails, and phone numbers from text files or direct input.
Project description
tcorpus - Text Corpus Analyser
A powerful command-line tool for analyzing text corpora. Extract palindromes, anagrams, word frequencies, pattern matches, emails, and phone numbers from text files or direct input.
Features
- 🔍 Word Analysis: Find palindromes, anagrams, and word frequencies
- 🎯 Pattern Matching: Search words using wildcard patterns and masks
- 📧 Email Extraction: Extract email addresses from text
- 📱 Phone Number Detection: Find phone numbers in various formats
- 🚫 Stopword Filtering: Filter out common words using config files or CLI options
- 📊 Multiple Output Formats: Save results as JSON or CSV
- ⚡ Fast & Lightweight: Zero external dependencies, pure Python standard library
Installation
From PyPI (Recommended)
pip install tcorpus
From Source
git clone <repository-url>
cd text-corpus-analyser
pip install -e .
Requirements
- Python 3.8 or higher
- No external dependencies (uses only Python standard library)
Quick Start
# Get help
tcorpus --help
# Analyze a text file
tcorpus all demo.txt output.json
# Use direct text input
tcorpus palindrome --text "A man a plan a canal Panama" output.json
Usage
Input Options
You can provide input in two ways:
-
From a file:
tcorpus <command> <input_file> <output_file>
-
Direct text (using
-tor--text):tcorpus <command> --text "Your text here" <output_file>
Common Options
All commands support these filtering options:
--stopwords <word1> <word2> ...- Additional stopwords to ignore (combined with config.ini)--starts-with <letter>- Keep only words starting with this letter--config <path>- Path to config file for stopwords (default:config.ini)-pw, --print-words- Print filtered results to terminal
Config File
Create a config.ini file in your working directory to define default stopwords:
[stopwords]
words = the,or,but,a,an,is,are,was,were
Commands
1. Palindrome Detection
Find all palindromes (words that read the same forwards and backwards) in your text.
Syntax:
tcorpus palindrome [input_file] <output_file> [options]
Examples:
# From file
tcorpus palindrome demo.txt palindromes.json
# Direct text
tcorpus palindrome --text "madam level cat radar" output.json
# With filters
tcorpus palindrome demo.txt output.json --starts-with m --print-words
# Filter out stopwords
tcorpus palindrome demo.txt output.json --stopwords the a an
Output Format:
{
"palindromes": ["level", "madam", "radar"]
}
2. Anagram Detection
Find groups of words that are anagrams of each other.
Syntax:
tcorpus anagram [input_file] <output_file> [options]
Examples:
# Basic usage
tcorpus anagram demo.txt anagrams.json
# With direct text
tcorpus anagram --text "cat act tac dog god" output.json
# Print results to terminal
tcorpus anagram demo.txt output.json --print-words
Output Format:
{
"anagrams": [
["act", "cat", "tac"],
["dog", "god"]
]
}
3. Word Frequency Analysis
Count how often each word appears in the text.
Syntax:
tcorpus freq [input_file] <output_file> [options]
Options:
-w, --words <word1> <word2> ...- Count specific words only, or useallfor all words
Examples:
# Count all words
tcorpus freq demo.txt frequencies.json
# Count specific words
tcorpus freq demo.txt output.json --words python code program
# Count all words (explicit)
tcorpus freq demo.txt output.json --words all
# Output as CSV
tcorpus freq demo.txt frequencies.csv
# With filters
tcorpus freq demo.txt output.json --starts-with p --print-words
Output Format (JSON):
{
"frequencies": {
"python": 5,
"code": 3,
"program": 2
}
}
Output Format (CSV):
word,count
code,3
program,2
python,5
4. Pattern Matching (Mask)
Find words matching a specific pattern using wildcards and special syntax.
Syntax:
tcorpus mask <pattern> [input_file] <output_file> [options]
Pattern Syntax:
*- Matches zero or more characters (wildcard)?- Matches exactly one characterword+- Words starting with "word" (e.g.,ram+matches "ram", "ramesh", "ramadan")+word- Words ending with "word" (e.g.,+ingmatches "running", "coding")+word+- Words containing "word" (e.g.,+ram+matches "program", "ramesh")
Options:
--min-length <n>- Minimum word length--max-length <n>- Maximum word length--length <n>- Exact word length--contains <substring>- Word must contain this substring
Examples:
# Wildcard pattern: words starting with 's' and ending with 'e'
tcorpus mask "s*e" demo.txt output.json
# Single character wildcard
tcorpus mask "c?t" demo.txt output.json
# Starts with pattern
tcorpus mask "ram+" demo.txt output.json
# Ends with pattern
tcorpus mask "+ing" demo.txt output.json
# Contains pattern
tcorpus mask "+ram+" demo.txt output.json
# With length filters
tcorpus mask "s*e" demo.txt output.json --min-length 4 --max-length 6
# Exact length
tcorpus mask "a*d" demo.txt output.json --length 5
# Must contain substring
tcorpus mask "s*e" demo.txt output.json --contains "al"
# Print matches to terminal
tcorpus mask "s*e" demo.txt output.json --print-words
Output Format:
{
"mask_matches": ["sale", "safe", "same", "sane", "site", "size"]
}
5. Email Extraction
Extract email addresses from text.
Syntax:
tcorpus email [input_file] <output_file> [options]
Examples:
# From file
tcorpus email demo.txt emails.json
# Direct text
tcorpus email --text "Contact us at info@example.com or support@test.org" output.json
# Print to terminal
tcorpus email demo.txt output.json --print-words
Output Format:
{
"emails": ["info@example.com", "support@test.org"]
}
6. Phone Number Extraction
Extract phone numbers from text in various formats.
Syntax:
tcorpus phone [input_file] <output_file> [options]
Options:
-d, --digits <n>- Minimum number of digits required (default: 10)
Examples:
# Basic usage (default: 10 digits minimum)
tcorpus phone demo.txt phones.json
# Custom minimum digits
tcorpus phone demo.txt output.json --digits 7
# Direct text
tcorpus phone --text "Call +1 (555) 123-4567 or 07123 456789" output.json
# Print to terminal
tcorpus phone demo.txt output.json --print-words
Output Format:
{
"phone_numbers": ["+1 (555) 123-4567", "07123 456789"]
}
7. All Analyses
Run all analyses at once: palindromes, anagrams, frequencies, emails, and phone numbers.
Syntax:
tcorpus all [input_file] <output_file> [options]
Options:
-w, --words <word1> <word2> ...- Words to count (useallfor all words)-d, --digits <n>- Minimum digits for phone numbers (default: 10)
Examples:
# Run all analyses
tcorpus all demo.txt complete_analysis.json
# With word frequency filter
tcorpus all demo.txt output.json --words python code
# Custom phone digit requirement
tcorpus all demo.txt output.json --digits 7
# With text filters
tcorpus all demo.txt output.json --starts-with a --stopwords the a an
# Print all results to terminal
tcorpus all demo.txt output.json --print-words
Output Format:
{
"palindromes": ["level", "madam"],
"anagrams": [["act", "cat", "tac"]],
"frequencies": {
"python": 5,
"code": 3
},
"emails": ["info@example.com"],
"phone_numbers": ["+1 (555) 123-4567"]
}
Advanced Examples
Combining Filters
# Find palindromes starting with 'm', excluding stopwords
tcorpus palindrome demo.txt output.json \
--starts-with m \
--stopwords the a an \
--print-words
Processing Compressed Files
The tool automatically handles .gz compressed files:
tcorpus all large_text.txt.gz output.json
Batch Processing
# Process multiple files
for file in *.txt; do
tcorpus all "$file" "${file%.txt}_analysis.json"
done
Output Formats
JSON (Default)
All commands output JSON by default. The structure varies by command:
- Single analysis:
{"palindromes": [...]} - All analyses:
{"palindromes": [...], "anagrams": [...], ...}
CSV (Frequency Only)
When using freq command with .csv extension:
tcorpus freq demo.txt frequencies.csv
Outputs:
word,count
python,5
code,3
Troubleshooting
Command Not Found
If tcorpus is not recognized after installation:
-
Check installation:
pip show tcorpus
-
Verify PATH includes Python Scripts directory:
python -m site --user-base
-
Use module execution:
python -m tcorpus --help
Config File Not Found
If config.ini is missing, the tool will still work but won't use default stopwords. Create a config.ini file in your working directory:
[stopwords]
words = the,or,but,a,an,is,are
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
Authors
- Raghu - raghu59770@gmail.com
Maintainers
- Santosh
- Raghuram
License
This project is open source and available for use.
Version
Current version: 0.1.5
Support
For issues, questions, or contributions, please open an issue on the project repository.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file tcorpus-0.1.6.tar.gz.
File metadata
- Download URL: tcorpus-0.1.6.tar.gz
- Upload date:
- Size: 14.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fa5da0261e41a9b49ebbdea7b933e7d50e205f332a61c10dd578eba2688cbd64
|
|
| MD5 |
e21584934d0c19531b8609d14372ec34
|
|
| BLAKE2b-256 |
7b264b89ffb1d8e05bef0a216bd634f92efa6a722760c543e2cbb7d27f2f4438
|
File details
Details for the file tcorpus-0.1.6-py3-none-any.whl.
File metadata
- Download URL: tcorpus-0.1.6-py3-none-any.whl
- Upload date:
- Size: 12.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
77871157e4e27642262d3f63f18a44fcf02b7e94422fa2833eabdc856620751c
|
|
| MD5 |
4d17b067d37b6663ad11a430d5094cdc
|
|
| BLAKE2b-256 |
eec324f29c1fcf34b7a2b1c9fafe257e1e1e7a1c580c0c9d9c48e920836b0ef4
|