Skip to main content

A powerful CLI-based text corpus analyser for extracting palindromes, anagrams, word frequencies, pattern matches, emails, and phone numbers from text files or direct input.

Project description

tcorpus - Text Corpus Analyser

Python Version License

A powerful, lightweight command-line tool for analyzing text corpora. Extract palindromes, anagrams, word frequencies, pattern matches, emails, and phone numbers from text files or direct input.

✨ Features

  • 🔍 Word Analysis: Find palindromes, anagrams, and word frequencies
  • 🎯 Pattern Matching: Search words using wildcard patterns and masks
  • 📧 Email Extraction: Extract email addresses from text
  • 📱 Phone Number Detection: Find phone numbers in various formats
  • 🚫 Stopword Filtering: Filter out common words using config files or CLI options
  • 📊 Multiple Output Formats: Save results as JSON or CSV
  • Zero Dependencies: Pure Python standard library, no external packages required
  • 🗜️ Compressed File Support: Automatically handles .gz compressed files

📦 Installation

From PyPI (Recommended)

pip install tcorpus

From Source

git clone https://github.com/Pegasus717/Text-Corpus-Analyser.git
cd text-corpus-analyser
pip install -e .

🚀 Quick Start

# Get help
tcorpus --help

# Analyze a text file with all features
tcorpus all demo.txt output.json

# Use direct text input
tcorpus palindrome --text "A man a plan a canal Panama" output.json

📖 Usage

Input Options

You can provide input in two ways:

  1. From a file:

    tcorpus <command> <input_file> <output_file>
    
  2. Direct text (using -t or --text):

    tcorpus <command> --text "Your text here" <output_file>
    

Common Options

All commands support these filtering options:

  • --stopwords <word1> <word2> ... - Additional stopwords to ignore (combined with config.ini)
  • --starts-with <letter> - Keep only words starting with this letter
  • --config <path> - Path to config file for stopwords (default: config.ini)
  • -pw, --print-words - Print filtered results to terminal

Config File

Create a config.ini file in your working directory to define default stopwords:

[stopwords]
words = the,or,but,a,an,is,are,was,were

📋 Commands

1. Palindrome Detection

Find all palindromes (words that read the same forwards and backwards).

# From file
tcorpus palindrome demo.txt palindromes.json

# Direct text
tcorpus palindrome --text "madam level cat radar" output.json

# With filters
tcorpus palindrome demo.txt output.json --starts-with m --print-words

Output:

{
  "palindromes": ["level", "madam", "radar"]
}

2. Anagram Detection

Find groups of words that are anagrams of each other.

# Basic usage
tcorpus anagram demo.txt anagrams.json

# With direct text
tcorpus anagram --text "cat act tac dog god" output.json

Output:

{
  "anagrams": [
    ["act", "cat", "tac"],
    ["dog", "god"]
  ]
}

3. Word Frequency Analysis

Count how often each word appears in the text.

# Count all words
tcorpus freq demo.txt frequencies.json

# Count specific words
tcorpus freq demo.txt output.json --words python code program

# Output as CSV
tcorpus freq demo.txt frequencies.csv

Output (JSON):

{
  "frequencies": {
    "python": 5,
    "code": 3,
    "program": 2
  }
}

Output (CSV):

word,count
code,3
program,2
python,5

4. Pattern Matching (Mask)

Find words matching a specific pattern using wildcards and special syntax.

Pattern Syntax:

  • * - Matches zero or more characters (wildcard)
  • ? - Matches exactly one character
  • word+ - Words starting with "word" (e.g., ram+ matches "ram", "ramesh", "ramadan")
  • +word - Words ending with "word" (e.g., +ing matches "running", "coding")
  • +word+ - Words containing "word" (e.g., +ram+ matches "program", "ramesh")

Options:

  • --min-length <n> - Minimum word length
  • --max-length <n> - Maximum word length
  • --length <n> - Exact word length
  • --contains <substring> - Word must contain this substring

Examples:

# Wildcard pattern: words starting with 's' and ending with 'e'
tcorpus mask "s*e" demo.txt output.json

# Starts with pattern
tcorpus mask "ram+" demo.txt output.json

# Ends with pattern
tcorpus mask "+ing" demo.txt output.json

# With length filters
tcorpus mask "s*e" demo.txt output.json --min-length 4 --max-length 6

Output:

{
  "mask_matches": ["sale", "safe", "same", "sane", "site", "size"]
}

5. Email Extraction

Extract email addresses from text.

# From file
tcorpus email demo.txt emails.json

# Direct text
tcorpus email --text "Contact us at info@example.com or support@test.org" output.json

Output:

{
  "emails": ["info@example.com", "support@test.org"]
}

6. Phone Number Extraction

Extract phone numbers from text in various formats.

# Basic usage (default: 10 digits minimum)
tcorpus phone demo.txt phones.json

# Custom minimum digits
tcorpus phone demo.txt output.json --digits 7

# Direct text
tcorpus phone --text "Call +1 (555) 123-4567 or 07123 456789" output.json

Output:

{
  "phone_numbers": ["+1 (555) 123-4567", "07123 456789"]
}

7. All Analyses

Run all analyses at once: palindromes, anagrams, frequencies, emails, and phone numbers.

# Run all analyses
tcorpus all demo.txt complete_analysis.json

# With word frequency filter
tcorpus all demo.txt output.json --words python code

# Custom phone digit requirement
tcorpus all demo.txt output.json --digits 7

Output:

{
  "palindromes": ["level", "madam"],
  "anagrams": [["act", "cat", "tac"]],
  "frequencies": {
    "python": 5,
    "code": 3
  },
  "emails": ["info@example.com"],
  "phone_numbers": ["+1 (555) 123-4567"]
}

🔧 Advanced Examples

Combining Filters

# Find palindromes starting with 'm', excluding stopwords
tcorpus palindrome demo.txt output.json \
  --starts-with m \
  --stopwords the a an \
  --print-words

Processing Compressed Files

The tool automatically handles .gz compressed files:

tcorpus all large_text.txt.gz output.json

Batch Processing

# Process multiple files (Unix/Linux/Mac)
for file in *.txt; do
  tcorpus all "$file" "${file%.txt}_analysis.json"
done

# Windows PowerShell
Get-ChildItem *.txt | ForEach-Object {
  tcorpus all $_.Name ($_.BaseName + "_analysis.json")
}

📤 Output Formats

JSON (Default)

All commands output JSON by default. The structure varies by command:

  • Single analysis: {"palindromes": [...]}
  • All analyses: {"palindromes": [...], "anagrams": [...], ...}

CSV (Frequency Only)

When using freq command with .csv extension, outputs CSV format suitable for spreadsheet applications.

🛠️ Requirements

  • Python 3.8 or higher
  • No external dependencies (uses only Python standard library)

🧪 Testing

Run the test suite:

python -m unittest discover -s tests -v

❓ Troubleshooting

Command Not Found

If tcorpus is not recognized after installation:

  1. Check installation:

    pip show tcorpus
    
  2. Verify PATH includes Python Scripts directory:

    python -m site --user-base
    
  3. Use module execution:

    python -m tcorpus --help
    

Config File Not Found

If config.ini is missing, the tool will still work but won't use default stopwords. Create a config.ini file in your working directory:

[stopwords]
words = the,or,but,a,an,is,are

🤝 Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

👥 Authors

👨‍💼 Maintainers

  • Santosh
  • Raghuram

📄 License

This project is open source and available under the MIT License.

📌 Version

Current version: 0.1.7

💬 Support

For issues, questions, or contributions, please open an issue on the project repository.


Made with ❤️ for text analysis enthusiasts

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tcorpus-0.1.7.tar.gz (14.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

tcorpus-0.1.7-py3-none-any.whl (11.8 kB view details)

Uploaded Python 3

File details

Details for the file tcorpus-0.1.7.tar.gz.

File metadata

  • Download URL: tcorpus-0.1.7.tar.gz
  • Upload date:
  • Size: 14.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.0

File hashes

Hashes for tcorpus-0.1.7.tar.gz
Algorithm Hash digest
SHA256 8ed2d9ddee29deb3cf563c9260032accc0ac96b63b16836bc6a991f8ce09c4f3
MD5 5cf46bc1b709875093bcd194f1fcbd9e
BLAKE2b-256 1c7cdc8d590105b102101d08daa63d73acd2615e9f5f9d791cfa9e274143232f

See more details on using hashes here.

File details

Details for the file tcorpus-0.1.7-py3-none-any.whl.

File metadata

  • Download URL: tcorpus-0.1.7-py3-none-any.whl
  • Upload date:
  • Size: 11.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.0

File hashes

Hashes for tcorpus-0.1.7-py3-none-any.whl
Algorithm Hash digest
SHA256 270fdefa01f0786d329b07a17bf1bdbab9fec4ac9a1cfa923dcb30c2911c2ce7
MD5 611e8f923f63926fde6c890fa8f0f213
BLAKE2b-256 d7658cbd828e0c6c15a0854cee1e53fe8fcafc60ec700f04dce784c8425e7395

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page