A Python package for detecting garbled text using multiple detection strategies with a scikit-learn-like interface
Project description
pygarble
Detect gibberish, garbled text, and nonsense with high precision.
A zero-dependency Python library for identifying random character sequences, keyboard mashing, encoding errors, and other forms of text corruption. Uses statistical analysis, phonotactic rules, and pattern matching to distinguish meaningful text from gibberish.
Installation
pip install pygarble
Quick Start
from pygarble import GarbleDetector, EnsembleDetector, Strategy
# Recommended: Use the default ensemble (99.2% precision, 85.6% recall)
detector = EnsembleDetector()
detector.predict("Hello world") # False - valid text
detector.predict("asdfghjkl") # True - keyboard mashing
detector.predict("qxzjkwp") # True - impossible letter combinations
# Get probability scores (0.0 = valid, 1.0 = gibberish)
detector.predict_proba("Hello world") # ~0.1
detector.predict_proba("xkqzjwp") # ~0.9
# Batch processing
texts = ["Hello world", "asdfghjkl", "Normal sentence here"]
results = detector.predict(texts) # [False, True, False]
Performance
Tested on 1,644 samples (dictionary words, sentences, random strings, keyboard mashing) via regression/benchmark.py:
| Detector | Precision | Recall | F1 Score |
|---|---|---|---|
| EnsembleDetector() | 99.2% | 85.6% | 91.9% |
| MARKOV_CHAIN | 99.2% | 84.3% | 91.2% |
| NGRAM_FREQUENCY | 97.6% | 75.0% | 84.8% |
| LOG_LIKELIHOOD_RATIO | 100% | 63.4% | 77.6% |
| WORD_ANOMALY | 100% | 52.8% | 69.1% |
The default ensemble is the union (voting="any") of MARKOV_CHAIN with two strategies that produced zero false positives on the benchmark (LOG_LIKELIHOOD_RATIO, WORD_ANOMALY), so each member only adds recall. It strictly dominates any single strategy.
Two useful variants:
# Strict precision: majority vote drove false positives to 0 on the
# benchmark (100% precision, 74.9% recall) at the cost of recall
detector = EnsembleDetector(
strategies=[
Strategy.MARKOV_CHAIN, Strategy.NGRAM_FREQUENCY, Strategy.WORD_LOOKUP,
Strategy.LOG_LIKELIHOOD_RATIO, Strategy.WORD_ANOMALY,
],
voting="majority",
)
# Specialist coverage: adds detection of hashes, mojibake, repeated junk,
# and homoglyph attacks that the character-model strategies don't target
detector = EnsembleDetector(
strategies=[
Strategy.LOG_LIKELIHOOD_RATIO, Strategy.WORD_ANOMALY,
Strategy.KEYBOARD_ADJACENCY, Strategy.MOJIBAKE,
Strategy.HEX_STRING, Strategy.REPETITION, Strategy.UNICODE_SCRIPT,
],
voting="any",
)
Detection Strategies
Recommended Strategies
| Strategy | Description | Precision |
|---|---|---|
MARKOV_CHAIN |
Character transition probabilities trained on English | 99.2% |
NGRAM_FREQUENCY |
Common English trigram analysis | 97.6% |
LOG_LIKELIHOOD_RATIO |
English-vs-random two-model bigram comparison | 100% |
WORD_ANOMALY |
Per-word scoring; catches one garbage token in a valid sentence | 100% |
WORD_LOOKUP |
Dictionary of 49K English words | high recall |
All Available Strategies
Statistical Models (v0.7.0+)
LOG_LIKELIHOOD_RATIO- Log-likelihood ratio of English vs uniform character models (length-normalized)WORD_ANOMALY- Fraction of individually-anomalous words; robust to garbage embedded in valid textKEYBOARD_ADJACENCY- Physical key-adjacency walks (catches mash the trigram lists miss)
High Precision (v0.5.0)
BIGRAM_PROBABILITY- Impossible letter pairsLETTER_POSITION- Invalid letter positionsCONSONANT_SEQUENCE- Too many consecutive consonantsVOWEL_PATTERN- Invalid vowel sequencesLETTER_FREQUENCY- Abnormal letter distributionRARE_TRIGRAM- Impossible trigrams
Core Strategies
MARKOV_CHAIN- Character-level Markov chain (best overall)NGRAM_FREQUENCY- Trigram frequency analysisWORD_LOOKUP- English dictionary lookupPRONOUNCEABILITY- English phonotactic rulesKEYBOARD_PATTERN- Keyboard row sequencesENTROPY_BASED- Shannon entropy analysisVOWEL_RATIO- Vowel to consonant ratio
Specialized Detectors
MOJIBAKE- Encoding corruption (UTF-8 as Latin-1)UNICODE_SCRIPT- Homoglyph/script mixing attacksHEX_STRING- Hash strings and UUIDsSYMBOL_RATIO- Excessive symbols/numbersREPETITION- Repeated patterns (ababab, repeated words)
Pattern Heuristics
PATTERN_MATCHING- Configurable regex patterns (keyboard rows, repeated/alternating chars)
Removed in v0.8.0:
CHARACTER_FREQUENCY,WORD_LENGTH,STATISTICAL_ANALYSIS,COMPRESSION_RATIO, andENGLISH_WORD_VALIDATION(the only strategy requiring a third-party dependency). All were dominated by the strategies above;WORD_LOOKUPreplacesENGLISH_WORD_VALIDATIONdependency-free.
Using Individual Strategies
from pygarble import GarbleDetector, Strategy
# Markov chain - best overall performance
detector = GarbleDetector(Strategy.MARKOV_CHAIN)
detector.predict("the quick brown fox") # False
detector.predict("xkqzjwpmv") # True
# High precision - zero false positives
detector = GarbleDetector(Strategy.BIGRAM_PROBABILITY)
detector.predict("hello world") # False
detector.predict("qxjjxz") # True (impossible: qx, jj, xz)
# Encoding corruption detection
detector = GarbleDetector(Strategy.MOJIBAKE)
detector.predict("Café") # False - valid UTF-8
detector.predict("Café") # True - mojibake
# Homoglyph attack detection
detector = GarbleDetector(Strategy.UNICODE_SCRIPT)
detector.predict("paypal") # False - all Latin
detector.predict("pаypal") # True - Cyrillic 'а'
Ensemble Detector
Combine multiple strategies for better accuracy:
from pygarble import EnsembleDetector, Strategy
# Default ensemble (recommended)
# Uses: MARKOV_CHAIN, LOG_LIKELIHOOD_RATIO, WORD_ANOMALY
# Voting: "any" - the two companions had zero benchmark false positives,
# so they only add recall on top of MARKOV_CHAIN
detector = EnsembleDetector()
# Custom strategies
detector = EnsembleDetector(
strategies=[
Strategy.MARKOV_CHAIN,
Strategy.BIGRAM_PROBABILITY,
Strategy.KEYBOARD_PATTERN,
]
)
# Different voting modes (default: "any" for the built-in strategy set,
# "majority" when you pass a custom strategies list)
detector = EnsembleDetector(voting="any") # High recall - flag if ANY strategy detects
detector = EnsembleDetector(voting="all") # High precision - flag only if ALL agree
detector = EnsembleDetector(voting="majority") # Balanced
detector = EnsembleDetector(voting="average") # Average probabilities
# Weighted voting
detector = EnsembleDetector(
strategies=[Strategy.MARKOV_CHAIN, Strategy.WORD_LOOKUP],
voting="weighted",
weights=[0.7, 0.3]
)
API Reference
GarbleDetector
GarbleDetector(
strategy: Strategy,
threshold: float = 0.5, # Probability threshold for predict()
**kwargs # Strategy-specific parameters
)
# Methods
detector.predict(text) # Returns bool or List[bool]
detector.predict_proba(text) # Returns float or List[float] (0.0-1.0)
EnsembleDetector
EnsembleDetector(
strategies: List[Strategy] = None, # Default: MARKOV_CHAIN + LLR + WORD_ANOMALY
threshold: float = 0.5,
voting: str = None, # "majority", "any", "all", "average", "weighted"
# default: "any" (built-in set) / "majority" (custom set)
weights: List[float] = None, # Required if voting="weighted"
)
# Methods (same as GarbleDetector)
detector.predict(text)
detector.predict_proba(text)
Common Use Cases
Filter User Input
detector = EnsembleDetector()
def validate_input(text):
if detector.predict(text):
return "Please enter valid text"
return None
Clean Data Pipeline
detector = GarbleDetector(Strategy.MARKOV_CHAIN)
clean_data = [text for text in raw_data if not detector.predict(text)]
Detect Encoding Issues
detector = GarbleDetector(Strategy.MOJIBAKE)
for text in documents:
if detector.predict(text):
print(f"Encoding issue detected: {text[:50]}...")
Detect Phishing/Homoglyphs
detector = GarbleDetector(Strategy.UNICODE_SCRIPT)
if detector.predict(domain_name):
print("Warning: Possible homoglyph attack")
Requirements
- Python 3.8+
- Zero dependencies
Development
git clone https://github.com/brightertiger/pygarble.git
cd pygarble
pip install -e ".[dev]"
pytest tests/ -v
License
MIT License
Changelog
0.8.0
- Breaking: removed legacy strategies CHARACTER_FREQUENCY, WORD_LENGTH, STATISTICAL_ANALYSIS, COMPRESSION_RATIO, ENGLISH_WORD_VALIDATION (and the
spellcheckerextra) - New default ensemble: MARKOV_CHAIN | LOG_LIKELIHOOD_RATIO | WORD_ANOMALY with
voting="any"(99.2% precision, 85.6% recall - strictly dominates any single strategy) votingnow defaults to "any" for the built-in strategy set, "majority" for custom sets
0.7.0
- Fixed ~45 verified bugs across all strategies (false positives on accented text, y-vowel words, proper nouns, Japanese, URLs, formatted text; false negatives on ALL-CAPS gibberish, cp1252 mojibake, repeated words)
- 3 new strategies: LOG_LIKELIHOOD_RATIO, WORD_ANOMALY, KEYBOARD_ADJACENCY
- Ensemble abstention: word-level strategies no longer dilute votes on short text
- Consistent TypeError contract; cleaned 670 junk entries from the word list
0.5.0
- 6 new high-precision strategies (BIGRAM_PROBABILITY, LETTER_POSITION, CONSONANT_SEQUENCE, VOWEL_PATTERN, LETTER_FREQUENCY, RARE_TRIGRAM)
- Redesigned default ensemble for 99.5% precision
- External validation benchmark (1,644 test cases)
0.4.0
- Added COMPRESSION_RATIO, MOJIBAKE, PRONOUNCEABILITY, UNICODE_SCRIPT strategies
0.3.0
- Zero-dependency core with embedded training data
- Added MARKOV_CHAIN, NGRAM_FREQUENCY, WORD_LOOKUP strategies
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pygarble-0.8.0.tar.gz.
File metadata
- Download URL: pygarble-0.8.0.tar.gz
- Upload date:
- Size: 238.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
223703c6a3f1e24641162edcc8e0c6acf791a417201105baefab2ba856ebe667
|
|
| MD5 |
ed889672ec88b3c9657a0f2993dbd50f
|
|
| BLAKE2b-256 |
d509611ce2102229455b54bd3cde7872caa32c54094bed6d92aaad6149f95d0d
|
File details
Details for the file pygarble-0.8.0-py3-none-any.whl.
File metadata
- Download URL: pygarble-0.8.0-py3-none-any.whl
- Upload date:
- Size: 227.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b01c53861d55427eb3f48f9ea451264a6fb1398d8432dfed0e119b3daeab7f93
|
|
| MD5 |
754b44e16c168cb924bf3c96d362080b
|
|
| BLAKE2b-256 |
e9a4225d64bfd3f996f0688b8d481fe59cff29f11de809290501e6e1b903c1d5
|