Human Readable Regular Expression
Project description
myre - Human Readable Regular Expression
myre is a composable regular expression library that transforms regex patterns into first-class objects supporting algebraic operations. It provides a human-friendly way to combine, filter, and transform pattern matching in Python, maintaining full compatibility with the standard re module interface.
Features
- Algebraic Operations: Combine patterns using
|(OR),&(AND),^(XOR/EXCLUDE),@(MASK) - Type-Safe: Protocol-based design with full type hints
- re-Compatible: Drop-in replacement for standard
remodule patterns - Composable: Patterns are immutable objects that can be combined like mathematical sets
- Zero Dependencies: Only requires
typing-extensionsfor modern type hints
Installation
pip install myre
Quick Start
import myre
# OR operation: match any pattern (default)
pattern = myre.compile(r"hello", r"world", r"foo")
for match in pattern.finditer("hello world foo"):
print(match.group()) # Prints: hello, world, foo
# AND operation: all patterns must match
from myre import Mode
pattern = myre.compile(r"hello", r"world", mode=Mode.ALL)
for match in pattern.finditer("hello world"):
print(match.group()) # Prints: hello world
# SEQUENCE operation: match patterns in order
pattern = myre.compile(r"hello") + myre.compile(r"world")
for match in pattern.finditer("hello world"):
print(match.group()) # Prints: hello world
# XOR operation: match pattern but exclude another
pattern = myre.compile(r"h\wllo") ^ r"ell"
for match in pattern.finditer("Hello Hallo Hillo"):
print(match.group()) # Prints: Hallo, Hillo (excludes "Hello")
# MASK operation: ignore specific patterns by masking with placeholders
pattern = myre.compile(r"test_{3}value") @ (r"\d+", "_")
for match in pattern.finditer("test123value"):
print(match.group()) # Prints: test123value
Core Concepts
Pattern Types
MatchAny (Mode.ANY - default)
Matches any of the given patterns, avoiding overlapping matches:
from myre import Mode
pattern = myre.compile(r"abc", r"bcd", r"cde", mode=Mode.ANY)
# or simply: pattern = myre.compile(r"abc", r"bcd", r"cde")
for match in pattern.finditer("abc bcd cde"):
print(match.group()) # Non-overlapping matches
MatchALL (Mode.ALL)
Requires all patterns to match somewhere in the search range. Returns a combined match:
pattern = myre.compile(r"hello", r"world", mode=Mode.ALL)
for match in pattern.finditer("hello world"):
print(match.group()) # "hello world"
Note: MatchALL collects the first match from each pattern and combines them into a single match result.
Operators
| Operator | Name | Description | Example |
|---|---|---|---|
| |
OR | Match any pattern | p1 | p2 → MatchAny(p1, p2) |
& |
AND | Match all patterns | p1 & p2 → MatchALL(p1, p2) |
+ |
SEQUENCE | Match patterns in order | p1 + p2 → MatchSeq(p1, p2) |
^ |
XOR | Exclude matches containing deny pattern | p1 ^ p2 → Match p1, exclude if text contains p2 |
@ |
MASK | Mask patterns with placeholders | p1 @ (mask, placeholder) → Replace mask with placeholder, match p1 |
Operator Usage Notes
^ (XOR): Currently operates at the string level - if the entire search string contains the deny pattern, all matches are rejected. This is different from traditional set XOR.
@ (MASK): Replaces matched text with placeholder characters (preserving length), then matches the pattern. Returns the original text (not masked text). Choose a placeholder that works with your pattern:
- Use
'x'if pattern usesx+orx{N}to match placeholders - Use
'_'if pattern uses_{3}etc. - Default is
'.'(dot character)
API Reference
PatternLike Protocol
All pattern objects implement the PatternLike protocol:
class PatternLike(Protocol):
def search(self, string: str, pos: int = 0, endpos: int = sys.maxsize) -> Optional[MatchLike]: ...
def findall(self, string: str, pos: int = 0, endpos: int = sys.maxsize) -> List[str]: ...
def finditer(self, string: str, pos: int = 0, endpos: int = sys.maxsize) -> Iterator[MatchLike]: ...
TODO: Standard library methods not yet implemented:
match()- Match at the beginning of the stringfullmatch()- Match the entire stringsplit()- Split string by pattern occurrencessub()/subn()- Replace pattern matchesflags- Access pattern flagspattern- Access pattern stringgroupindex- Access group name mapping
MatchLike Protocol
All match objects implement the MatchLike protocol:
class MatchLike(Protocol):
re: PatternLike
string: str
def start(self, group: int | str = 0) -> int: ...
def end(self, group: int | str = 0) -> int: ...
def span(self, group: int | str = 0) -> Tuple[int, int]: ...
def group(self, group: int | str = 0) -> str: ...
def groups(self) -> tuple[str, ...]: ...
TODO: Standard library methods not yet implemented:
groupdict()- Return dict of named groupsexpand()- Expand template using groupslastindex- Last matched group indexlastgroup- Last matched group namepos/endpos- Search boundaries
Pattern Classes
Base
@dataclass
class Base:
p_trim: PatternLike # Pre-process filter
p_deny: PatternLike # Post-process filter
def finditer(self, string: str, pos: int = 0, endpos: int = sys.maxsize) -> Iterator[MatchLike]
def findall(self, string: str, pos: int = 0, endpos: int = sys.maxsize) -> List[str]
def search(self, string: str, pos: int = 0, endpos: int = sys.maxsize) -> Optional[MatchLike]
MatchAny
@dataclass
class MatchAny(Base):
patterns: tuple[PatternLike, ...]
@classmethod
def compile(cls, *patterns: Any, flag: int = 0) -> Self
MatchALL
@dataclass
class MatchALL(MatchAny):
patterns: tuple[PatternLike, ...]
@classmethod
def compile(cls, *patterns: Any, flag: int = 0) -> Self
Advanced Usage
Working with re.Pattern
import re
import myre
# Mix regex strings and compiled patterns
regex1 = re.compile(r"hello")
pattern = myre.compile(regex1, r"world", flag=re.IGNORECASE)
Nested Operations
# Complex pattern: (hello OR world) BUT NOT (foo OR bar)
base = myre.compile(r"hello", r"world")
exclude = myre.compile(r"foo", r"bar")
pattern = base ^ exclude
Position-Aware Matching
# Mask punctuation before matching
pattern = myre.compile(r"\bhello\b") @ (r"[.,!?;:]", " ")
for match in pattern.finditer("Hello, world! Hello."):
print(match.group()) # Prints both "Hello," and "Hello."
Architecture
The library is built on three architectural layers:
- Protocol Layer (
protocol.py): Defines structural types for pattern/match compatibility - Pattern Layer (
pattern.py): Implements algebraic operations using Template Method pattern - Match Layer (
match.py): Provides position-aware match result wrappers
Key Design Patterns
- Template Method:
Base.finditer()orchestrates the matching pipeline - Protocol-based Polymorphism: Structural subtyping over nominal inheritance
- Immutable Operators: All operations create new objects via
deepcopy
Development
# Install dependencies
poetry install
# Run tests
pytest tests/
# Type checking
mypy myre/
# Linting and formatting
ruff check .
ruff format .
Contributing
Contributions are welcome! Please ensure:
- All tests pass:
pytest tests/ - Type checking passes:
mypy myre/ - Code is formatted:
ruff format .
License
MIT License - see LICENSE for details.
Acknowledgments
Inspired by the need for composable, human-readable regular expressions in Python.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file myre-0.0.8.tar.gz.
File metadata
- Download URL: myre-0.0.8.tar.gz
- Upload date:
- Size: 55.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.24 {"installer":{"name":"uv","version":"0.9.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
56a8e88b8b64d133456e2846b2a881d5ff84e2b7f148c4225c08ba12b4221026
|
|
| MD5 |
32a65499d1239e37dcf229b335eb393a
|
|
| BLAKE2b-256 |
ed87e496e20c494c17bea8f50e9b596c11020706fd8b8a0827d838dc2372ca99
|
File details
Details for the file myre-0.0.8-py3-none-any.whl.
File metadata
- Download URL: myre-0.0.8-py3-none-any.whl
- Upload date:
- Size: 10.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.24 {"installer":{"name":"uv","version":"0.9.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
50d6f85b531c3e9a3f95af67d41eb1338fb1c08562c4ed222303b26f75601dae
|
|
| MD5 |
842591b8838d9204a21c03abfcb255ec
|
|
| BLAKE2b-256 |
5dc8ae481dbe9f7953ad74d61d01c8bff5cd34092f4a79b38c030a2873eead2d
|