This release is a pre-release and may not be stable for production use.
cerberos
A local, Python-based tool for approximate homology-aware splitting and descriptive diagnostics of peptide/protein FASTA datasets.
Cerberos reads a FASTA file, generates approximate sequence-similarity clusters, and assigns complete clusters to train, validation, and test splits. It can also describe an existing set of three splits. It is designed to help identify obvious near-duplicate leakage. It is not a substitute for a validated alignment-based homology search and does not guarantee leakage-free evaluation.
The package requires NumPy. Plotting is an optional Matplotlib extra. KS and Wasserstein descriptive distances are computed with NumPy and do not require SciPy. Cerberos runs locally and does not call external services or external sequence-search binaries.
Contents
- Installation
- Quick start
- Input and output
- CLI reference
- Method and interpretation
- Scientific definitions and limitations
- Python API
- Development and tests
- References
Installation
Python 3.10 or later is supported. The PyPI distribution name is cerberos-peptide-splitter, the Python import name is cerberos_peptide_splitter, and the installed command is cerberos-peptide-splitter. This project has not yet been published to PyPI; install from a source checkout until the first release.
After the first release to PyPI:
python -m pip install cerberos-peptide-splitter
Optional plotting support:
python -m pip install 'cerberos-peptide-splitter[plots]'
Optional full runtime extras:
python -m pip install 'cerberos-peptide-splitter[full]'
Optional developer toolchain:
python -m pip install 'cerberos-peptide-splitter[dev]'
To work from a source checkout:
git clone https://github.com/AI-Meat-Lab/cerberos.git
cd cerberos
python -m venv .venv
# Linux/macOS: source .venv/bin/activate
# Windows: .venv\Scripts\activate
python -m pip install -e '.[dev]'
Quick start
Split one FASTA file using defaults:
cerberos-peptide-splitter split --input peptides.fasta --output-dir splits/
For a sensitivity-oriented reduced-alphabet pass followed by exact raw-sequence edit-distance verification:
cerberos-peptide-splitter split \
--input peptides.fasta \
--reduced-alphabet groups5 \
--verify-with levenshtein \
--similarity-threshold 0.80 \
--prefilter-threshold 0.30 \
--kmer-sizes 2,3 \
--output-dir splits/
Audit an existing train/validation/test directory:
cerberos-peptide-splitter audit --dir splits/ --output-dir audit/
Use python -m cerberos_peptide_splitter instead of cerberos-peptide-splitter if running from an unpacked source tree without installing the console entry point.
Input and output
FASTA input
Records may be unlabeled or have a binary label as the final header field:
>peptide_1|1
KWKLFKKIEKVGQNIRDGIIK
>peptide_2|0
MKTIIALSYIFCLVFA
Rules:
- A final
|0or|1is parsed as a label; other header text is retained as the identifier. - Sequence lines are uppercased and all whitespace within them is removed.
- Empty records are ignored. Sequence text before the first
>header and empty headers are errors. - Other symbols, such as
X,-, or*, are retained and are not validated against a particular alphabet. With a reduced alphabet enabled, every non-canonical symbol is mapped to the sameXgroup. - Mixed labeled and unlabeled files are accepted. Label stratification is used only when every input record has a label.
Outputs
split writes train.fasta, val.fasta, test.fasta, report.txt, stats.json, and, when Matplotlib is installed, a plots/ directory. audit reads FASTAs and writes only report, statistics, and plots; it does not overwrite the input FASTAs.
The audit command recognizes train.fasta/.fa/.faa, val.fasta/.fa/.faa or valid.fasta/.fa, and test.fasta/.fa/.faa.
stats.json contains the effective configuration, candidate-generation counts and limits for split runs, split sizes and label counts, length and feature summaries, sampled pairwise k-mer Jaccard summaries, empirical KS D statistics, Wasserstein distances, and heuristic warnings. Empty samples are represented with JSON null, never non-standard NaN literals.
For large files, reduce the search ceiling explicitly, for example with --max-candidate-pairs 100000. This bounds candidate-pair verification, but a smaller value increases the chance that related sequences are not grouped. Always check candidate_generation in stats.json and the warning in report.txt to see whether sampling or the cap affected a run.
CLI reference
cerberos-peptide-splitter split
| Option | Default | Meaning |
|---|---|---|
--input PATH |
required | Input FASTA. |
--output-dir DIR |
cerberos_out |
Output directory. |
--seed N |
42 |
Integer seed in NumPy's supported 32-bit range. |
--train-pct F |
0.80 |
Train proportion. The three non-negative proportions are normalized to sum to one. |
--val-pct F |
0.10 |
Validation proportion. |
--test-pct F |
0.10 |
Test proportion. |
--no-clustering |
off | Use random splitting instead of grouping candidate homologs. |
--similarity-threshold F |
0.70 |
MinHash estimate threshold if no verifier is selected; final score threshold otherwise. Must be in [0, 1]. |
--kmer-sizes CSV |
3 |
Positive comma-separated k-mer sizes. |
--num-hashes N |
128 |
MinHash signature size, at least 4. |
--max-candidate-pairs N |
250000 |
Hard cap on candidate pairs examined across configured k-mer searches. Dense buckets are sampled deterministically. Lower caps use less time and memory but can miss related pairs. |
--reduced-alphabet NAME |
none |
none, groups5, or groups7; transform residues before candidate generation. |
--verify-with NAME |
none |
none, kmer-exact, or levenshtein. |
--prefilter-threshold F |
0.30 |
MinHash estimate threshold before deterministic verification; cannot exceed the final threshold. |
cerberos-peptide-splitter audit
| Option | Default | Meaning |
|---|---|---|
--dir DIR |
required | Directory containing the three split FASTAs. |
--output-dir DIR |
cerberos_out |
Diagnostics output directory. |
--seed N |
42 |
Seed for reproducible diagnostic sampling. |
--kmer-sizes CSV |
3 |
Positive k-mer sizes; diagnostics use the first value. |
Invalid CLI values and missing input files are reported as concise command-line errors. The process exits non-zero on failure.
Method and interpretation
- Optional alphabet transform.
groups5maps aliphatic and aromatic residues together, polar residues together, basic residues together, acidic residues together, and G/P/C together.groups7keeps aliphatic, aromatic, polar, basic, acidic, G/P, and C groups separate. These are heuristic groupings, not a substitution matrix or evolutionary model. - K-mer sketches. For each configured k, unique contiguous k-mers are hashed into seeded MinHash signatures. CRC32 provides stable input hashes. The observed sketch agreement estimates set Jaccard similarity.
- Bounded LSH candidate generation. The 128-hash default uses 32 bands of four rows. One band's buckets are processed at a time instead of retaining all pairs in memory. Under the ideal independent-MinHash model, a pair with transformed k-mer Jaccard
swould be proposed with probability approximately1 - (1 - s^4)^32. This is a model-based retrieval probability, not a guarantee. Finite hash families, a partial final band for non-multiples of four, dense-bucket sampling, and the candidate budget change realized recall. Buckets of at most 128 sequences are fully enumerated; larger buckets use a deterministic sampled neighborhood. A default budget of 250,000 candidate pairs, deduplicated within each k search, is distributed over the configured k values. Dense-bucket sampling and budget exhaustion are reported in the log,stats.json, andreport.txt. This avoids quadratic pair-set memory and limits verification work, but it can miss related pairs. A low prefilter threshold does not make candidate recall exhaustive. - Optional deterministic scoring.
kmer-exactcalculates exact unique-k-mer Jaccard on transformed sequences and takes the maximum across configured k values.levenshteincalculates global unit-cost edit similarity on raw sequences,1 - edit_distance / max(len(seq_a), len(seq_b)). It is not alignment-based percent identity. Verification only scores pairs already proposed by LSH. - Connected components. Accepted pair edges are merged with Union-Find. A connected component may include a pair of endpoints whose direct score is below the edge threshold, due to transitivity.
- Whole-cluster split assignment. Clusters are greedily assigned to minimize deviation from target split sizes. Preserving clusters takes priority over exact proportions. Clustered assignment does not stratify clusters by label. If
--no-clusteringis used, random splitting is applied; when every record has a label, label stratification is used. - Descriptive diagnostics. Raw sequences are used for pairwise k-mer distances and biochemical features. Pairwise similarity is sampled per split for cost control. Feature distances are descriptive, not inferential tests of model generalization.
Scientific definitions and limitations
- Jaccard: unique-k-mer set intersection divided by union. Repeated occurrences do not receive extra weight. A sequence shorter than k has no k-mers; Cerberos treats that as no similarity evidence rather than declaring two empty sets a match.
- Reduced alphabets: grouping changes the sequence representation and can increase similarity among chemically grouped residues. It can also merge unrelated sequences. It is a heuristic sensitivity/specificity tradeoff; evaluate settings for the target domain.
- Levenshtein score: global edit distance gives substitutions, insertions, and deletions unit cost. The normalized score is length-dependent, does not use BLOSUM/PAM scores, does not model local alignment, and should not be reported as biological percent identity.
- MinHash and LSH: MinHash agreement approximates Jaccard under assumptions about the hash family; LSH is a probabilistic candidate filter. Dense buckets are sampled and the pair budget may stop a run early; either can omit a pair. The report flags those cases. A missed candidate cannot be recovered by verification. For high-assurance leakage control, use an alignment-based clustering or search method and inspect thresholds on a representative validation set.
- Clusters and thresholds: threshold edges are transitively closed; therefore maximum pairwise within-cluster similarity is not guaranteed to meet the threshold. The cluster is an operational grouping for split assignment, not a statement that all members are homologous.
- Hydropathy: mean Kyte-Doolittle score over canonical residues, using the original residue scale. This is a sequence summary, not a membrane topology prediction.
- Charge: approximate peptide net charge at pH 7.0, including free termini and standard approximate side-chain pKa values via Henderson-Hasselbalch fractions. Local environment, terminal modifications, pKa shifts, and non-canonical residues are not modeled.
- Molecular mass: average free-amino-acid masses are summed and 18.01528 Da is subtracted per peptide bond. Unknown symbols are assigned an approximate 110 Da free-residue mass. This is not monoisotopic mass and does not account for modifications, cyclization, disulfide formation, or unusual termini.
- Unknown symbols: canonical frequencies and group fractions are per total sequence length;
unknown_fractionreports positions outside the 20 canonical amino acids. Hydropathy averages canonical residues only. - KS and Wasserstein: the report stores empirical two-sample KS D and 1-D first Wasserstein distance per feature. No p-values or multiple-testing correction are computed. The
>0.20warning is a heuristic and depends on sample size, feature scaling, and data context. - No leakage guarantee: no warning is not proof of independence, no distribution test establishes OOD safety, and this software has not been validated as a clinical or regulatory assay.
Python API
from cerberos_peptide_splitter import RunConfig, parse_fasta, split_records
records = parse_fasta("peptides.fasta")
config = RunConfig(
train_pct=0.8,
val_pct=0.1,
test_pct=0.1,
kmer_sizes=[2, 3],
reduced_alphabet="groups5",
verify_with="levenshtein",
similarity_threshold=0.8,
prefilter_threshold=0.3,
seed=42,
)
splits = split_records(records, config, verbose=False)
print({name: len(items) for name, items in splits.items()})
Top-level exports also include write_fasta, assign_clusters, stratified_random_split, summarize, build_report, run_split, and run_audit. Lower-level modules expose the reduced-alphabet, k-mer, MinHash, and edit-distance helpers.
Development and tests
python -m pip install -e '.[dev]'
pytest -q
pytest -q -m 'not optional' # NumPy-only CI subset
ruff check cerberos_peptide_splitter tests
ruff format --check cerberos_peptide_splitter tests
black --check cerberos_peptide_splitter tests
isort --check-only cerberos_peptide_splitter tests
mypy cerberos_peptide_splitter --ignore-missing-imports --no-strict-optional
python -m build
python -m twine check dist/*
GitHub Actions checks a NumPy-only test path and a full-dependency test path across Ubuntu, macOS, and Windows. It exercises all 3 × 3 alphabet/verifier combinations, runs CLI and reproducibility smoke tests, builds distributions, and runs formatting, lint, and type checks. CI results, not this README, are the source of truth for currently passing environments.
References
- Kyte, J. & Doolittle, R. F. (1982). A simple method for displaying the hydropathic character of a protein. Journal of Molecular Biology 157(1), 105–132. doi:10.1016/0022-2836(82)90515-0; PubMed.
- LibreTexts, What is a protein? — peptide bond formation, residue masses, and water loss. Section 1A.
- Grimsley, G. R., Scholtz, J. M. & Pace, C. N. (2009). A summary of the measured pK values of the ionizable groups in folded proteins. Protein Science 18, 247–251. doi:10.1002/pro.19; PMC.
Release maintainers
The release workflow publishes through PyPI Trusted Publishing (OIDC). Before the first release, configure the PyPI project cerberos-peptide-splitter with a trusted publisher for the GitHub repository AI-Meat-Lab/cerberos and workflow .github/workflows/release.yml. Releases use v* tags and fail unless the tag exactly matches project.version in pyproject.toml.
License and project links
MIT licensed; see LICENSE. Report issues or contribute through the GitHub repository, issue tracker, and pull requests.
Metadata
Release files for cerberos-peptide-splitter 0.0.1b1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| cerberos_peptide_splitter-0.0.1b1.tar.gz | 48.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cerberos_peptide_splitter-0.0.1b1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 82.7 kB
Release files / cerberos_peptide_splitter-0.0.1b1.tar.gz
| Download URL | cerberos_peptide_splitter-0.0.1b1.tar.gz |
|---|---|
| Size | 48.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
23b2b443447a59620075d071bf340b51db9866c6d8926d740274068cb5cab046
|
|
BLAKE2b-256 checksum How to use checksums |
851b709be08ef7686d344c0d08c9a19086b8ee63f9b47985179c530efd6832ae
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|
Release files / cerberos_peptide_splitter-0.0.1b1-py3-none-any.whl
| Download URL | cerberos_peptide_splitter-0.0.1b1-py3-none-any.whl |
|---|---|
| Size | 34.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e3cd07fe795d87e410573282fd6afd6b78216cfe4eb90453eed21e916ece6e35
|
|
BLAKE2b-256 checksum How to use checksums |
89fd81babe0d3888aa5d8f614304e5a2e58d181f068a0d6859d7ea9e8dd1e85c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|