Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

cerberos

A local, Python-based tool for approximate homology-aware splitting and descriptive diagnostics of peptide/protein FASTA datasets.

CI Lint CodeQL Python License: MIT

Cerberos reads a FASTA file, generates approximate sequence-similarity clusters, and assigns complete clusters to train, validation, and test splits. It can also describe an existing set of three splits. It is designed to help identify obvious near-duplicate leakage. It is not a substitute for a validated alignment-based homology search and does not guarantee leakage-free evaluation.

The package requires NumPy. Plotting is an optional Matplotlib extra. KS and Wasserstein descriptive distances are computed with NumPy and do not require SciPy. Cerberos runs locally and does not call external services or external sequence-search binaries.

Contents

Installation

Python 3.10 or later is supported. The PyPI distribution name is cerberos-peptide-splitter, the Python import name is cerberos_peptide_splitter, and the installed command is cerberos-peptide-splitter. This project has not yet been published to PyPI; install from a source checkout until the first release.

After the first release to PyPI:

python -m pip install cerberos-peptide-splitter

Optional plotting support:

python -m pip install 'cerberos-peptide-splitter[plots]'

Optional full runtime extras:

python -m pip install 'cerberos-peptide-splitter[full]'

Optional developer toolchain:

python -m pip install 'cerberos-peptide-splitter[dev]'

To work from a source checkout:

git clone https://github.com/AI-Meat-Lab/cerberos.git
cd cerberos
python -m venv .venv
# Linux/macOS: source .venv/bin/activate
# Windows:    .venv\Scripts\activate
python -m pip install -e '.[dev]'

Quick start

Split one FASTA file using defaults:

cerberos-peptide-splitter split --input peptides.fasta --output-dir splits/

For a sensitivity-oriented reduced-alphabet pass followed by exact raw-sequence edit-distance verification:

cerberos-peptide-splitter split \
  --input peptides.fasta \
  --reduced-alphabet groups5 \
  --verify-with levenshtein \
  --similarity-threshold 0.80 \
  --prefilter-threshold 0.30 \
  --kmer-sizes 2,3 \
  --output-dir splits/

Audit an existing train/validation/test directory:

cerberos-peptide-splitter audit --dir splits/ --output-dir audit/

Use python -m cerberos_peptide_splitter instead of cerberos-peptide-splitter if running from an unpacked source tree without installing the console entry point.

Input and output

FASTA input

Records may be unlabeled or have a binary label as the final header field:

>peptide_1|1
KWKLFKKIEKVGQNIRDGIIK
>peptide_2|0
MKTIIALSYIFCLVFA

Rules:

  • A final |0 or |1 is parsed as a label; other header text is retained as the identifier.
  • Sequence lines are uppercased and all whitespace within them is removed.
  • Empty records are ignored. Sequence text before the first > header and empty headers are errors.
  • Other symbols, such as X, -, or *, are retained and are not validated against a particular alphabet. With a reduced alphabet enabled, every non-canonical symbol is mapped to the same X group.
  • Mixed labeled and unlabeled files are accepted. Label stratification is used only when every input record has a label.

Outputs

split writes train.fasta, val.fasta, test.fasta, report.txt, stats.json, and, when Matplotlib is installed, a plots/ directory. audit reads FASTAs and writes only report, statistics, and plots; it does not overwrite the input FASTAs.

The audit command recognizes train.fasta/.fa/.faa, val.fasta/.fa/.faa or valid.fasta/.fa, and test.fasta/.fa/.faa.

stats.json contains the effective configuration, candidate-generation counts and limits for split runs, split sizes and label counts, length and feature summaries, sampled pairwise k-mer Jaccard summaries, empirical KS D statistics, Wasserstein distances, and heuristic warnings. Empty samples are represented with JSON null, never non-standard NaN literals.

For large files, reduce the search ceiling explicitly, for example with --max-candidate-pairs 100000. This bounds candidate-pair verification, but a smaller value increases the chance that related sequences are not grouped. Always check candidate_generation in stats.json and the warning in report.txt to see whether sampling or the cap affected a run.

CLI reference

cerberos-peptide-splitter split

Option Default Meaning
--input PATH required Input FASTA.
--output-dir DIR cerberos_out Output directory.
--seed N 42 Integer seed in NumPy's supported 32-bit range.
--train-pct F 0.80 Train proportion. The three non-negative proportions are normalized to sum to one.
--val-pct F 0.10 Validation proportion.
--test-pct F 0.10 Test proportion.
--no-clustering off Use random splitting instead of grouping candidate homologs.
--similarity-threshold F 0.70 MinHash estimate threshold if no verifier is selected; final score threshold otherwise. Must be in [0, 1].
--kmer-sizes CSV 3 Positive comma-separated k-mer sizes.
--num-hashes N 128 MinHash signature size, at least 4.
--max-candidate-pairs N 250000 Hard cap on candidate pairs examined across configured k-mer searches. Dense buckets are sampled deterministically. Lower caps use less time and memory but can miss related pairs.
--reduced-alphabet NAME none none, groups5, or groups7; transform residues before candidate generation.
--verify-with NAME none none, kmer-exact, or levenshtein.
--prefilter-threshold F 0.30 MinHash estimate threshold before deterministic verification; cannot exceed the final threshold.

cerberos-peptide-splitter audit

Option Default Meaning
--dir DIR required Directory containing the three split FASTAs.
--output-dir DIR cerberos_out Diagnostics output directory.
--seed N 42 Seed for reproducible diagnostic sampling.
--kmer-sizes CSV 3 Positive k-mer sizes; diagnostics use the first value.

Invalid CLI values and missing input files are reported as concise command-line errors. The process exits non-zero on failure.

Method and interpretation

  1. Optional alphabet transform. groups5 maps aliphatic and aromatic residues together, polar residues together, basic residues together, acidic residues together, and G/P/C together. groups7 keeps aliphatic, aromatic, polar, basic, acidic, G/P, and C groups separate. These are heuristic groupings, not a substitution matrix or evolutionary model.
  2. K-mer sketches. For each configured k, unique contiguous k-mers are hashed into seeded MinHash signatures. CRC32 provides stable input hashes. The observed sketch agreement estimates set Jaccard similarity.
  3. Bounded LSH candidate generation. The 128-hash default uses 32 bands of four rows. One band's buckets are processed at a time instead of retaining all pairs in memory. Under the ideal independent-MinHash model, a pair with transformed k-mer Jaccard s would be proposed with probability approximately 1 - (1 - s^4)^32. This is a model-based retrieval probability, not a guarantee. Finite hash families, a partial final band for non-multiples of four, dense-bucket sampling, and the candidate budget change realized recall. Buckets of at most 128 sequences are fully enumerated; larger buckets use a deterministic sampled neighborhood. A default budget of 250,000 candidate pairs, deduplicated within each k search, is distributed over the configured k values. Dense-bucket sampling and budget exhaustion are reported in the log, stats.json, and report.txt. This avoids quadratic pair-set memory and limits verification work, but it can miss related pairs. A low prefilter threshold does not make candidate recall exhaustive.
  4. Optional deterministic scoring. kmer-exact calculates exact unique-k-mer Jaccard on transformed sequences and takes the maximum across configured k values. levenshtein calculates global unit-cost edit similarity on raw sequences, 1 - edit_distance / max(len(seq_a), len(seq_b)). It is not alignment-based percent identity. Verification only scores pairs already proposed by LSH.
  5. Connected components. Accepted pair edges are merged with Union-Find. A connected component may include a pair of endpoints whose direct score is below the edge threshold, due to transitivity.
  6. Whole-cluster split assignment. Clusters are greedily assigned to minimize deviation from target split sizes. Preserving clusters takes priority over exact proportions. Clustered assignment does not stratify clusters by label. If --no-clustering is used, random splitting is applied; when every record has a label, label stratification is used.
  7. Descriptive diagnostics. Raw sequences are used for pairwise k-mer distances and biochemical features. Pairwise similarity is sampled per split for cost control. Feature distances are descriptive, not inferential tests of model generalization.

Scientific definitions and limitations

  • Jaccard: unique-k-mer set intersection divided by union. Repeated occurrences do not receive extra weight. A sequence shorter than k has no k-mers; Cerberos treats that as no similarity evidence rather than declaring two empty sets a match.
  • Reduced alphabets: grouping changes the sequence representation and can increase similarity among chemically grouped residues. It can also merge unrelated sequences. It is a heuristic sensitivity/specificity tradeoff; evaluate settings for the target domain.
  • Levenshtein score: global edit distance gives substitutions, insertions, and deletions unit cost. The normalized score is length-dependent, does not use BLOSUM/PAM scores, does not model local alignment, and should not be reported as biological percent identity.
  • MinHash and LSH: MinHash agreement approximates Jaccard under assumptions about the hash family; LSH is a probabilistic candidate filter. Dense buckets are sampled and the pair budget may stop a run early; either can omit a pair. The report flags those cases. A missed candidate cannot be recovered by verification. For high-assurance leakage control, use an alignment-based clustering or search method and inspect thresholds on a representative validation set.
  • Clusters and thresholds: threshold edges are transitively closed; therefore maximum pairwise within-cluster similarity is not guaranteed to meet the threshold. The cluster is an operational grouping for split assignment, not a statement that all members are homologous.
  • Hydropathy: mean Kyte-Doolittle score over canonical residues, using the original residue scale. This is a sequence summary, not a membrane topology prediction.
  • Charge: approximate peptide net charge at pH 7.0, including free termini and standard approximate side-chain pKa values via Henderson-Hasselbalch fractions. Local environment, terminal modifications, pKa shifts, and non-canonical residues are not modeled.
  • Molecular mass: average free-amino-acid masses are summed and 18.01528 Da is subtracted per peptide bond. Unknown symbols are assigned an approximate 110 Da free-residue mass. This is not monoisotopic mass and does not account for modifications, cyclization, disulfide formation, or unusual termini.
  • Unknown symbols: canonical frequencies and group fractions are per total sequence length; unknown_fraction reports positions outside the 20 canonical amino acids. Hydropathy averages canonical residues only.
  • KS and Wasserstein: the report stores empirical two-sample KS D and 1-D first Wasserstein distance per feature. No p-values or multiple-testing correction are computed. The >0.20 warning is a heuristic and depends on sample size, feature scaling, and data context.
  • No leakage guarantee: no warning is not proof of independence, no distribution test establishes OOD safety, and this software has not been validated as a clinical or regulatory assay.

Python API

from cerberos_peptide_splitter import RunConfig, parse_fasta, split_records

records = parse_fasta("peptides.fasta")
config = RunConfig(
    train_pct=0.8,
    val_pct=0.1,
    test_pct=0.1,
    kmer_sizes=[2, 3],
    reduced_alphabet="groups5",
    verify_with="levenshtein",
    similarity_threshold=0.8,
    prefilter_threshold=0.3,
    seed=42,
)
splits = split_records(records, config, verbose=False)
print({name: len(items) for name, items in splits.items()})

Top-level exports also include write_fasta, assign_clusters, stratified_random_split, summarize, build_report, run_split, and run_audit. Lower-level modules expose the reduced-alphabet, k-mer, MinHash, and edit-distance helpers.

Development and tests

python -m pip install -e '.[dev]'
pytest -q
pytest -q -m 'not optional'  # NumPy-only CI subset
ruff check cerberos_peptide_splitter tests
ruff format --check cerberos_peptide_splitter tests
black --check cerberos_peptide_splitter tests
isort --check-only cerberos_peptide_splitter tests
mypy cerberos_peptide_splitter --ignore-missing-imports --no-strict-optional
python -m build
python -m twine check dist/*

GitHub Actions checks a NumPy-only test path and a full-dependency test path across Ubuntu, macOS, and Windows. It exercises all 3 × 3 alphabet/verifier combinations, runs CLI and reproducibility smoke tests, builds distributions, and runs formatting, lint, and type checks. CI results, not this README, are the source of truth for currently passing environments.

References

  1. Kyte, J. & Doolittle, R. F. (1982). A simple method for displaying the hydropathic character of a protein. Journal of Molecular Biology 157(1), 105–132. doi:10.1016/0022-2836(82)90515-0; PubMed.
  2. LibreTexts, What is a protein? — peptide bond formation, residue masses, and water loss. Section 1A.
  3. Grimsley, G. R., Scholtz, J. M. & Pace, C. N. (2009). A summary of the measured pK values of the ionizable groups in folded proteins. Protein Science 18, 247–251. doi:10.1002/pro.19; PMC.

Release maintainers

The release workflow publishes through PyPI Trusted Publishing (OIDC). Before the first release, configure the PyPI project cerberos-peptide-splitter with a trusted publisher for the GitHub repository AI-Meat-Lab/cerberos and workflow .github/workflows/release.yml. Releases use v* tags and fail unless the tag exactly matches project.version in pyproject.toml.

MIT licensed; see LICENSE. Report issues or contribute through the GitHub repository, issue tracker, and pull requests.

Metadata

Release files for cerberos-peptide-splitter 0.0.1b1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cerberos-peptide-splitter 0.0.1b1
File Size Uploaded
cerberos_peptide_splitter-0.0.1b1.tar.gz 48.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cerberos-peptide-splitter 0.0.1b1
File Interpreter ABI Platform
cerberos_peptide_splitter-0.0.1b1-py3-none-any.whl Python 3 none any Details

Total release size: 82.7 kB

Release files / cerberos_peptide_splitter-0.0.1b1.tar.gz

Download URL cerberos_peptide_splitter-0.0.1b1.tar.gz
Size 48.1 kB
Tags Source
SHA-256 checksum
How to use checksums
23b2b443447a59620075d071bf340b51db9866c6d8926d740274068cb5cab046
BLAKE2b-256 checksum
How to use checksums
851b709be08ef7686d344c0d08c9a19086b8ee63f9b47985179c530efd6832ae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release files / cerberos_peptide_splitter-0.0.1b1-py3-none-any.whl

Download URL cerberos_peptide_splitter-0.0.1b1-py3-none-any.whl
Size 34.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e3cd07fe795d87e410573282fd6afd6b78216cfe4eb90453eed21e916ece6e35
BLAKE2b-256 checksum
How to use checksums
89fd81babe0d3888aa5d8f614304e5a2e58d181f068a0d6859d7ea9e8dd1e85c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.0

Release history Release notifications | RSS feed

This release

0.0.1b1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page