This release is a pre-release and may not be stable for production use.
cerberos
A local, Python-based tool for approximate homology-aware splitting and descriptive diagnostics of peptide/protein FASTA datasets.
Cerberos reads a FASTA file, generates approximate sequence-similarity clusters, and assigns complete clusters to train, validation, and test splits. It can also describe an existing set of three splits. It is designed to help identify obvious near-duplicate leakage; it is not a substitute for a validated alignment-based homology search or a guarantee of leakage-free evaluation.
The package requires NumPy. Plotting is an optional Matplotlib extra; KS and Wasserstein descriptive distances are computed with NumPy and do not require SciPy. It runs locally and does not call external services or external sequence-search binaries.
Contents
- Installation
- Quick start
- Input and output
- CLI reference
- Method and interpretation
- Scientific definitions and limitations
- Python API
- Development and tests
- References
Installation
Python 3.10 or later is supported. The PyPI distribution name is cerberos-peptide-splitter, the Python import name is cerberos_peptide_splitter, and the installed command is cerberos-peptide-splitter. This project has not yet been published to PyPI; install from the source checkout below until the first release.
# After the first release to PyPI:
python -m pip install cerberos-peptide-splitter
Optional plotting support (after the first PyPI release):
python -m pip install 'cerberos-peptide-splitter[plots]'
Optional plotting and the full developer toolchain (after the first PyPI release):
python -m pip install 'cerberos-peptide-splitter[full]'
python -m pip install 'cerberos-peptide-splitter[dev]'
To work from a source checkout:
git clone https://github.com/AI-Meat-Lab/cerberos.git
cd cerberos
python -m venv .venv
# Linux/macOS: source .venv/bin/activate
# Windows: .venv\Scripts\activate
python -m pip install -e '.[dev]'
Quick start
Split one FASTA file using defaults:
cerberos-peptide-splitter split --input peptides.fasta --output-dir splits/
For a sensitivity-oriented reduced-alphabet pass followed by exact raw-sequence edit-distance verification:
cerberos-peptide-splitter split \
--input peptides.fasta \
--reduced-alphabet groups5 \
--verify-with levenshtein \
--similarity-threshold 0.80 \
--prefilter-threshold 0.30 \
--kmer-sizes 2,3 \
--output-dir splits/
Audit an existing train/validation/test directory:
cerberos-peptide-splitter audit --dir splits/ --output-dir audit/
Use python -m cerberos_peptide_splitter instead of cerberos-peptide-splitter if running from an unpacked source tree without installing the console entry point.
Input and output
FASTA input
Records may be unlabeled or have a binary label as the final header field:
>peptide_1|1
KWKLFKKIEKVGQNIRDGIIK
>peptide_2|0
MKTIIALSYIFCLVFA
Rules:
- A final
|0or|1is parsed as a label; other header text is retained as the identifier. - Sequence lines are uppercased and all whitespace within them is removed.
- Empty records are ignored. Sequence text before the first
>header and empty headers are errors. - Other symbols (for example
X,-, or*) are retained; they are not validated against a particular alphabet. With a reduced alphabet enabled, every non-canonical symbol is mapped to the sameXgroup. - Mixed labeled and unlabeled files are accepted. Label stratification is used only when every input record has a label.
Outputs
split writes train.fasta, val.fasta, test.fasta, report.txt, stats.json, and—when Matplotlib is installed—a plots/ directory. audit reads FASTAs and writes only report/statistics/plots; it does not overwrite the input FASTAs.
The audit command recognizes train.fasta/.fa/.faa, val.fasta/.fa/.faa or valid.fasta/.fa, and test.fasta/.fa/.faa.
stats.json contains effective configuration, split sizes and label counts, length/feature summaries, sampled pairwise k-mer Jaccard summaries, empirical KS D statistics, Wasserstein distances, and heuristic warnings. Empty samples are represented with JSON null, never non-standard NaN literals.
CLI reference
cerberos-peptide-splitter split
| Option | Default | Meaning |
|---|---|---|
--input PATH |
required | Input FASTA. |
--output-dir DIR |
cerberos_out |
Output directory. |
--seed N |
42 |
Integer seed in NumPy's supported 32-bit range. |
--train-pct F |
0.80 |
Train proportion. The three non-negative proportions are normalized to sum to one. |
--val-pct F |
0.10 |
Validation proportion. |
--test-pct F |
0.10 |
Test proportion. |
--no-clustering |
off | Use random splitting instead of grouping candidate homologs. |
--similarity-threshold F |
0.70 |
MinHash estimate threshold if no verifier is selected; final score threshold otherwise. Must be in [0, 1]. |
--kmer-sizes CSV |
3 |
Positive comma-separated k-mer sizes. |
--num-hashes N |
128 |
MinHash signature size (at least 4). |
--reduced-alphabet NAME |
none |
none, groups5, or groups7; transform residues before candidate generation. |
--verify-with NAME |
none |
none, kmer-exact, or levenshtein. |
--prefilter-threshold F |
0.30 |
MinHash estimate threshold before deterministic verification; cannot exceed final threshold. |
cerberos-peptide-splitter audit
| Option | Default | Meaning |
|---|---|---|
--dir DIR |
required | Directory containing the three split FASTAs. |
--output-dir DIR |
cerberos_out |
Diagnostics output directory. |
--seed N |
42 |
Seed for reproducible diagnostic sampling. |
--kmer-sizes CSV |
3 |
Positive k-mer sizes; diagnostics use the first value. |
Invalid CLI values and missing input files are reported as concise command-line errors. The process exits non-zero on failure.
Method and interpretation
- Optional alphabet transform.
groups5maps aliphatic and aromatic residues together, polar residues together, basic residues together, acidic residues together, and G/P/C together.groups7keeps aliphatic, aromatic, polar, basic, acidic, G/P, and C groups separate. These are heuristic groupings, not a substitution matrix or evolutionary model. - K-mer sketches. For each configured k, unique contiguous k-mers are hashed into seeded MinHash signatures. CRC32 provides stable input hashes. The observed sketch agreement estimates set Jaccard similarity.
- LSH candidate generation. Signatures are partitioned into bands; pairs sharing a band are considered. Candidate retrieval is approximate and depends on signature length, band size, sequence composition, k, and threshold. A low prefilter threshold does not make candidate recall exhaustive.
- Optional deterministic scoring.
kmer-exactcalculates exact unique-k-mer Jaccard on transformed sequences and takes the maximum across configured k values.levenshteincalculates global unit-cost edit similarity on raw sequences,1 - edit_distance / max(sequence_length). It is not alignment-based percent identity. Verification only scores pairs already proposed by LSH. - Connected components. Accepted pair edges are merged with Union-Find. A connected component may include a pair of endpoints whose direct score is below the edge threshold (transitivity).
- Whole-cluster split assignment. Clusters are greedily assigned to minimize deviation from target split sizes; preserving clusters takes priority over exact proportions. Clusters are not stratified by label.
- Descriptive diagnostics. Raw sequences are used for pairwise k-mer distances and biochemical features. Pairwise similarity is sampled per split for cost control. Feature distances are descriptive, not inferential tests of model generalization.
Scientific definitions and limitations
- Jaccard: unique-k-mer set intersection divided by union. Repeated occurrences do not receive extra weight. A sequence shorter than k has no k-mers; Cerberos treats that as no similarity evidence rather than declaring two empty sets a match.
- Reduced alphabets: grouping changes the sequence representation and can increase similarity among chemically grouped residues. It can also merge unrelated sequences. It is a heuristic sensitivity/specificity tradeoff; evaluate settings for the target domain.
- Levenshtein score: global edit distance gives substitutions, insertions, and deletions unit cost. The normalized score is length-dependent, does not use BLOSUM/PAM scores, does not model local alignment, and should not be reported as biological percent identity.
- MinHash and LSH: MinHash agreement approximates Jaccard under assumptions about the hash family; LSH is a probabilistic candidate filter. A missed candidate cannot be recovered by verification. For high-assurance leakage control, use an alignment-based clustering/search method and inspect thresholds on a representative validation set.
- Clusters and thresholds: threshold edges are transitively closed; therefore maximum pairwise within-cluster similarity is not guaranteed to meet the threshold. The cluster is an operational grouping for split assignment, not a statement that all members are homologous.
- Hydropathy: mean Kyte-Doolittle score over canonical residues, using the original residue scale. This is a sequence summary, not a membrane topology prediction.
- Charge: approximate peptide net charge at pH 7.0, including free termini and standard approximate side-chain pKa values via Henderson-Hasselbalch fractions. Local environment, terminal modifications, pKa shifts, and non-canonical residues are not modeled.
- Molecular mass: average free-amino-acid masses are summed and 18.01528 Da is subtracted per peptide bond. Unknown symbols are assigned an approximate 110 Da free-residue mass. This is not monoisotopic mass and does not account for modifications, cyclization, disulfide formation, or unusual termini.
- Unknown symbols: canonical frequencies and group fractions are per total sequence length;
unknown_fractionreports positions outside the 20 canonical amino acids. Hydropathy averages canonical residues only. - KS and Wasserstein: the report stores empirical two-sample KS D and 1-D first Wasserstein distance per feature. No p-values or multiple-testing correction are computed. The
>0.20warning is a heuristic and depends on sample size, feature scaling, and data context. - No leakage guarantee: no warning is not proof of independence, no distribution test establishes OOD safety, and this software has not been validated as a clinical or regulatory assay.
Python API
from cerberos_peptide_splitter import RunConfig, parse_fasta, split_records
records = parse_fasta("peptides.fasta")
config = RunConfig(
train_pct=0.8,
val_pct=0.1,
test_pct=0.1,
kmer_sizes=[2, 3],
reduced_alphabet="groups5",
verify_with="levenshtein",
similarity_threshold=0.8,
prefilter_threshold=0.3,
seed=42,
)
splits = split_records(records, config, verbose=False)
print({name: len(items) for name, items in splits.items()})
Top-level exports also include write_fasta, assign_clusters, stratified_random_split, summarize, build_report, run_split, and run_audit. Lower-level modules expose the reduced-alphabet, k-mer, MinHash, and edit-distance helpers.
Development and tests
python -m pip install -e '.[dev]'
pytest -q
pytest -q -m 'not optional' # NumPy-only CI subset
ruff check cerberos_peptide_splitter tests
ruff format --check cerberos_peptide_splitter tests
black --check cerberos_peptide_splitter tests
isort --check-only cerberos_peptide_splitter tests
mypy cerberos_peptide_splitter --ignore-missing-imports --no-strict-optional
python -m build
python -m twine check dist/*
GitHub Actions checks a NumPy-only test path and a full-dependency test path across Ubuntu, macOS, and Windows, exercises all 3 × 3 alphabet/verifier combinations, runs CLI and reproducibility smoke tests, builds distributions, and runs formatting/lint/type checks. CI results—not this README—are the source of truth for currently passing environments.
References
- Kyte, J. & Doolittle, R. F. (1982). A simple method for displaying the hydropathic character of a protein. Journal of Molecular Biology 157(1), 105–132. doi:10.1016/0022-2836(82)90515-0; PubMed.
- LibreTexts, What is a protein?—peptide bond formation, residue masses, and water loss. Section 1A.
- Grimsley, G. R., Scholtz, J. M. & Pace, C. N. (2009). A summary of the measured pK values of the ionizable groups in folded proteins. Protein Science 18, 247–251. doi:10.1002/pro.19; PMC.
Release maintainers
The release workflow publishes through PyPI Trusted Publishing (OIDC). Before the first release, configure the PyPI project cerberos-peptide-splitter with a trusted publisher for the GitHub repository AI-Meat-Lab/cerberos and workflow .github/workflows/release.yml. Releases use v* tags and fail unless the tag exactly matches project.version in pyproject.toml.
License and project links
MIT licensed; see LICENSE. Report issues or contribute through the GitHub repository, issue tracker, and pull requests.
Metadata
Release files for cerberos-peptide-splitter 0.0.1b0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| cerberos_peptide_splitter-0.0.1b0.tar.gz | 43.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| cerberos_peptide_splitter-0.0.1b0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 75.3 kB
Release files / cerberos_peptide_splitter-0.0.1b0.tar.gz
| Download URL | cerberos_peptide_splitter-0.0.1b0.tar.gz |
|---|---|
| Size | 43.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3256141c03fa64aa17a24c77328d704af5b96a81dbf5d0e4ac1e54fb335c4b4f
|
|
BLAKE2b-256 checksum How to use checksums |
20ad488562dc113fef85278bb439ff0c2e894d6b6f3df85ecec00a6dda4320a6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|
Release files / cerberos_peptide_splitter-0.0.1b0-py3-none-any.whl
| Download URL | cerberos_peptide_splitter-0.0.1b0-py3-none-any.whl |
|---|---|
| Size | 31.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
81e3ae52bf53b9e22fe15f986790e2cbb69770320d317fb5461ebf4281531550
|
|
BLAKE2b-256 checksum How to use checksums |
b4b31aa21f20011a0b9be12a8ac306f49435967433b6738c75880dc10a92446b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.0
|