Satellome
A comprehensive bioinformatics tool for analyzing satellite DNA (tandem repeats) in telomere-to-telomere (T2T) genome assemblies.
Overview
Satellome uses FasTAN (Fast Tandem Repeat Finder) as its default tandem repeat detection engine, providing fast and accurate identification of repetitive DNA sequences. The tool classifies and visualizes tandem repeats with a focus on centromeric and telomeric regions.
The tool is designed to work with various genome assembly projects including:
- T2T (Telomere-to-Telomere) Consortium assemblies
- DNA Zoo chromosome-length assemblies
- VGP (Vertebrate Genome Project) assemblies
- NCBI RefSeq and GenBank assemblies
Features
- Fast Tandem Repeat Detection: Uses FasTAN by default for rapid, accurate detection
- Smart Classification: Categorizes repeats into microsatellites, complex repeats, and other types
- Rich Visualizations: Generates karyotype plots and chromosome-level visualizations
- Annotation Integration: Supports GFF3 and RepeatMasker annotations
- Parallel Processing: Efficient handling of large genomes
- Smart Pipeline: Automatically skips completed steps (override with
--force) - Compressed File Support: Direct processing of .gz compressed FASTA files
- Optional TRF Support: Traditional TRF analysis available with
--run-trfflag
Quick Start
# Install from PyPI
pip install satellome
# Install required binaries (FasTAN, tanbed)
satellome --install-all
# Run on a genome
satellome -i genome.fasta -o output_dir -p project_name -t 8
Installation
From PyPI (Recommended)
pip install satellome
# Install external tools
satellome --install-all
From Source
git clone https://github.com/aglabx/satellome.git
cd satellome
pip install -e .
satellome --install-all
External Tools
Satellome requires FasTAN and tanbed for default operation. Install them automatically:
# Install all tools (FasTAN, tanbed, modified TRF)
satellome --install-all
# Or install individually
satellome --install-fastan
satellome --install-tanbed
satellome --install-trf-large # For genomes with chromosomes >2GB
Build requirements: git, make, C compiler (gcc/clang)
# Ubuntu/Debian
sudo apt-get install build-essential git
# macOS
xcode-select --install
Usage
Basic Command
satellome -i genome.fasta -o output_dir -p project_name -t 8
Common Options
# With GFF3 annotations
satellome -i genome.fasta -o output_dir -p project_name -t 8 --gff annotations.gff3
# With RepeatMasker annotations
satellome -i genome.fasta -o output_dir -p project_name -t 8 --rm repeatmasker.out
# Force rerun all steps
satellome -i genome.fasta -o output_dir -p project_name -t 8 --force
# Also run traditional TRF analysis
satellome -i genome.fasta -o output_dir -p project_name -t 8 --run-trf
Parameters
| Parameter | Description | Default |
|---|---|---|
-i, --input |
Input FASTA file (.fa, .fasta, .gz) | Required |
-o, --output |
Output directory | Required |
-p, --project |
Project name | Required |
-t, --threads |
Number of threads | 1 |
--gff |
GFF3 annotation file | None |
--rm |
RepeatMasker output file | None |
--run-trf |
Also run TRF analysis | False |
--force |
Force rerun all steps | False |
--taxid |
NCBI taxonomy ID | None |
Output Structure
output_dir/
├── genome.sat # Main SAT output (all arrays)
├── genome.1kb.sat # Arrays >1kb
├── genome.3kb.sat # Arrays >3kb
├── genome.10kb.sat # Arrays >10kb
├── genome.micro.sat # Microsatellites (1-9 bp monomers)
├── genome.complex.sat # Complex repeats (>9 bp monomers)
├── genome.pmicro.sat # Potential microsatellites
├── genome.tssr.sat # Tandem simple sequence repeats
├── genome.gaps.bed # Gaps annotation
├── results.yaml # Analysis statistics
├── run_manifest.json # What the run produced (written last) + step statuses
├── fastan/ # FasTAN intermediate files
│ ├── genome.1aln # FasTAN alignment output
│ └── genome.bed # FasTAN BED format
├── fasta/ # FASTA sequences
│ └── genome.arrays.fasta # All array sequences
├── gff3/ # GFF3 annotations
│ ├── genome.1kb.gff
│ ├── genome.complex.gff
│ └── ...
├── images/ # Visualizations
│ └── *.png
└── reports/ # HTML reports
└── satellome_report.html
Taxon Names in File Names
Karyotype charts are named after the taxon (--taxon, or the name resolved from
--taxid). Organism names are free text and NCBI strain designations often
contain slashes — Leishmania braziliensis MHOM/BR/75/M2904, Chlorella vulgaris CCAP 1055/1. Such a name is reduced to a single safe file-name
component (..._MHOM_BR_75_M2904.karyo.*) and the substitution is logged; the
plot titles keep the original name. Without this the slash was read as a path
separator and the drawing step died with FileNotFoundError after the whole
pipeline had already written its data.
Verifying a Run
Do not decide that an output directory is complete by checking that some files
exist. A file that was read or copied while satellome was still writing it exists
just as hard as a complete one, and a gzip of such a partial read is a valid
archive that gzip -t accepts — so truncated data can enter downstream analysis
unnoticed.
Every run writes run_manifest.json last, recording each file it produced
with its byte size plus the status of every step. Verify against it:
satellome --verify-run output_dir
- exit
0— the directory matches its manifest and no step failed - exit
1— not a verifiably complete run: no/corrupt manifest, a failed step, a missing file, a leftover*.partial, or a file whose size no longer matches what the run wrote (the truncated-copy case) - exit
2— the argument is not a directory
Files already compressed by your own pipeline are still checked: for a missing
X with an X.gz next to it, the gzip ISIZE trailer (the uncompressed length
the compressor actually consumed) is compared to the recorded size, which catches
a .gz made from an incomplete read. Above 4 GiB that comparison is modulo
4 GiB and the report says so.
Two other guarantees back this up:
- Atomic outputs — files are written as
<path>.partialand renamed into place, so a final name never refers to a half-written file. If you compress or copy an output directory concurrently, you either get the complete file or no file, never a truncated one. - Output-directory lock — a second satellome run into the same
-ois refused, naming the pid and host that holds it, instead of overwriting the first run's files mid-write. Override with--ignore-lockonly if you are sure.
A run whose drawing step fails still writes a manifest — with drawing: failed —
and exits non-zero. The data files are complete and usable; the failed step is
recorded in the run's own artifact rather than only in an exit code.
SAT File Format
The SAT format is a tab-delimited file with the following columns:
| Column | Description |
|---|---|
| project | Project name |
| trf_id | Unique array ID |
| trf_head | Chromosome/scaffold name |
| trf_l_ind | Left coordinate (1-based) |
| trf_r_ind | Right coordinate |
| trf_period | Monomer period length |
| trf_n_copy | Number of copies |
| trf_pmatch | Percent match |
| trf_pvar | Percent variation |
| trf_entropy | Shannon entropy |
| trf_consensus | Consensus monomer sequence |
| trf_array | Full array sequence |
| trf_array_gc | Array GC content |
| trf_consensus_gc | Consensus GC content |
| trf_array_length | Array length in bp |
| trf_joined | Join status |
| trf_family | Repeat family |
| trf_ref_annotation | Reference annotation |
Classification System
Satellome classifies tandem repeats into four categories:
| Category | Description | Criteria |
|---|---|---|
| micro | Microsatellites | Monomer 1-9 bp |
| complex | Complex repeats | Monomer >9 bp, entropy >1.82 |
| pmicro | Potential microsatellites | Intermediate characteristics |
| tssr | Tandem simple sequence repeats | Simple patterns |
Example Results
Analysis of CHM13 v2.0 human genome (3.1 GB):
| Category | Arrays | % Genome |
|---|---|---|
| Total | 614,616 | - |
| Complex | 20,373 | 5.27% |
| Microsatellites | 319,489 | 1.96% |
| TSSR | 296,475 | 0.47% |
| >1kb | 14,438 | 7.69% |
| >10kb | 1,223 | 6.67% |
Utility Scripts
Format Conversion
python scripts/trf_to_fasta.py -i repeats.sat -o repeats.fasta
python scripts/trf_to_gff3.py -i repeats.sat -o repeats.gff3
Analysis Tools
python scripts/trf_get_large.py -i repeats.sat -m 1000 -o large_repeats.sat
python scripts/trf_get_micro_stat.py -i repeats.sat -o micro_stats.txt
python scripts/check_telomeres.py -i genome.fasta -t repeats.sat
Testing
pytest tests/unit/ -v
Contributing
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
Citation
If you use Satellome in your research, please cite:
Komissarov A. et al. (2026). Satellome: A comprehensive tool for satellite DNA
analysis in T2T genome assemblies. [Publication details]
License
This project is licensed under the MIT License - see the LICENSE file for details.
Support
- Issues: GitHub Issues
- Documentation: Wiki
- Email: ad3002@gmail.com
Acknowledgments
- FasTAN by Gene Myers
- Tandem Repeat Finder by Gary Benson
- T2T Consortium
- DNA Zoo
- Vertebrate Genome Project
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file satellome-1.8.0.tar.gz.
File metadata
- Download URL: satellome-1.8.0.tar.gz
- Upload date:
- Size: 263.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
197867ccc49075c3ec1ea3a0a014a8862581bac13ee19c192759f46a080e02b3
|
|
| MD5 |
f464b2cc45e93a178c1d9d2313b548a6
|
|
| BLAKE2b-256 |
00ce7e29066bbae61f545d0a27c14688909cbb2e9e12ea3766c65e822cba7e38
|
Provenance
The following attestation bundles were made for satellome-1.8.0.tar.gz:
Publisher:
publish.yml on aglabx/satellome
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
satellome-1.8.0.tar.gz -
Subject digest:
197867ccc49075c3ec1ea3a0a014a8862581bac13ee19c192759f46a080e02b3 - Sigstore transparency entry: 2498467234
- Sigstore integration time:
-
Permalink:
aglabx/satellome@b30e60f153d796a89bd8109028bee00ffc3acbe7 -
Branch / Tag:
refs/tags/v1.8.0 - Owner: https://github.com/aglabx
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b30e60f153d796a89bd8109028bee00ffc3acbe7 -
Trigger Event:
push
-
Statement type:
File details
Details for the file satellome-1.8.0-py3-none-any.whl.
File metadata
- Download URL: satellome-1.8.0-py3-none-any.whl
- Upload date:
- Size: 200.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7621f6777ef92188e9e701f7fb246b40c7bfaef27a4082821d6b79766b5a4808
|
|
| MD5 |
e1ea4dbc76090e9d4a2315707be12f94
|
|
| BLAKE2b-256 |
393668e9f63989eeb2cc8ca8e6180d60d6db5f3e60c33df18ada6518f97073b9
|
Provenance
The following attestation bundles were made for satellome-1.8.0-py3-none-any.whl:
Publisher:
publish.yml on aglabx/satellome
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
satellome-1.8.0-py3-none-any.whl -
Subject digest:
7621f6777ef92188e9e701f7fb246b40c7bfaef27a4082821d6b79766b5a4808 - Sigstore transparency entry: 2498467249
- Sigstore integration time:
-
Permalink:
aglabx/satellome@b30e60f153d796a89bd8109028bee00ffc3acbe7 -
Branch / Tag:
refs/tags/v1.8.0 - Owner: https://github.com/aglabx
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@b30e60f153d796a89bd8109028bee00ffc3acbe7 -
Trigger Event:
push
-
Statement type: