Skip to main content

Satellome

Tests codecov Python Version License: MIT PyPI version

A comprehensive bioinformatics tool for analyzing satellite DNA (tandem repeats) in telomere-to-telomere (T2T) genome assemblies.

Overview

Satellome uses FasTAN (Fast Tandem Repeat Finder) as its default tandem repeat detection engine, providing fast and accurate identification of repetitive DNA sequences. The tool classifies and visualizes tandem repeats with a focus on centromeric and telomeric regions.

The tool is designed to work with various genome assembly projects including:

  • T2T (Telomere-to-Telomere) Consortium assemblies
  • DNA Zoo chromosome-length assemblies
  • VGP (Vertebrate Genome Project) assemblies
  • NCBI RefSeq and GenBank assemblies

Features

  • Fast Tandem Repeat Detection: Uses FasTAN by default for rapid, accurate detection
  • Smart Classification: Categorizes repeats into microsatellites, complex repeats, and other types
  • Rich Visualizations: Generates karyotype plots and chromosome-level visualizations
  • Annotation Integration: Supports GFF3 and RepeatMasker annotations
  • Parallel Processing: Efficient handling of large genomes
  • Smart Pipeline: Automatically skips completed steps (override with --force)
  • Compressed File Support: Direct processing of .gz compressed FASTA files
  • Optional TRF Support: Traditional TRF analysis available with --run-trf flag

Quick Start

# Install from PyPI
pip install satellome

# Install required binaries (FasTAN, tanbed, Rust helpers)
satellome --install-all

# Run on a genome
satellome -i genome.fasta -o output_dir -p project_name -t 8

Installation

From PyPI (Recommended)

pip install satellome

# Install external tools
satellome --install-all

Checking the installation

satellome --doctor

--doctor reports where this install lives, whether your shell can actually see the launcher, and where each external tool resolves from. It exits 0 when everything is healthy and 1 when it finds a problem, so it can be used as a setup gate in a driver script.

Missing tools are reported by consequence, not merely as "not installed", because that is the part that decides whether a finished run is complete:

Tool Missing means
fastan, tanbed, arraysplitter the analysis cannot run
sat-family family clustering is skipped — no families output
telomere-check the telomere check is skipped — no telomere output
find-gaps, bed-extract, genome-size a Python fallback gives the same result, slower
trf only needed with --run-trf

A tool whose absence removes results makes --doctor exit non-zero, and is announced at the start of a run — before hours of compute, rather than as a line that scrolls past in the middle of it. Tools that only cost speed are listed separately and are not treated as problems.

Install the Rust helpers (they are compiled from the rust/ sources in this repository, so a Rust toolchain is required):

satellome --install-rust-tools    # or: satellome --install-all

It is the answer to the most common post-install surprise:

WARNING: The script satellome is installed in '/home/user/.local/bin'
which is not on PATH.

That means pip used a user install and your shell cannot see the launcher — satellome will report command not found even though the package is fine. Satellome repairs this itself:

python -m satellome --fix-path

python -m satellome always works, because it does not depend on PATH at all — which is what makes it the way in when the satellome command is exactly the thing you cannot reach. --fix-path appends one marked block to the startup file your shell already reads:

# >>> satellome path >>>
export PATH="$HOME/.local/bin:$PATH"
# <<< satellome path <<<

It is idempotent (the marker means it is written once and never again), it will not duplicate an entry you added by hand, and it appends to a file that already exists rather than creating a new one — creating ~/.bash_profile where none existed would stop a login shell from reading ~/.profile.

A process cannot change its parent shell's environment, so the block applies to new shells. For the shell you are in: source ~/.bashrc.

The same repair runs automatically at the start of any run that detects the problem, so it does not have to be found first. To keep the warning but never let satellome touch a shell file, set SATELLOME_NO_PATH_FIX=1. Manual equivalents, if you prefer:

export PATH="$HOME/.local/bin:$PATH"                       # this shell only
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bashrc   # permanent
python -m satellome ...                                    # no PATH change needed

Two competing installations (a satellome on PATH that is not the one your active interpreter owns) are reported but never "fixed" — deciding which one to delete is yours to make.

python -m satellome is the most robust form in batch scripts: it always follows the active environment, whereas a console script keeps the interpreter its shebang was written with. Satellome resolves pip-installed companion tools (such as arraysplitter) through the same interpreter-aware lookup, so a directory missing from PATH degrades nothing silently — it is reported and the tool is still used.

Set SATELLOME_NO_ENV_CHECK=1 to silence the startup check entirely on machines where this layout is deliberate; --doctor still reports it.

From Source

git clone https://github.com/aglabx/satellome.git
cd satellome
pip install -e .
satellome --install-all

External Tools

Satellome requires FasTAN and tanbed for default operation. Install them automatically:

# Install all tools (FasTAN, tanbed, modified TRF)
satellome --install-all

# Or install individually
satellome --install-fastan
satellome --install-tanbed
satellome --install-trf-large  # For genomes with chromosomes >2GB

Build requirements: git, make, C compiler (gcc/clang)

# Ubuntu/Debian
sudo apt-get install build-essential git

# macOS
xcode-select --install

Usage

Basic Command

satellome -i genome.fasta -o output_dir -p project_name -t 8

Common Options

# With GFF3 annotations
satellome -i genome.fasta -o output_dir -p project_name -t 8 --gff annotations.gff3

# With RepeatMasker annotations
satellome -i genome.fasta -o output_dir -p project_name -t 8 --rm repeatmasker.out

# Force rerun all steps
satellome -i genome.fasta -o output_dir -p project_name -t 8 --force

# Also run traditional TRF analysis
satellome -i genome.fasta -o output_dir -p project_name -t 8 --run-trf

Parameters

Parameter Description Default
-i, --input Input FASTA file (.fa, .fasta, .gz) Required
-o, --output Output directory Required
-p, --project Project name Required
-t, --threads Number of threads 1
--gff GFF3 annotation file None
--rm RepeatMasker output file None
--run-trf Also run TRF analysis False
--force Force rerun all steps False
--taxid NCBI taxonomy ID None

Output Structure

output_dir/
├── genome.sat                    # Main SAT output (all arrays)
├── genome.1kb.sat                # Arrays >1kb
├── genome.3kb.sat                # Arrays >3kb
├── genome.10kb.sat               # Arrays >10kb
├── genome.micro.sat              # Microsatellites (1-9 bp monomers)
├── genome.complex.sat            # Complex repeats (>9 bp monomers)
├── genome.pmicro.sat             # Potential microsatellites
├── genome.tssr.sat               # Tandem simple sequence repeats
├── genome.gaps.bed               # Gaps annotation
├── results.yaml                  # Analysis statistics
├── run_manifest.json             # What the run produced (written last) + step statuses
├── fastan/                       # FasTAN intermediate files
│   ├── genome.1aln               # FasTAN alignment output
│   └── genome.bed                # FasTAN BED format
├── fasta/                        # FASTA sequences
│   └── genome.arrays.fasta       # All array sequences
├── gff3/                         # GFF3 annotations
│   ├── genome.1kb.gff
│   ├── genome.complex.gff
│   └── ...
├── images/                       # Visualizations
│   └── *.png
└── reports/                      # HTML reports
    └── satellome_report.html

Taxon Names in File Names

Karyotype charts are named after the taxon (--taxon, or the name resolved from --taxid). Organism names are free text and NCBI strain designations often contain slashes — Leishmania braziliensis MHOM/BR/75/M2904, Chlorella vulgaris CCAP 1055/1. Such a name is reduced to a single safe file-name component (..._MHOM_BR_75_M2904.karyo.*) and the substitution is logged; the plot titles keep the original name. Without this the slash was read as a path separator and the drawing step died with FileNotFoundError after the whole pipeline had already written its data.

Verifying a Run

Do not decide that an output directory is complete by checking that some files exist. A file that was read or copied while satellome was still writing it exists just as hard as a complete one, and a gzip of such a partial read is a valid archive that gzip -t accepts — so truncated data can enter downstream analysis unnoticed.

Every run writes run_manifest.json last, recording each file it produced with its byte size plus the status of every step. Verify against it:

satellome --verify-run output_dir
  • exit 0 — the directory matches its manifest and no step failed
  • exit 1 — not a verifiably complete run: no/corrupt manifest, a failed step, a missing file, a leftover *.partial, or a file whose size no longer matches what the run wrote (the truncated-copy case)
  • exit 2 — the argument is not a directory

Files already compressed by your own pipeline are still checked: for a missing X with an X.gz next to it, the gzip ISIZE trailer (the uncompressed length the compressor actually consumed) is compared to the recorded size, which catches a .gz made from an incomplete read. Above 4 GiB that comparison is modulo 4 GiB and the report says so.

Two other guarantees back this up:

  • Atomic outputs — files are written as <path>.partial and renamed into place, so a final name never refers to a half-written file. If you compress or copy an output directory concurrently, you either get the complete file or no file, never a truncated one.
  • Output-directory lock — a second satellome run into the same -o is refused, naming the pid and host that holds it, instead of overwriting the first run's files mid-write. Override with --ignore-lock only if you are sure.

A run whose drawing step fails still writes a manifest — with drawing: failed — and exits non-zero. The data files are complete and usable; the failed step is recorded in the run's own artifact rather than only in an exit code.

SAT File Format

The SAT format is a tab-delimited file with the following columns:

Column Description
project Project name
trf_id Unique array ID
trf_head Chromosome/scaffold name
trf_l_ind Left coordinate (1-based)
trf_r_ind Right coordinate
trf_period Monomer period length
trf_n_copy Number of copies
trf_pmatch Percent match
trf_pvar Percent variation
trf_entropy Shannon entropy
trf_consensus Consensus monomer sequence
trf_array Full array sequence
trf_array_gc Array GC content
trf_consensus_gc Consensus GC content
trf_array_length Array length in bp
trf_joined Join status
trf_family Repeat family
trf_ref_annotation Reference annotation

Classification System

Satellome classifies tandem repeats into four categories:

Category Description Criteria
micro Microsatellites Monomer 1-9 bp
complex Complex repeats Monomer >9 bp, entropy >1.82
pmicro Potential microsatellites Intermediate characteristics
tssr Tandem simple sequence repeats Simple patterns

Example Results

Analysis of CHM13 v2.0 human genome (3.1 GB):

Category Arrays % Genome
Total 614,616 -
Complex 20,373 5.27%
Microsatellites 319,489 1.96%
TSSR 296,475 0.47%
>1kb 14,438 7.69%
>10kb 1,223 6.67%

Utility Scripts

Format Conversion

python scripts/trf_to_fasta.py -i repeats.sat -o repeats.fasta
python scripts/trf_to_gff3.py -i repeats.sat -o repeats.gff3

Analysis Tools

python scripts/trf_get_large.py -i repeats.sat -m 1000 -o large_repeats.sat
python scripts/trf_get_micro_stat.py -i repeats.sat -o micro_stats.txt
python scripts/check_telomeres.py -i genome.fasta -t repeats.sat

Testing

pytest tests/unit/ -v

Contributing

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'Add amazing feature')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

Citation

If you use Satellome in your research, please cite:

Komissarov A. et al. (2026). Satellome: A comprehensive tool for satellite DNA
analysis in T2T genome assemblies. [Publication details]

License

This project is licensed under the MIT License - see the LICENSE file for details.

Support

Acknowledgments

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

satellome-1.11.0.tar.gz (288.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

satellome-1.11.0-py3-none-any.whl (216.2 kB view details)

Uploaded Python 3

File details

Details for the file satellome-1.11.0.tar.gz.

File metadata

  • Download URL: satellome-1.11.0.tar.gz
  • Upload date:
  • Size: 288.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for satellome-1.11.0.tar.gz
Algorithm Hash digest
SHA256 3d880cb453a0d993f26942a2b96ae78c79991bde3e3f9792e48cb9a62043b9d5
MD5 0c736591a9007fc09fb7b01d6bac19b1
BLAKE2b-256 32f36ee6f7bb5674fafe302a2722940d120ffa9e19863a4a7a5acaecc64badf4

See more details on using hashes here.

Provenance

The following attestation bundles were made for satellome-1.11.0.tar.gz:

Publisher: publish.yml on aglabx/satellome

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file satellome-1.11.0-py3-none-any.whl.

File metadata

  • Download URL: satellome-1.11.0-py3-none-any.whl
  • Upload date:
  • Size: 216.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for satellome-1.11.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e0e7084b09878f0628b09c1b52564c7552bdc1a7c226fcd76d11ac76ee1fd284
MD5 c0e96ec7312754ce081c8a2e2f3c490a
BLAKE2b-256 ec909a54be989b9fac8bdc7f88308e0d6854881e63af1c2909b30f00030bee67

See more details on using hashes here.

Provenance

The following attestation bundles were made for satellome-1.11.0-py3-none-any.whl:

Publisher: publish.yml on aglabx/satellome

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.17.2

2 files

1.17.1

2 files

1.17.0

2 files

1.16.1

2 files

1.16.0

2 files

1.15.1

2 files

1.15.0

2 files

1.14.0

2 files

1.13.0

2 files

1.12.1

2 files

1.12.0

2 files

This release

1.11.0 This release

2 files

1.10.0

2 files

1.9.0

2 files

1.8.0

2 files

1.7.0

2 files

1.6.1

2 files

1.6.0

2 files

1.5.2

2 files

1.5.1

2 files

1.5.0

2 files

1.4.3

2 files

1.4.1

2 files

1.4.0

2 files

1.1.0

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page