VarTrace
VarTrace estimates, for each variant observed in single-cell RNA sequencing, the probability that the variant is not present in the genome of that cell. It works from the RNA data alone, without paired DNA sequencing.
Why
A variant seen in RNA may be present in the genome, or it may arise only at the RNA level. Telling the two apart normally requires sequencing the DNA of the same cell, which is expensive and often not possible. The common shortcut, checking whether the variant is listed in public databases, misclassifies a large share of cases in both directions.
VarTrace answers the question from the RNA data and returns a calibrated probability rather than a yes or no. A score of 0.87 means that of a hundred variants scored that way, about 87 really are absent from the genome.
Installing
pipx install vartrace
Or, for the latest state of the source:
pipx install "git+https://codeberg.org/KamenYovchev/vartrace.git"
The trained models come with the package and the web interface is included, so nothing else has to be downloaded or installed separately.
Using it
Run it with nothing after it and the web interface opens in the browser. It listens only on your own machine, so the files you load stay on it.
vartrace
Give it files instead and it works from the command line, one file per cell or one file holding many cells:
vartrace cell_01.vcf
vartrace cells/*.tsv --threshold 0.867 --output results.tsv
Options:
--threshold 0.5 by default. 0.867 is the high confidence mode, where
precision is about 0.99 at the cost of some recall.
--no-annotate do not ask Ensembl for dbSNP and gnomAD when the input
carries no annotation of its own.
--output where to write the table. A second file with the same name
and .info.txt records what produced it.
Input
A plain VCF, as it comes from the variant caller, or a table annotated with ANNOVAR, where the VCF fields sit in the Otherinfo columns.
A VCF may hold one sample or many. One column per cell barcode is read as many cells, and a cell is counted for a variant only when its genotype or its allele depth says it carries it.
Only biallelic single base substitutions are used, because that is what the models were trained on. Anything else is counted and skipped, and the count is reported.
If the input carries no database annotation, VarTrace looks up dbSNP and gnomAD through the public Ensembl service. That needs an internet connection and can be switched off. Of the five database features these two carry nearly all of the weight, which is why the other two are used when present but never fetched.
Output
A table with one row per variant: the cell, the position, the alleles, the score, the call at the chosen threshold, and which model was used.
Alongside it a short record of what produced those numbers, including the model file and its fingerprint, and where the annotation came from.
The tool also states in words why it chose the model it did, for example that the input had dbSNP and gnomAD but not REDIportal or COSMIC.
The models
Five models come with the package and one is chosen automatically from what the input supports. No database annotation at all, or all four databases, is served by a model trained for exactly that case, with or without the cross-cell features depending on whether more than one cell was given. Anything in between goes to a model trained with feature groups withheld at random, which is the only one that can take a partial input.
Each model carries a description in JSON next to it: the algorithm, its settings, the features it expects in order, what it was trained on, and its validation metrics.
Scope and limitations
The models were trained on 45 single cells from three cell lines, sequenced with one technology, with paired DNA sequencing as the ground truth. They have not been validated on primary tissue or on other protocols.
The score answers one question: is the variant present in the genome. It does not say which biological process produced a variant that is absent from it.
The models were trained on gnomAD as ANNOVAR supplies it. When VarTrace fetches the frequencies from Ensembl instead, the release may differ; which one answered is recorded in the output.
Licence
Code: PolyForm Noncommercial License 1.0.0, see LICENSE. Trained models: Creative Commons Attribution-NonCommercial 4.0, see LICENSE-MODELS.
Free for research and any other non-commercial use. Commercial use is not permitted.
Citation
To follow.
Metadata
Release files for vartrace 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vartrace-0.1.0.tar.gz | 1.7 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vartrace-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 3.4 MB
Release files / vartrace-0.1.0.tar.gz
| Download URL | vartrace-0.1.0.tar.gz |
|---|---|
| Size | 1.7 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f95cfe4efd89e6f5931d29d1be5b786bedaa7c5f5e1a5849d78a96acfbda9a08
|
|
BLAKE2b-256 checksum How to use checksums |
bafa60aaca920ae8b5333e0f5a731bbba02312a022060a9e1e4014b27f2a04bf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / vartrace-0.1.0-py3-none-any.whl
| Download URL | vartrace-0.1.0-py3-none-any.whl |
|---|---|
| Size | 1.7 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
d53ec4192b7d935e2fdebde738d616518cf39d80ff763fd60b471c8b1b427296
|
|
BLAKE2b-256 checksum How to use checksums |
647ae2af033778ee9057fe500094ca061ddf77d5cca5dbea5488559042b46b36
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|