osteosarc
Python library and command-line tool for the public osteosarc.com dataset: one patient's osteosarcoma sequencing, variant calls, cancer vaccines and clinical history. Find files, pick variants and vaccine peptides, and fetch the reads around a variant without downloading a whole BAM.
Documentation · Key concepts · Command line · Python API · Changelog
Features
- Browse without downloading. Search nearly 400,000 files by sample, timepoint and assay: bulk and single-cell RNA, exome, genome, Oxford Nanopore and PacBio.
- Variants and vaccine peptides. The site's variants with checked alleles, read counts, which pipelines found them, vaccine peptides and ELISPOT results.
- Reads around a variant. Copy just the reads you need out of a remote BAM into a small local one.
- Corrected by default. 35 fixes to known problems in the published data, each with its evidence, such as the MAP2 vaccine target's allele. Every load checks them against the snapshot's sources, and you can turn them off.
- Clinical timeline. Treatments, procedures, imaging, MRD and lab results as a text chart, or in an interactive terminal explorer.
- Reproducible. The site's metadata is saved as dated snapshots that reopen offline. The website changes; your results don't, until you sync again.
- OpenVax integration. Works with Varcode, Isovar, Topiary and Vaxrank, builds reproducible test BAMs, and includes a list of 637 candidate structural variants.
Install
python -m pip install osteosarc
Needs Python 3.9+ on Linux or macOS, and SAMtools on your PATH to fetch reads.
Quickstart
from osteosarc import Dataset
data = Dataset.sync() # Save today's website metadata (about 57 MB)
print(data.describe_samples())
rna = data.assets_for_sample("T0_tumor", kind="alignment", assay="rna-seq")
targets = data.variants(gene="DYNC1H1", status="ready")
source = rna["rna-seq/reprocessed/BG003082/BG003082.Aligned.sortedByCoord.out.md.bam"]
reads = data.extract_reads(source, variants=targets, padding=100)
print(reads.path) # Local indexed BAM
The BAM key is the file's path in the dataset's public S3 bucket. Only the reads near the variants are downloaded, into a small indexed BAM in your local cache.
Later, Dataset.open() reopens your most recent snapshot without a network
connection. The website changes over time, so snapshots are saved by download date:
osteosarc snapshots lists them, and Dataset.open(date="2026-09") or
--snapshot 2026-09 picks the newest from that month.
Get started explains each step.
The same workflow from the terminal:
osteosarc sync
osteosarc samples
osteosarc assets --sample T0_tumor --kind alignment --assay rna-seq
osteosarc variants --gene DYNC1H1 --status ready
osteosarc reads rna-seq/reprocessed/BG003082/BG003082.Aligned.sortedByCoord.out.md.bam --variant DYNC1H1-chr14-101980529 --variant DYNC1H1-chr14-102030200 --padding 100
osteosarc explore
osteosarc explore opens an interactive browser for specimens, files, variants and
the timeline. Type help for commands and quit to leave.
Guides
| I want to… | Read |
|---|---|
| Understand sample IDs, variant status, coordinates and corrections | Key concepts |
| Find RNA, DNA, single-cell or long-read files and read tables | Find samples and files |
| Get alleles, read counts or vaccine peptides | Select variants |
| Fetch, filter or pair reads by variant or region | Extract reads |
| Browse treatments, specimens, MRD and labs | Browse the timeline |
| Pass data to Varcode, Isovar, Topiary or Vaxrank | Use other libraries |
| Build small, verifiable test BAMs | Read fixtures |
| Explore candidate structural variants | SV interest catalogue |
Data, license and citation
Osteosarc applies source corrections
by default. Use Dataset.open(corrections=False) or
osteosarc --no-corrections to see the published values.
Code is Apache-2.0. The dataset is listed as CC0-1.0 in the AWS Open Data Registry. Cite the dataset and your access date when using it.
Development
python -m pip install -e '.[test]'
ruff check osteosarc tests scripts
python -m pytest -q
See testing for documentation builds and live-example checks.
Release files for osteosarc 0.7.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| osteosarc-0.7.0.tar.gz | 883.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| osteosarc-0.7.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.5 MB
Release files / osteosarc-0.7.0.tar.gz
| Download URL | osteosarc-0.7.0.tar.gz |
|---|---|
| Size | 883.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
eb7f154f33988e6074494924693fb82305f60a36c92d25e15eb95b82c1442fde
|
|
BLAKE2b-256 checksum How to use checksums |
0322df9ccd9965fb1590823565c6133799e828fdd2592b77ed01681fc0f7b23a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.6
|
Release files / osteosarc-0.7.0-py3-none-any.whl
| Download URL | osteosarc-0.7.0-py3-none-any.whl |
|---|---|
| Size | 569.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b9fd5e8c1e01491ca2f9f8a27ea5a6ce4be38a3da2c7d8ce28c5d57bfb5dbfce
|
|
BLAKE2b-256 checksum How to use checksums |
60f59d3a27359abbe18067debf922e7700fb362a224151ac4807948e61295a2f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.6
|