Skip to main content

osteosarc

Python library and command-line tool for the public osteosarc.com dataset: one patient's osteosarcoma sequencing, variant calls, cancer vaccines and clinical history. Find files, pick variants and vaccine peptides, and fetch the reads around a variant without downloading a whole BAM.

Documentation · Key concepts · Command line · Python API · Changelog

Features

  • Browse without downloading. Search nearly 400,000 files by sample, timepoint and assay: bulk and single-cell RNA, exome, genome, Oxford Nanopore and PacBio.
  • Variants and vaccine peptides. The site's variants with checked alleles, read counts, which pipelines found them, vaccine peptides and ELISPOT results.
  • Reads around a variant. Copy just the reads you need out of a remote BAM into a small local one.
  • Corrected by default. 35 fixes to known problems in the published data, each with its evidence, such as the MAP2 vaccine target's allele. Every load checks them against the snapshot's sources, and you can turn them off.
  • Clinical timeline. Treatments, procedures, imaging, MRD and lab results as a text chart, or in an interactive terminal explorer.
  • Reproducible. The site's metadata is saved as dated snapshots that reopen offline. The website changes; your results don't, until you sync again.
  • OpenVax integration. Works with Varcode, Isovar, Topiary and Vaxrank, builds reproducible test BAMs, and includes a list of 637 candidate structural variants.

Install

python -m pip install osteosarc

Needs Python 3.9+ on Linux or macOS, and SAMtools on your PATH to fetch reads.

Explore the data

osteosarc sync      # Once: about 57 MB of the website's metadata
osteosarc explore

The explorer opens with a summary of the data. Try samples, specimen T1_tumor, variants MAP2, timeline 2024-05 2024-09 and help; quit leaves. Running osteosarc on its own lists more commands to try.

In Python or a notebook, everything shows a readable preview:

from osteosarc import Dataset

data = Dataset.sync()
print(data.summary())            # What's here, and what to try next
print(data.variants(gene="MAP2"))

From samples to reads

from osteosarc import Dataset

data = Dataset.sync()  # Save today's website metadata (about 57 MB)
print(data.describe_samples())

rna = data.assets_for_sample("T0_tumor", kind="alignment", assay="rna-seq")
targets = data.variants(gene="DYNC1H1", status="ready")

source = rna["rna-seq/reprocessed/BG003082/BG003082.Aligned.sortedByCoord.out.md.bam"]
reads = data.extract_reads(source, variants=targets, padding=100)
print(reads.path)  # Local indexed BAM

The BAM key is the file's path in the dataset's public S3 bucket. Only the reads near the variants are downloaded, into a small indexed BAM in your local cache.

Later, Dataset.open() reopens your most recent snapshot without a network connection. The website changes over time, so snapshots are saved by download date: osteosarc snapshots lists them, and Dataset.open(date="2026-09") or --snapshot 2026-09 picks the newest from that month. Get started explains each step.

The same workflow from the terminal:

osteosarc sync
osteosarc samples
osteosarc assets --sample T0_tumor --kind alignment --assay rna-seq
osteosarc variants --gene DYNC1H1 --status ready
osteosarc reads rna-seq/reprocessed/BG003082/BG003082.Aligned.sortedByCoord.out.md.bam --variant DYNC1H1-chr14-101980529 --variant DYNC1H1-chr14-102030200 --padding 100
osteosarc explore

osteosarc explore opens an interactive browser for specimens, files, variants and the timeline. Type help for commands and quit to leave.

Guides

I want to… Read
Understand sample IDs, variant status, coordinates and corrections Key concepts
Find RNA, DNA, single-cell or long-read files and read tables Find samples and files
Get alleles, read counts or vaccine peptides Select variants
Fetch, filter or pair reads by variant or region Extract reads
Browse treatments, specimens, MRD and labs Browse the timeline
Pass data to Varcode, Isovar, Topiary or Vaxrank Use other libraries
Build small, verifiable test BAMs Read fixtures
Explore candidate structural variants SV interest catalogue

Data, license and citation

Osteosarc applies source corrections by default. Use Dataset.open(corrections=False) or osteosarc --no-corrections to see the published values.

Code is Apache-2.0. The dataset is listed as CC0-1.0 in the AWS Open Data Registry. Cite the dataset and your access date when using it.

Development

python -m pip install -e '.[test]'
ruff check osteosarc tests scripts
python -m pytest -q

See testing for documentation builds and live-example checks.

Release files for osteosarc 0.8.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for osteosarc 0.8.0
File Size Uploaded
osteosarc-0.8.0.tar.gz 888.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for osteosarc 0.8.0
File Interpreter ABI Platform
osteosarc-0.8.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.5 MB

Release files / osteosarc-0.8.0.tar.gz

Download URL osteosarc-0.8.0.tar.gz
Size 888.3 kB
Tags Source
SHA-256 checksum
How to use checksums
5fb92aae82b8440829286f57ac8b1a6836c798fac1989fd45edd837e2291a7ef
BLAKE2b-256 checksum
How to use checksums
ec03fa34c1da01d318e2c40089ed81812336c5890a01b66e01adaf6c4db9ef4e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release files / osteosarc-0.8.0-py3-none-any.whl

Download URL osteosarc-0.8.0-py3-none-any.whl
Size 571.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
dbdb8483aac858c000343b7cce6047a755a796bc287a579fcde4806afe49116c
BLAKE2b-256 checksum
How to use checksums
1db0471663cccc1e0218cd703e8a70983db6c492c6c18796c689e7aa1973f664
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release history Release notifications | RSS feed

0.9.0

2 release files

This release

0.8.0 This release

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page