Skip to main content

osteosarc

Python library and command-line tool for the public osteosarc.com dataset: one patient's osteosarcoma sequencing, variant calls, cancer vaccines and clinical history. Find a sample's files, pick variants and vaccine peptides, and fetch the reads around a variant without downloading a whole BAM.

Documentation · Key concepts · Command line · Python API · Changelog

Features

  • Samples and their files. Every tumor, organoid and blood sample, what was sequenced, and its BAMs and FASTQ folders, with the commands that fetch them.
  • Browse without downloading. Search nearly 400,000 files by sample, kind, assay and folder, and see which you've already downloaded.
  • Variants and vaccine peptides. The site's variants with checked alleles, read counts, which pipelines found them, vaccine peptides and ELISPOT results.
  • Reads around a variant. Copy just the reads you need out of a remote BAM into a small local one.
  • Corrected by default. 35 fixes to known problems in the published data, each with its evidence, such as the MAP2 vaccine target's allele. Every load checks them against the snapshot's sources, and you can turn them off.
  • Clinical timeline. Every treatment, procedure, scan and MRD result on one chart, with a row per drug.
  • Reproducible. The site's metadata is saved as dated snapshots that reopen offline. The website changes; your results don't, until you sync again.
  • OpenVax integration. Works with Varcode, Isovar, Topiary and Vaxrank, builds reproducible test BAMs, and includes a list of 637 candidate structural variants.

Install

python -m pip install osteosarc

Needs Python 3.9+ on Linux or macOS, and SAMtools on your PATH to fetch reads.

Look around

osteosarc sync                # Once: about 57 MB of the website's metadata
osteosarc                     # Which snapshot you're using, and every command
osteosarc samples             # Samples and what was sequenced
osteosarc samples T1_tumor    # One sample's files, and commands to get them
osteosarc files               # What's in the bucket, by kind and folder
osteosarc variants --gene MAP2
osteosarc timeline            # Treatments, procedures, scans and MRD

Every command prints text for people, and --json for scripts. osteosarc repl opens Python with the data loaded as data, and in Python or a notebook everything shows a readable preview:

from osteosarc import Dataset

data = Dataset.sync()   # or Dataset.open() to reopen your newest snapshot offline
data                    # What's here, and how to get it
data.samples            # Samples and their sequencing
data.samples["T1_tumor"]

From samples to reads

from osteosarc import Dataset

data = Dataset.sync()  # Save today's website metadata (about 57 MB)

rna = data.samples["T0_tumor"].files.select(kind="alignment", assay="rna-seq")
targets = data.variants(gene="DYNC1H1", status="ready")

source = rna["rna-seq/reprocessed/BG003082/BG003082.Aligned.sortedByCoord.out.md.bam"]
reads = data.extract_reads(source, variants=targets, padding=100)
print(reads.path)  # Local indexed BAM

The BAM key is the file's path in the dataset's public S3 bucket. Only the reads near the variants are downloaded, into a small indexed BAM in your local cache. data.download(key, to=".") fetches a whole file instead, and data.downloads() lists what you have.

Later, Dataset.open() reopens your most recent snapshot without a network connection. The website changes over time, so snapshots are saved by download date: osteosarc snapshots lists them, and Dataset.open(date="2026-09") or --snapshot 2026-09 picks the newest from that month. Get started explains each step.

The same workflow from the terminal:

osteosarc sync
osteosarc samples T0_tumor
osteosarc files --sample T0_tumor --kind alignment --assay rna-seq
osteosarc variants --gene DYNC1H1 --status ready
osteosarc reads rna-seq/reprocessed/BG003082/BG003082.Aligned.sortedByCoord.out.md.bam --variant DYNC1H1-chr14-101980529 --variant DYNC1H1-chr14-102030200 --padding 100
osteosarc downloads

Guides

I want to… Read
Understand samples, files, variant status, coordinates and corrections Key concepts
Find a sample's RNA, DNA, single-cell or long-read files, and download them Samples and files
Get alleles, read counts or vaccine peptides Select variants
Fetch, filter or pair reads by variant or region Extract reads
Browse treatments, MRD and labs Browse the timeline
Pass data to Varcode, Isovar, Topiary or Vaxrank Use other libraries
Build small, verifiable test BAMs Read fixtures
Explore candidate structural variants SV interest catalogue

Data, license and citation

Osteosarc applies source corrections by default. Use Dataset.open(corrections=False) or osteosarc --no-corrections to see the published values.

Code is Apache-2.0. The dataset is listed as CC0-1.0 in the AWS Open Data Registry. Cite the dataset and your access date when using it.

Development

python -m pip install -e '.[test]'
ruff check osteosarc tests scripts
python -m pytest -q

See testing for documentation builds and live-example checks.

Release files for osteosarc 0.9.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for osteosarc 0.9.0
File Size Uploaded
osteosarc-0.9.0.tar.gz 904.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for osteosarc 0.9.0
File Interpreter ABI Platform
osteosarc-0.9.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.5 MB

Release files / osteosarc-0.9.0.tar.gz

Download URL osteosarc-0.9.0.tar.gz
Size 904.8 kB
Tags Source
SHA-256 checksum
How to use checksums
f2c13f44b6cb0f6a2c2021e547c37d717c6f2488d5cbb2ac08092c99fcef3c94
BLAKE2b-256 checksum
How to use checksums
988fa5c1cbea7fb2c9e95423400f950e78bc663fe9235a32bf6ca30ff8153f73
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release files / osteosarc-0.9.0-py3-none-any.whl

Download URL osteosarc-0.9.0-py3-none-any.whl
Size 582.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
237afa0aba81b0eb76ba6f7e01b5c514268f6a8625c9eace08e9cf37b8b0cf1e
BLAKE2b-256 checksum
How to use checksums
816128d715b95e3c36997f91420456eee1c3a3e8bbdc0119f3de9aa243914d57
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.6

Release history Release notifications | RSS feed

0.10.0

2 release files

This release

0.9.0 This release

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page