Skip to main content

Cancer Omics Drug Experiment Response Dataset

There is a recent explosion of deep learning algorithms that to tackle the computational problem of predicting drug treatment outcome from baseline molecular measurements. To support this,we have built a benchmark dataset that harmonizes diverse datasets to better assess algorithm performance.

This package collects diverse sets of paired molecular datasets with corresponding drug sensitivity data. All data here is reprocessed and standardized so it can be easily used as a benchmark dataset for the This repository leverages existing datasets to collect the data required for deep learning model development. Since each deep learning model requires distinct data capabilities, the goal of this repository is to collect and format all data into a schema that can be leveraged for existing models.

Coderdata Motivation

The goal of this repository is two-fold: First, it aims to collate and standardize the data for the broader community. This requires running a series of scripts to build and append to a standardized data model. Second, it has a series of scripts that pull from the data model to create model-specific data files that can be run by the data infrastructure.

Data access

For the access to the latest version of CoderData, please visit our documentation site which provides access to Figshare and instructions for using the Python package to download the data.

Data format

All coderdata files are in text format - either comma delimited or tab delimited (depending on data type). Each dataset can be evaluated individually according to the CoderData schema that is maintained in LinkML and can be udpated via a commit to the repository. For more details, please see the schema description.

Building a local version

The build process can be found in our coderbuild directory. Here you can follow the instructions to build your own local copy of the data on your machine.

Adding a new dataset

We have standardized the build (coderbuild) process so an additional dataset can be built locally or as part of the next version of coder. Here are the steps to follow:

  1. First visit the coderbuild directory and ensure you can build a local copy of CoderData.

  2. Checkout this repository and create a subdirectory of the coderbuild directory with your own build files.

  3. Develop your scripts to build the data files according to our LinkML Schema. This will require collecting the following metadata:

  • entrez gene identifiers (or you can use the genes.csv file
  • sample information such as species and model system type
  • drug name that can be searched on PubChem

You can validate each file by using the linkML validator together with our schema file.

You can use the following scripts as part of your build process:

  1. Wrap your scripts in standard shell scripts with the following names and arguments:
shell script arguments description
build_samples.sh [latest_samples] Latest version of samples generated by coderbuild
build_omics.sh [gene file] [samplefile] This includes the genes.csv that was generated in the original build as well as the sample file generated above.
build_drugs.sh [drugfile1,drugfile2,...] This includes a comma-delimited list of all drugs files generated from previous build
build_exp.sh [samplfile ] [drugfile] sample file and drug file generated by previous scripts
  1. Put the Docker container file inside the Docker directory with the name Dockerfile.[datasetname].

  2. Run build_all.py from the root directory, which should now add in your Dockerfile in the mix and call the scripts in your Docker container to build the files.

Metadata

Release files for coderdata 2.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for coderdata 2.2.1
File Size Uploaded
coderdata-2.2.1.tar.gz 25.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for coderdata 2.2.1
File Interpreter ABI Platform
coderdata-2.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 53.8 kB

Release files / coderdata-2.2.1.tar.gz

Download URL coderdata-2.2.1.tar.gz
Size 25.2 kB
Tags Source
SHA-256 checksum
How to use checksums
1f69cf16b36b7a7b42bc623f403959ad44ae5d40c56934fa9202e1be1cfa4c05
BLAKE2b-256 checksum
How to use checksums
2ecff8517065bd124e41c9814738c28c08399dadf090827677db1d08317338a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via python-httpx/0.27.2

Release files / coderdata-2.2.1-py3-none-any.whl

Download URL coderdata-2.2.1-py3-none-any.whl
Size 28.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b94da797d10c97ecf5029d1bb7a6c55bf2cb105e94c4291eb91625de30e444d7
BLAKE2b-256 checksum
How to use checksums
81618da70b5e72caa2f4f95c77296bacefa60d9c2619d89ff23e56b1e789f817
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via python-httpx/0.27.2

Release history Release notifications | RSS feed

This release

2.2.1 This release

2 release files

2.2.0

2 release files

2.1.0

2 release files

2.0.4

2 release files

2.0.3

2 release files

2.0.2

2 release files

2.0.1

2 release files

2.0.0

2 release files

0.1.40

2 release files

0.1.29

2 release files

0.1.28

2 release files

0.1.27

2 release files

0.1.26

2 release files

0.1.25

2 release files

0.1.22

2 release files

0.1.21

2 release files

0.1.20

2 release files

0.1.19

2 release files

0.1.18

2 release files

0.1.17

2 release files

0.1.16

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page