Clinical trial data normalization and report generation toolkit for MatchMiner integration
Project description
Michigan Medicine Clinical Trial Integration Toolkit ⚒️
This repo prepares data from various sources to integrate with popular open-source clinical trial matching software (MatchMiner supported, TrialMatchAI at some point). Beyond providing input sources for popular clinical trial matching software, this repo also provides post-matching processing scripts to generate reports. For instance, once clinical trial matching has completed, it interfaces with the structured output (from MatchMiner) to provide a report for the patient.
Quick Start
# Step 1: Normalize patient data from Excel template
ctm-mm patients <patient_data_template.xlsx> --pt-uuid 1234 --out pt_1234.json
# Step 2: Normalize trial data from AMC XML or ClinicalTrials.gov
ctm-mm trials --amc <amc_trials.xml> --sparrow <sparrow.xlsx> --west <west.xlsx> --out trials.json
# Note - We can also fetch trials directly from ClinicalTrials.gov without any template:
# Step 2 (alternate): Fetch a single trial from ClinicalTrials.gov (raw or normalized)
ctm-fetch --nct NCT03067181 --output nct-raw.json
# Step 2.5 (alternate): Normalize clinicaltrials.gov data to matchminer (mm)
ctm-fetch --nct NCT03067181 --output nct-normalized.json --fmt-mm
# Step 3
# Ensure you do the MANUAL PROCESSING from the normalized format to MATCHMINER!!!
# This involves translating text eligiblity criteria into MatcherMiner format
# Step 4: Load trial data into MatchMiner
python -m matchengine.main load -t trials.json --trial-format json --db test
# Step 5: Run MatchMiner
python -m Matcher.main --config source/Matcher/config/config.json
# Step 6: Export match results from MatchMiner's Mongo database
SECRETS_JSON=SECRETS_JSON.json python export_matches.py --patient 7439568 --output export/ --db v1
# Step 7: Build report from match results
ctm-report --pt data/patient.json --matches data/matchminer_export.json --engine mm --out output.pdf
Warning for Patient and Clinical Trial Data
Patient Data
Patient data is based upon manual recording of patient data into a template excel sheet. Until we can automate the process of pulling a patient's file and normalizing it, this will have to do. See example at /data/raw/pt_template.xlsx.
After you fill out the template excel sheet, please be careful of:
- ONCOTREE_PRIMARY_DIAGNOSIS_NAME passes through directly from the Excel oncotree_primary_diagnosis column - if that cell has a free-text diagnosis instead of an Oncotree code, MatchMiner won't match on it correctly
- TRUE_HUGO_SYMBOL passes through from the gene column as-is - no validation against HGNC so be careful to use a HUGO symbol
- variant_type in the Excel drives VARIANT_CATEGORY mapping - if someone enters an unrecognized value it gets skipped with a warning
Clinical Trial Data
Clinical trial data comes from automated sources for AMC and Sparrow. For AMC, an export of the OnCORE database is retreived in the form of an XML file (see data/raw/amc_trails_raw.xml). For Sparrow, an Excel document is emailed each month (see data/sparrow_trials_raw.xlsx).
For UMH-West, an Excel sheet is emailed is a format that is incomplete and requires manual processing immediately. Thus, this repo holds a template for recording West trial data instead of the raw data file we receive. See the template for West's trial data at data/raw/west_trial_template.xlsx --> note the *_template.xlsx instead of *_raw.xlsx ;-)
The automation from AMC, Sparrow, West, and ClinicalTrials.gov clinical trial data requires some manual massaging of the data before it can be put into MatchMiner:
- treatment_list.step[0].match is auto-populated with age constraints only (AMC) or age + gender for sex-restricted trials (CTGov) - all genomic and oncotree criteria must be filled in manually
- Eligibility criteria are parsed and preserved structurally (inclusion/exclusion with nesting) but are not converted into MatchMiner and/or match nodes automatically
- AMC trials have no arm structure - a single stub arm (ARM 1) is generated; if the trial has multiple arms or dose cohorts they must be added manually
- status is normalized automatically (OPEN TO ACCRUAL → open to accrual) but verify any trial with an unusual status value, as unrecognized values pass through unchanged
- AMC data has no drug/intervention field - _summary.drugs will always be empty and must be populated manually if needed for display
- protocol_no is only populated for AMC trials; CTGov trials will have null there unless manually set
Background
Creating a clinical trial match involves two key categories of data:
- patient information (query)
- clinical trial information (reference)
Patient data must match up to trial eligibility requirements to make a match. Trial eligibility often relies on some combination of patient demographics, diagnoses, prior treatments, as well as molecular and genetic data. Combining all of this information into one spot can be difficult, thus the purpose of this toolkit is to make it easier to combine all the data sources needed into the formats required for popular clinical trial engines and report building.
For matching a patient to trials, we define 2 categories of patient data:
- clinical data
- genomic data
Clinical data is kind of an overloaded term, but it refers to general patient demographics as well as diagnoses. Examples include Age, Sex, Primary Diagnosis, Weight, etc. Genomic data comes from molecular testing, which is either performed in-house (AMC's Division of Diagnostic Genetics and Genomics) or from an external company (Tempus, Caris, Foundation, to name a few).
Data pipeline
- Raw --> Normalized Patient Data
- Fill out patient_data_template.xlsx
- Run
$ ctm-mm patients ...
- Raw --> Normalized Trial Data. This is much more complex since we have 4 data sources (Sparrow, West, AMC, and ClinicalTrials.gov)
- For AMC, you can point to the XML file to create a structure very similar to Matchminer:
$ ctm-mm trials --amc nct-raw.json --out to-normalized.json- Above produces a .JSON file that needs a little bit of manual curation to be ingested by MatchMiner
- For ClinicalTrials.gov data, you can fetch a particular trial by NCT number and then normalize to the same format as AMC
- Run
$ ctm-fetch --nct NCT03067181 --output nct-raw.json --fmt-mm# fetch clinical trials and normalize output to matchminer format - You can also run
$ ctm-mm trials --ct <raw-ctgov.json> --out to-normalized.json# given a JSON of unnormalized clinicaltrials.gov trials, this will normalize it to matchminer format - Please note 2.1 and 2.2 have overlapping functionality. the --ct <raw-ctgov.json> in 2.2 would come from running
ctm-fetch --nct NCT03067181 --output nct-raw.json, which omitting the--fmt-mm, stored the data in a non-normalized way (raw from ct.gov api).
- Run
- Ensure you finish manually curating eligibility criteria for (1) and (2)
- For AMC, you can point to the XML file to create a structure very similar to Matchminer:
- Read MatchMiner docs to prepare for the next 2 steps!
- Load in the new data to MatchMiner
$ python -m matchengine.main load ...
- Execute the match!
$ python -m Matcher.main ...
- Export match data from MatchMiner database
$ python export_matches.py ...
- Build the report from a combination of patient and match data
$ ctm-report ...
Schemas
Schema levels:
schemas/raw/- Source-faithful models: one per data source (PT Excel sheets, AMC XML, Sparrow XLSX, West XLSX, and CTGov API). Fields mirror source column names exactly.schemas/matchminer/- MatchMiner-ready models.- Patient data (
MMClinical,MMGenomic) - Normalized trial data (
ClinicalTrialNormalized. Trial documents require manual curation of the match tree before loading into MatchMiner. - Trial match data (
MMTrialMatchExport). This is the output generated after MatchMiner matches a patient to a trial and the data is exported out of MongoDB.
- Patient data (
schemas/trialmatchai/- Placeholder for future TrialMatchAI integration.
Normalized Trials and CTML Curation
ctm-mm trials outputs a list of ClinicalTrialNormalized documents (schema: src/ctm/schemas/matchminer/clinical_trial.py). Each document has the following top-level keys:
| Key | Description |
|---|---|
protocol_no |
Internal protocol number. Populated for AMC trials only; null for CTGov-sourced trials unless set manually. |
nct_id |
ClinicalTrials.gov identifier. |
status |
Normalized status string: open to accrual, closed to accrual, or suspended. |
entity |
Source tag: amc, ctgov, sparrow, or west. |
treatment_list |
CTML match tree. Contains step → arm → match nodes. This is what MatchMiner queries against. |
eligibility |
Parsed inclusion/exclusion criteria with full nesting preserved. For reference only — not automatically wired into the match tree. |
_summary |
Display-friendly summary fields: phase, age group, sponsor, drugs, conditions, PI, etc. |
_raw |
Complete source dump for zero data loss. Pattern B sources (Sparrow, West) also include a _raw._sparrow or _raw._west sub-key with the original Excel row. |
Integration patterns
- Pattern A (AMC): XML → raw schema →
ClinicalTrialNormalizeddirectly. No CTGov lookup. - Pattern B (Sparrow, West): Excel provides NCT IDs → each trial is fetched from CTGov → normalized via the CTGov pipeline → source metadata merged into
_raw.
What is auto-populated vs. what needs manual curation
ctm-mm trials gets you most of the way there, but the match tree requires manual work before MatchMiner can use it:
treatment_list.step[0].matchis seeded with age constraints only (AMC), or age + gender for sex-restricted CTGov trials- All genomic criteria (Hugo symbol, protein change, variant category, etc.) must be added manually as
genomicmatch nodes - All oncotree/cancer type criteria must be added manually as
clinicalmatch nodes - Eligibility free text is parsed and stored in
eligibilityas a reference, but is not automatically converted into match nodes - AMC trials have no arm structure — a single stub arm (
ARM 1) is generated; multi-arm or dose-escalation trials must be expanded manually
Loading into MatchMiner
Once the match tree is curated, load the trial documents:
python -m matchengine.main load -t trials.json --trial-format json --db <your-db>
MatchMiner expects each document to conform to CTML format. See the MatchMiner docs for the full field reference.
Setup
Installation
uv pip install -r requirements.txt --python .venv/bin/python
Running tests
.venv/bin/python -m pytest
On macOS, WeasyPrint also needs the native Pango library, which isn't installed by default:
brew install pango
Disclaimer about MatchMiner and MongoDB
Matchminer uses MongoDB to faciliate patient-trial matches. In a nutshell, it stores all patient data in 2 collections: i) clinical and ii) genomic. It stores all clinical trials in the trial collection. When you run $ python -m matchengine.main load ..., it stores all the patient and trial data in these collection.
When we run $ python -m Matcher.main ..., this connects to the database and matches all clinical+genomic docs for each patient with eligible trials in the trial collection and saves the results in the trial_match collection.
We need to connect to MongoDB to run Matchminer. This requires us to provide a username, password, and connection url to Mongo. This info is stored in a file called SECRETS_JSON.json. We also need to ensure we have Mongo installed with an instance running for us to make this work. See our forked MatchMiner README for more info.
Build a report
To build a PDF report from the mock data:
$ ctm-report --pt data/mock/pt-data.json --matches data/mock/mm-matches.json --engine mm --out output.pdf
The data/mock/pt-data.json file follows the same format that is output from ctm-mm patients <patient_data_template.xlsx> --out pt-data.json. This is MatchMiner-normalized data that is in JSON with 3 high-level keys: i) clinical ii) genomic iii) extras. The schema can be found at src/ctm/schemas/matchminer/patient.py in MMClinical and MMGenomic. The report actually only ever touches the extras sub-key, since all the rest of the data can come from the MatchMiner match output.
The data/mock/mm-matches.json file follows the schema found in src/ctm/schemas/matchminer/trial_match.py, specifically in the model MMTrialMatchExport. It is the JSON export produced by MatchMiner after running the matching engine against a patient. It has 3 top-level keys: i) pt_uuid ii) clinical (a minimal patient reference with SAMPLE_ID, MRN, etc.) and iii) trial_match (a list of match records). The primary match is selected by preferring match_level: arm over step, then reason_type: genomic over clinical, then sort_order. All remaining unique trials appear as secondary matches.
For a live preview in your browser:
$ ctm-report --pt data/mock/pt-data.json --matches data/mock/mm-matches.json --engine mm --preview
This opens a browser tab at http://localhost:5500/report.html. Any edit to a
template in templates/, the stylesheet in static/report.css, or the active
data directory (data/mock/ or data/real/, depending on the flag)
automatically re-renders the report and refreshes the page.
Project layout
Folders
data/- see Data directories table abovetemplates/-report.htmlis the base page,_*.htmlare the per-section includesstatic/report.css- shared styling, including the@pagerule for PDF page size/marginssrc/ctm/reports/builder.py- loads JSON data and renders the Jinja2 template to HTMLpreview.py- live-reload dev serverbuild_pdf.py- renders the HTML and exports it to PDF via WeasyPrint
Data directories
We have some example data we store in this repo as a nice reference.
| Directory | Purpose |
|---|---|
data/content/ |
Static content used to build the report. |
data/dump/ |
Ignore for now. It's where we dump random raw trial and patient data. |
data/mock/ |
Synthetic data that shows the format of patient, trial, and match data and is used to build a mock report to show the report format. Patient and trial data are normalized, while match data is exported from the match engine (MatchMiner only for now) |
data/raw/ |
Templates for manually creating initial patient and clinical trial data. This is the input for the normalization step |
Extensions
MatchMiner's matching fields are config-driven, making it straightforward to add new matchable fields without touching the core engine. The config lives at matchengine/config/dfci_config.json in the matchengine-V2 repo.
Adding a new genomic field
To add a boolean or freeform string field, add an entry to ctml_collection_mappings.genomic.trial_key_mappings, add the field to projections.genomic, and add it to indices.genomic. Use "nomap" when the trial and patient values should match exactly with no transformation.
Example — boolean field:
"ALIEN_MUTATION": {
"sample_key": "ALIEN_MUTATION",
"sample_value": "nomap"
}
A trial requiring this field would include in its match node:
{ "genomic": { "alien_mutation": true } }
And the patient's genomic document would need:
{ "ALIEN_MUTATION": true }
nomap works equally well for freeform strings — the trial value and patient value must simply match exactly. If you control both sides (i.e. you populate the patient data and curate the trial CTML), consistency is easy to maintain.
Adding a new clinical field
Same pattern, but under ctml_collection_mappings.clinical.trial_key_mappings. MatchMiner will query the clinical collection instead of genomic. Also add the field to MMClinical in src/ctm/schemas/matchminer/patient.py and populate it from the patient Excel template.
Example — patient stage:
"STAGE": {
"sample_key": "STAGE",
"sample_value": "nomap"
}
Adding a field with a custom value transform
When source data uses inconsistent representations of the same value, use a named transform function instead of "nomap". The function is defined in matchengine/plugins/DFCIQueryTransformers.py and translates the trial's criterion value into the MongoDB query before matching.
Example — REVERSE_KEY field that reverses a string before matching:
In dfci_config.json:
"REVERSE_KEY": {
"sample_key": "REVERSE_KEY",
"sample_value": "reverse_map"
}
In DFCIQueryTransformers.py, add a method to the transformers class:
def reverse_map(self, sample_value, trial_value):
return {"REVERSE_KEY": trial_value[::-1]}
A trial with "reverse_key": "olleh" would then match a patient whose REVERSE_KEY field is "hello". In practice, transforms are used to normalize value variants — e.g. "MSI-High", "MSI_H", and "MSI-H" all resolving to the same canonical match — rather than to reverse strings, but the mechanism is identical.
How patient data enters the match
Patient data never flows through the Python transform layer. The actual match works in three steps:
- Transform function — receives only the trial's criterion value and returns a MongoDB query object, e.g.
{ MMR_STATUS: "Deficient (MMR-D / MSI-H)" } - Engine — appends a
$in: [clinical_ids]filter to that query and fires it directly at MongoDB - MongoDB — executes the comparison against stored patient documents and returns matching IDs
The transform function has no knowledge of any specific patient. It only translates trial language into MongoDB query syntax. The patient document is a MongoDB concern exclusively.
Current limitations of the transform layer
Because transform functions only receive the trial's criterion value and never see the patient document, any eligibility criterion whose threshold depends on other patient fields cannot be expressed in CTML and must be flagged for manual clinician review.
Examples of criteria that cannot be automated:
- Serum creatinine by age/gender table — the acceptable creatinine threshold varies by both age and gender. Evaluating this requires knowing the patient's age and gender at query time, which the transform function cannot access.
- Lab value ≤ 1.5x ULN — the upper limit of normal (ULN) is age-dependent, so
1.5x ULNis not a fixed number. A fixed approximation (e.g. adult ULN) will be wrong for pediatric patients.
The general pattern: any criterion of the form "threshold depends on [other patient attribute]" falls outside what the plugin layer can handle without custom engine-level code. These criteria are preserved in eligibility.inclusion and eligibility.exclusion for human review, which is why that field is kept even though it is not wired into the match tree.
Exception — age/gender-conditional thresholds can be expressed as a large or block:
Rather than requiring a custom transform, each row of a lookup table (e.g. serum creatinine by age and gender) becomes one and branch inside an or. No new engine code is needed — only a new clinical field (e.g. CREATININE) with a range transform, and verbose but correct CTML curation:
{
"match": [
{
"or": [
{
"and": [
{ "clinical": { "age_numerical": ">=1", "age_numerical": "<6" } },
{ "clinical": { "creatinine": "<=0.8" } }
]
},
{
"and": [
{ "clinical": { "age_numerical": ">=6", "age_numerical": "<10" } },
{ "clinical": { "creatinine": "<=1.0" } }
]
},
{
"and": [
{ "clinical": { "age_numerical": ">=16", "gender": "Male" } },
{ "clinical": { "creatinine": "<=1.7" } }
]
}
]
}
]
}
8 age bands × 2 genders = up to 16 and branches, but it is fully correct and requires no engine changes — just verbose curation.
TODO
As it stands, the pipeline has multiple manual processing steps. Ideally, we can automatically integrate from the raw data sources and take it all the way to matching and report building.
- make a pipeline that points to raw patient+trial info and spits out match report
- docker-ize this process
- deploy this onto PS1A server
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ctm_toolkit-0.1.0.tar.gz.
File metadata
- Download URL: ctm_toolkit-0.1.0.tar.gz
- Upload date:
- Size: 146.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
167f1f782d54c5a9460252c822a897537d971b79189d9d268ba25f034315137f
|
|
| MD5 |
f60bc76b9a14d2501eb420fa9db77e6a
|
|
| BLAKE2b-256 |
115d3b4f8647d6037edfec5d9ae53b9e3a53fb3aa69fdd689233d84c1ea02547
|
File details
Details for the file ctm_toolkit-0.1.0-py3-none-any.whl.
File metadata
- Download URL: ctm_toolkit-0.1.0-py3-none-any.whl
- Upload date:
- Size: 45.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
5baf52ed5d2a7cf787687c64c301b67e16a513f4e3fb7e165d09284cd8c174db
|
|
| MD5 |
c469c7489104de55943f6e18aa537848
|
|
| BLAKE2b-256 |
a55661d1e2b5d1a191f44db989edf25f430c3d58858b0c19aca87c2ef7252161
|