Skip to main content

edc_lab_results_import

Imports lab results from a folder of PDF reports into the Result model, resolves each result to the subject, timepoint and requisition it belongs to, and reports what could not be resolved or does not agree with the CRF it was transcribed onto.

See also download-gmail-pdfs. If you are using it, run download-gmail-pdfs before running import_results.

Your import_results configuration will need a parser for your specific PDF format. See also parse-trial-labs.

Settings

# parser callable per laboratory, keyed by laboratory name
EDC_LAB_RESULTS_PARSERS = {"MNH": "..."}

# utest id and unit mapping files per laboratory
EDC_LAB_RESULTS_MAPPING_FILES = {"MNH": "..."}

# where result PDFs are uploaded for import
EDC_LAB_RESULTS_IMPORT_UPLOAD_DIR = "~/edc/lab_results/upload"

# where source PDFs are archived after import
EDC_LAB_RESULTS_IMPORT_STORAGE_DIR = "~/edc/lab_results/storage"

# which requisition panel an analyte panel is drawn under
EDC_LAB_RESULTS_IMPORT_REQUISITION_PANEL_MAP = {"wbc_diff": "fbc"}

The upload and storage folders are kept apart:

  • Upload folder (EDC_LAB_RESULTS_IMPORT_UPLOAD_DIR). Put result PDFs here, by hand or with a tool like download-gmail-pdfs. import_results reads the PDFs from this folder and leaves them in place.

  • Storage folder (EDC_LAB_RESULTS_IMPORT_STORAGE_DIR). The archive of original PDFs behind SourceDocument. On import, each PDF is copied here, stored by its sha256, and served from here to users with permission. The application manages this folder. Do not put files here by hand.

Both folders must exist, and the upload folder may not be the storage folder or inside it. System checks E004 to E008 report otherwise. The storage folder holds PII and should sit outside of MEDIA_ROOT.

The panel map needs explaining. The panel a result is reported under is not always the panel it was drawn under. The white cell differentials are their own analyte panel, but no visit schedule requires a wbc_diff requisition because they are collected on the FBC requisition. Without this map, looking a requisition up by the analyte panel can only fail. It is empty by default and an unlisted panel maps to itself. Result.panel_name keeps the analyte panel either way, which is the truthful description of what was measured.

A first import

manage.py import_results --laboratory MNH --dry-run
manage.py import_results --laboratory MNH

Both read from the upload folder. Pass a folder as the first argument to read from somewhere else.

Then check what did not resolve:

from edc_lab_results_import.dataframes import get_df_orphan_results

df = get_df_orphan_results()
df.bucket.value_counts()

On a healthy import:

panel_unknown

Should be 0. Anything here is a utest id that is in no registered panel, so the result can never be matched to a requisition. Fix the mapping before going further. See get_mappings.

resolver_miss

Should be 0 or near it. A requisition already exists for this result and the importer did not find it, which means the importer missed something rather than the data being wrong. Investigate before linking it away.

requisition_not_keyed

The genuine data manager worklist: the panel was expected at this timepoint and nobody keyed the requisition.

visit_not_found

No timepoint. Check candidate_rule for the ones the baseline rule can place.

panel_not_expected

A real panel at a timepoint that did not call for it. An ad hoc draw, or a mapping worth re-examining.

A data manager keys a requisition, not an analyte, so group by subject, timepoint and panel for the size of the actual work:

df.groupby(
    ["subject_identifier", "visit_code", "visit_code_sequence", "panel_name"]
).ngroups

Then compare the imported values against the CRFs they were transcribed onto, which is the point of all of it:

from edc_lab_results_import.dataframes import get_df_result_comparison

df = get_df_result_comparison()
df[df.value_status == "differs"].sort_values("pct_diff", ascending=False)
df[df.ratio.between(9.5, 10.5)]                            # decimal point slips
df[(df.n_imported_for_key == 0) & df.crf_value.notna()]    # keyed, nothing imported
df[df.n_imported_for_key > 1]                              # corrected reports

Repairing a database imported before these fixes

Results imported before panel_name was persisted carry an empty panel, so every join keyed on panel is dead: the requisition lookup, the requisition metadata lookup, and the related visit fallback in get_df_result_comparison.

Re-importing does not repair them. save_to_model skips a row whose unique key already exists and never updates it. The repair path is the backfill.

manage.py backfill_panel_name --dry-run
manage.py backfill_panel_name

Set EDC_LAB_RESULTS_IMPORT_REQUISITION_PANEL_MAP, then re-read the report. The backfill writes the analyte panel and does not consult the map, so the two can be done in either order, but the map must be set before the report or the linker mean anything.

manage.py link_orphan_results --dry-run
manage.py link_orphan_results

link_orphan_results consumes the resolver_miss bucket only: results carrying no requisition where one is already keyed at their timepoint for their panel. RequisitionModelMixin.Meta constrains panel and related visit to be unique together, so the target is determined rather than guessed, and no date is matched on. It reads the report itself rather than an exported worklist, so it cannot act on a stale one, and it re-reads each result before writing so one linked in the meantime is left alone.

Both commands save one row at a time so simple_history records every change and a run stays reversible from the audit trail. Both are resumable: an interrupted run picks up where it stopped.

What the repair does not fix

Two things stay as they are for rows imported before these changes.

Results with no timepoint. The baseline pass in ResultImporter.match_baseline_visits runs at import time only. A specimen collected on or before a subject’s first visit cannot belong to a later timepoint, so baseline is the only candidate, bounded by MAX_DAYS_BEFORE_BASELINE. Existing rows keep their empty subject_visit. get_df_orphan_results proposes one in candidate_subject_visit_id and candidate_rule, and names the requisition waiting there in candidate_requisition_id, but nothing writes it. A writer would follow the shape of ResultLinker.

Converted values. converted_result_value and converted_units are declared on the model and read back by model_to_dataframe, but nothing computes them: apply_unit_mapping_after_resolve only rewrites units in place. So the units fallback in the comparison rule never fires, and a result whose units differ from the CRF reads not_compared rather than being converted and compared.

Knowing when a frame has gone stale

The frames are queries, not tables, so calling one again always reflects the database. What goes stale is a dataframe held in a notebook or a spreadsheet someone exported days ago. result_expected in particular is a user driven change, so a worklist pulled before a site edited its requisitions will overstate the work.

from edc_lab_results_import.dataframes import changed_since_pulled

df = get_df_orphan_results()
changed_since_pulled(df)

Any non-zero count means read it again. Note that DataFrame.attrs survives a notebook but not a round trip through CSV, so an exported worklist loses its timestamp.

The comparison rule

comparison_rules.add_comparison_columns is the single implementation of whether an imported result agrees with the CRF. ResultComparison applies it to the rows behind one CRF on the result search page, and get_df_result_comparison applies it to every result CRF in the trial, so the page and the dataframe cannot disagree about the same row.

Value, units and the abnormal flag are compared separately, so a row that agrees on the value but not the units still reads as a difference. Differences are taken at the precision the CRF stores, so a value differing only in decimal places the CRF does not hold reads as exactly 0.0 and sorts to the bottom.

Metadata

Release files for edc-lab-results-import 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for edc-lab-results-import 1.0.0
File Size Uploaded
edc_lab_results_import-1.0.0.tar.gz 53.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for edc-lab-results-import 1.0.0
File Interpreter ABI Platform
edc_lab_results_import-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 128.6 kB

Release files / edc_lab_results_import-1.0.0.tar.gz

Download URL edc_lab_results_import-1.0.0.tar.gz
Size 53.2 kB
Tags Source
SHA-256 checksum
How to use checksums
46402aa25eb3d034632d191ed593e5e7406a9ed4375af42f15b413a25594bd0f
BLAKE2b-256 checksum
How to use checksums
321b376fd00637a3d5abc660961404275162aff4ee87eb01ee3b100933ad6644
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.18 {"installer":{"name":"uv","version":"0.12.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / edc_lab_results_import-1.0.0-py3-none-any.whl

Download URL edc_lab_results_import-1.0.0-py3-none-any.whl
Size 75.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c63296ffd3319dd94389dd12cfc685a30cb69dddf1f679022ee9f7ba1f0e4099
BLAKE2b-256 checksum
How to use checksums
480c250acde6bffc1af96eb51bf558ff50c2d56bc31e4d70cd4ced933a9840a9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.18 {"installer":{"name":"uv","version":"0.12.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page