edc_lab_results_import
Imports lab results from a folder of PDF reports into the Result model, resolves each result to the subject, timepoint and requisition it belongs to, and reports what could not be resolved or does not agree with the CRF it was transcribed onto.
See also download-gmail-pdfs. If you are using it, run download-gmail-pdfs before running import_results.
Your import_results configuration will need a parser for your specific PDF format. See also parse-trial-labs.
Settings
# parser callable per laboratory, keyed by laboratory name
EDC_LAB_RESULTS_PARSERS = {"MNH": "..."}
# utest id and unit mapping files per laboratory
EDC_LAB_RESULTS_MAPPING_FILES = {"MNH": "..."}
# where result PDFs are uploaded for import
EDC_LAB_RESULTS_IMPORT_UPLOAD_DIR = "~/edc/lab_results/upload"
# where source PDFs are archived after import
EDC_LAB_RESULTS_IMPORT_STORAGE_DIR = "~/edc/lab_results/storage"
# which requisition panel an analyte panel is drawn under
EDC_LAB_RESULTS_IMPORT_REQUISITION_PANEL_MAP = {"wbc_diff": "fbc"}
The upload and storage folders are kept apart:
Upload folder (EDC_LAB_RESULTS_IMPORT_UPLOAD_DIR). Put result PDFs here, by hand or with a tool like download-gmail-pdfs. import_results reads the PDFs from this folder and leaves them in place.
Storage folder (EDC_LAB_RESULTS_IMPORT_STORAGE_DIR). The archive of original PDFs behind SourceDocument. On import, each PDF is copied here, stored by its sha256, and served from here to users with permission. The application manages this folder. Do not put files here by hand.
Both folders must exist, and the upload folder may not be the storage folder or inside it. System checks E004 to E008 report otherwise. The storage folder holds PII and should sit outside of MEDIA_ROOT.
The panel map needs explaining. The panel a result is reported under is not always the panel it was drawn under. The white cell differentials are their own analyte panel, but no visit schedule requires a wbc_diff requisition because they are collected on the FBC requisition. Without this map, looking a requisition up by the analyte panel can only fail. It is empty by default and an unlisted panel maps to itself. Result.panel_name keeps the analyte panel either way, which is the truthful description of what was measured.
A first import
manage.py import_results --laboratory MNH --dry-run
manage.py import_results --laboratory MNH
Both read from the upload folder. Pass a folder as the first argument to read from somewhere else.
Then check what did not resolve:
from edc_lab_results_import.dataframes import get_df_orphan_results
df = get_df_orphan_results()
df.bucket.value_counts()
On a healthy import:
- panel_unknown
Should be 0. Anything here is a utest id that is in no registered panel, so the result can never be matched to a requisition. Fix the mapping before going further. See get_mappings.
- resolver_miss
Should be 0 or near it. A requisition already exists for this result and the importer did not find it, which means the importer missed something rather than the data being wrong. Investigate before linking it away.
- requisition_not_keyed
The genuine data manager worklist: the panel was expected at this timepoint and nobody keyed the requisition.
- visit_not_found
No timepoint. Check candidate_rule for the ones the baseline rule can place.
- panel_not_expected
A real panel at a timepoint that did not call for it. An ad hoc draw, or a mapping worth re-examining.
A data manager keys a requisition, not an analyte, so group by subject, timepoint and panel for the size of the actual work:
df.groupby(
["subject_identifier", "visit_code", "visit_code_sequence", "panel_name"]
).ngroups
Then compare the imported values against the CRFs they were transcribed onto, which is the point of all of it:
from edc_lab_results_import.dataframes import get_df_result_comparison
df = get_df_result_comparison()
df[df.value_status == "differs"].sort_values("pct_diff", ascending=False)
df[df.ratio.between(9.5, 10.5)] # decimal point slips
df[(df.n_imported_for_key == 0) & df.crf_value.notna()] # keyed, nothing imported
df[df.n_imported_for_key > 1] # corrected reports
Repairing a database imported before these fixes
Results imported before panel_name was persisted carry an empty panel, so every join keyed on panel is dead: the requisition lookup, the requisition metadata lookup, and the related visit fallback in get_df_result_comparison.
Re-importing does not repair them. save_to_model skips a row whose unique key already exists and never updates it. The repair path is the backfill.
manage.py backfill_panel_name --dry-run
manage.py backfill_panel_name
Set EDC_LAB_RESULTS_IMPORT_REQUISITION_PANEL_MAP, then re-read the report. The backfill writes the analyte panel and does not consult the map, so the two can be done in either order, but the map must be set before the report or the linker mean anything.
manage.py link_orphan_results --dry-run
manage.py link_orphan_results
link_orphan_results consumes the resolver_miss bucket only: results carrying no requisition where one is already keyed at their timepoint for their panel. RequisitionModelMixin.Meta constrains panel and related visit to be unique together, so the target is determined rather than guessed, and no date is matched on. It reads the report itself rather than an exported worklist, so it cannot act on a stale one, and it re-reads each result before writing so one linked in the meantime is left alone.
Both commands save one row at a time so simple_history records every change and a run stays reversible from the audit trail. Both are resumable: an interrupted run picks up where it stopped.
What the repair does not fix
Two things stay as they are for rows imported before these changes.
Results with no timepoint. The baseline pass in ResultImporter.match_baseline_visits runs at import time only. A specimen collected on or before a subject’s first visit cannot belong to a later timepoint, so baseline is the only candidate, bounded by MAX_DAYS_BEFORE_BASELINE. Existing rows keep their empty subject_visit. get_df_orphan_results proposes one in candidate_subject_visit_id and candidate_rule, and names the requisition waiting there in candidate_requisition_id, but nothing writes it. A writer would follow the shape of ResultLinker.
Converted values. converted_result_value and converted_units are declared on the model and read back by model_to_dataframe, but nothing computes them: apply_unit_mapping_after_resolve only rewrites units in place. So the units fallback in the comparison rule never fires, and a result whose units differ from the CRF reads not_compared rather than being converted and compared.
Knowing when a frame has gone stale
The frames are queries, not tables, so calling one again always reflects the database. What goes stale is a dataframe held in a notebook or a spreadsheet someone exported days ago. result_expected in particular is a user driven change, so a worklist pulled before a site edited its requisitions will overstate the work.
from edc_lab_results_import.dataframes import changed_since_pulled
df = get_df_orphan_results()
changed_since_pulled(df)
Any non-zero count means read it again. Note that DataFrame.attrs survives a notebook but not a round trip through CSV, so an exported worklist loses its timestamp.
The comparison rule
comparison_rules.add_comparison_columns is the single implementation of whether an imported result agrees with the CRF. ResultComparison applies it to the rows behind one CRF on the result search page, and get_df_result_comparison applies it to every result CRF in the trial, so the page and the dataframe cannot disagree about the same row.
Value, units and the abnormal flag are compared separately, so a row that agrees on the value but not the units still reads as a difference. Differences are taken at the precision the CRF stores, so a value differing only in decimal places the CRF does not hold reads as exactly 0.0 and sorts to the bottom.
Metadata
Release files for edc-lab-results-import 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| edc_lab_results_import-1.0.0.tar.gz | 53.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| edc_lab_results_import-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 128.6 kB
Release files / edc_lab_results_import-1.0.0.tar.gz
| Download URL | edc_lab_results_import-1.0.0.tar.gz |
|---|---|
| Size | 53.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
46402aa25eb3d034632d191ed593e5e7406a9ed4375af42f15b413a25594bd0f
|
|
BLAKE2b-256 checksum How to use checksums |
321b376fd00637a3d5abc660961404275162aff4ee87eb01ee3b100933ad6644
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.18 {"installer":{"name":"uv","version":"0.12.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / edc_lab_results_import-1.0.0-py3-none-any.whl
| Download URL | edc_lab_results_import-1.0.0-py3-none-any.whl |
|---|---|
| Size | 75.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c63296ffd3319dd94389dd12cfc685a30cb69dddf1f679022ee9f7ba1f0e4099
|
|
BLAKE2b-256 checksum How to use checksums |
480c250acde6bffc1af96eb51bf558ff50c2d56bc31e4d70cd4ced933a9840a9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.18 {"installer":{"name":"uv","version":"0.12.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|