Skip to main content

patient_frame

Build the EHR step's patient frame from the caches, not from the files.

ehr_criteria_query needs eight header fields per scan -- name, date of birth, patient ID, sex, study date, model, laterality and the file the row came from -- to identify each patient. scan_metadata already holds those fields, extracted header-only by a background runtime and keyed on the same resolved path the cohort is named by, so serving from there is one SQL query and touches no files. Reading them through the datasource instead costs a parse-cache lookup per file and, on a miss, a full file read, which on a network share dominates the step.

_disposition decides per file which route it takes. Only one disposition drops a file: a row saying the extractor read it and found no series in it (an unsupported format, an empty file), which re-reading would fail the same way. Every bucket is counted, so a run can report where its cohort went rather than leave a shortfall to be inferred from a patient count.

Module​

Functions​

build_patient_frame​

def build_patient_frame(    *, datasource: _Datasource, filenames: Sequence[str], cache: CacheProtocol | None,) ‑> PatientFrameResult:

Build the patient metadata frame for filenames.

Files are served from scan_metadata where it can answer for them and read through datasource where it cannot; see the module docstring for which is which.

Arguments

  • datasource: The imaging datasource, read only for the files scan_metadata cannot answer for.
  • filenames: The cohort's file paths, as the file_metadata inventory spells them.
  • cache: The pod's cache. When None there is nothing to read, so every file is read through the datasource — the behaviour callers had before this module existed.

Returns The frame and the per-bucket counts.

Classes​

PatientFrameResult​

class PatientFrameResult(    frame: pd.DataFrame,    served_from_scan_metadata: int,    fell_back: int,    fallback_absent: int,    fallback_stale: int,    fallback_processing_failed: int,    fallback_unknown_reason: int,    skipped_no_series: int,):

The patient frame, and an account of how each file contributed to it.

The counters exist so a run can say where its cohort went; a patient count alone leaves a shortfall to be reconciled by eye against the file count.

Variables​

  • static fallback_absent : int
  • static fallback_processing_failed : int
  • static fallback_stale : int
  • static fallback_unknown_reason : int
  • static fell_back : int
  • static served_from_scan_metadata : int
  • static skipped_no_series : int

Methods​


log_summary​

def log_summary(self, task_hash: str) ‑> None:

Log where the cohort went, at a level proportionate to the news.

Arguments

  • task_hash: The run's task hash, for correlating with other lines.