patient_frame
Build the EHR step's patient frame from the caches, not from the files.
ehr_criteria_query needs eight header fields per scan -- name, date of birth,
patient ID, sex, study date, model, laterality and the file the row came from --
to identify each patient. scan_metadata already holds those fields, extracted
header-only by a background runtime and keyed on the same resolved path the
cohort is named by, so serving from there is one SQL query and touches no files.
Reading them through the datasource instead costs a parse-cache lookup per file
and, on a miss, a full file read, which on a network share dominates the step.
_disposition decides per file which route it takes. Only one disposition
drops a file: a row saying the extractor read it and found no series in it (an
unsupported format, an empty file), which re-reading would fail the same way.
Every bucket is counted, so a run can report where its cohort went rather than
leave a shortfall to be inferred from a patient count.
Module
Functions
build_patient_frame
def build_patient_frame( *, datasource: _Datasource, filenames: Sequence[str], cache: CacheProtocol | None,) ‑> PatientFrameResult:Build the patient metadata frame for filenames.
Files are served from scan_metadata where it can answer for them and read
through datasource where it cannot; see the module docstring for which is
which.
Arguments
datasource: The imaging datasource, read only for the filesscan_metadatacannot answer for.filenames: The cohort's file paths, as thefile_metadatainventory spells them.cache: The pod's cache. WhenNonethere is nothing to read, so every file is read through the datasource — the behaviour callers had before this module existed.
Returns The frame and the per-bucket counts.
Classes
PatientFrameResult
class PatientFrameResult( frame: pd.DataFrame, served_from_scan_metadata: int, fell_back: int, fallback_absent: int, fallback_stale: int, fallback_processing_failed: int, fallback_unknown_reason: int, skipped_no_series: int,):The patient frame, and an account of how each file contributed to it.
The counters exist so a run can say where its cohort went; a patient count alone leaves a shortfall to be reconciled by eye against the file count.
Variables
- static
fallback_absent : int
- static
fallback_processing_failed : int
- static
fallback_stale : int
- static
fallback_unknown_reason : int
- static
fell_back : int
- static
frame : pandas.core.frame.DataFrame
- static
served_from_scan_metadata : int
- static
skipped_no_series : int
Methods
log_summary
def log_summary(self, task_hash: str) ‑> None:Log where the cohort went, at a level proportionate to the news.
Arguments
task_hash: The run's task hash, for correlating with other lines.