Skip to main content

functions

Choosing which of a run's files are worth paying to infer.

Pure selection, with no cache and no datasource: the task builds ReduceCandidates from the stored scan rows and this decides which survive. Keeping it free of both is what lets the interesting cases — ties, undated scans, a patient who cannot qualify — be tested without a database.

The one rule the whole module keeps: a file is removed only because a newer scan of the same patient supersedes it, or because its patient cannot reach min_scans anyway. A file this cannot reason about is kept, never dropped.

Module​

Functions​

select_files​

def select_files(    candidates: Sequence[ReduceCandidate], *, num_latest: int, min_scans: int = 1,) ‑> list[str]:

Return the files worth inferring, sorted.

Keeps the newest num_latest files in each (patient, *group_key) group, then drops any patient left holding fewer than min_scans of them. A candidate with no patient id is passed through untouched: it belongs to no group, so nothing can supersede it and no scan count applies to it.

A group whose members state two different patient_keys is not reduced at all. bitfount_patient_id drops middle names, prefixes and suffixes, so twins sharing a birthday collapse into one group, and superseding across them would leave one of them with no imaging at all. Two vendor ids are evidence the group is two people, and this module removes a file only because a newer scan of the same patient supersedes it.

Arguments

  • candidates: Every file in the run's selection, with its identity and acquisition time.
  • num_latest: How many files to keep per group. A ceiling, not a quota — a group holding fewer keeps them all.
  • min_scans: How many kept files a patient must hold to be worth inferring at all. 1 keeps every patient.

Returns The selected file paths, sorted.

Raises

  • UnsatisfiableScopeError: If the arguments cannot select anything (see _validate).

Classes​

ReduceCandidate​

class ReduceCandidate(    file_path: str,    bitfount_patient_id: str | None,    scan_datetime: datetime | None,    group_key: tuple[str, ...] = (),    patient_key: str | None = None,):

One file, with what the scan rows say about it.

Arguments

  • file_path: The file, as the inventory and every cache table key it.
  • bitfount_patient_id: The derived patient id, or None when the file's scan rows carry no usable name and date of birth.
  • scan_datetime: The file's most recent acquisition time, or None.
  • group_key: The configured grouping values beside the patient — the file's laterality, protocol, or both. Empty when the reduction groups by patient alone.
  • patient_key: The source's own patient identifier (a DICOM Patient ID, or the vendor's key), or None when the header stated none. Not the grouping key — it is not shared across vendors for one patient — but a finer distinction than bitfount_patient_id, which drops middle names and so collapses twins. Two of these inside one group are evidence the group holds two people.

Variables​

  • static bitfount_patient_id : str | None
  • static file_path : str
  • static group_key : tuple[str, ...]
  • static patient_key : str | None