functions
Choosing which of a run's files are worth paying to infer.
Pure selection, with no cache and no datasource: the task builds
ReduceCandidates from the stored scan rows and this decides which survive.
Keeping it free of both is what lets the interesting cases — ties, undated
scans, a patient who cannot qualify — be tested without a database.
The one rule the whole module keeps: a file is removed only because a newer
scan of the same patient supersedes it, or because its patient cannot reach
min_scans anyway. A file this cannot reason about is kept, never dropped.
Module
Functions
select_files
def select_files( candidates: Sequence[ReduceCandidate], *, num_latest: int, min_scans: int = 1,) ‑> list[str]:Return the files worth inferring, sorted.
Keeps the newest num_latest files in each (patient, *group_key) group,
then drops any patient left holding fewer than min_scans of them. A
candidate with no patient id is passed through untouched: it belongs to no
group, so nothing can supersede it and no scan count applies to it.
A group whose members state two different patient_keys is not reduced at
all. bitfount_patient_id drops middle names, prefixes and suffixes, so
twins sharing a birthday collapse into one group, and superseding across
them would leave one of them with no imaging at all. Two vendor ids are
evidence the group is two people, and this module removes a file only
because a newer scan of the same patient supersedes it.
Arguments
candidates: Every file in the run's selection, with its identity and acquisition time.num_latest: How many files to keep per group. A ceiling, not a quota — a group holding fewer keeps them all.min_scans: How many kept files a patient must hold to be worth inferring at all.1keeps every patient.
Returns The selected file paths, sorted.
Raises
UnsatisfiableScopeError: If the arguments cannot select anything (see_validate).
Classes
ReduceCandidate
class ReduceCandidate( file_path: str, bitfount_patient_id: str | None, scan_datetime: datetime | None, group_key: tuple[str, ...] = (), patient_key: str | None = None,):One file, with what the scan rows say about it.
Arguments
file_path: The file, as the inventory and every cache table key it.bitfount_patient_id: The derived patient id, orNonewhen the file's scan rows carry no usable name and date of birth.scan_datetime: The file's most recent acquisition time, orNone.group_key: The configured grouping values beside the patient — the file's laterality, protocol, or both. Empty when the reduction groups by patient alone.patient_key: The source's own patient identifier (a DICOMPatient ID, or the vendor's key), orNonewhen the header stated none. Not the grouping key — it is not shared across vendors for one patient — but a finer distinction thanbitfount_patient_id, which drops middle names and so collapses twins. Two of these inside one group are evidence the group holds two people.
Variables
- static
bitfount_patient_id : str | None
- static
file_path : str
- static
group_key : tuple[str, ...]
- static
patient_key : str | None
- static
scan_datetime : datetime.datetime | None