date_order
How a two-field date of birth is read when deriving a patient identity.
04/03/1952 is 3 April read month-first and 4 March read day-first, and the
two derive different BitfountPatientIDs for one patient — which splits them
across every cache keyed on the value. Something has to settle the order, and
it has to settle it the same way on every run, because those IDs are cache
keys joined across runs and appear in published exports.
This lives here, rather than beside either of its two callers, because both of
them need it: bitfount.steps.patient_identity is the entry point every
name-and-DOB route goes through, and
bitfount.federated.algorithms.ophthalmology.dataframe_generation_extensions
holds the derivation itself and imports from neither direction cleanly — the
first already imports the second, so putting this in the first would be
circular.
Module
Functions
check_date_order
def check_date_order(values: Iterable[Any], *, source: str) ‑> None:Report a column whose own values disprove the resolved field order.
The column does not get to change the order — that is what would make the order a fact about a batch — but a file written the other way round keys every ambiguous patient in it to the wrong date, so it must not be silent.
Arguments
values: The column's cells, as supplied.source: What to name in the log line. Never a cell value: a date of birth is PHI.
infer_date_order
def infer_date_order(values: Iterable[Any]) ‑> bool | None:Read the field order off a column of dates of birth.
A value settles its own order only when the two readings agree — which
means the parser had no choice, because one of the fields cannot be a
month. Which field it put the day in is then the answer: 25/12/1950
resolves to the 25th, and 25 is the first field, so the column is
day-first. 12/25/1950 resolves to the same day from the second field,
so that column is month-first. Everything else — an ambiguous two-field
date, an ISO date, a DICOM DA, a blank — carries no evidence either way.
Arguments
values: The column's cells, as supplied.
Returns
True for day-first, False for month-first, or None when the
column proves neither. A column proving both also returns None:
two orders in one export is a broken source, not a verdict, and
guessing between them would key some of its patients wrongly either
way.
reset_date_order
def reset_date_order() ‑> None:Forget the resolved field order, so the next call derives it again.
Only tests need this. The order is deliberately resolved once for the
process: it decides BitfountPatientIDs, which are cache keys joined
across runs, so it must not vary with what a given batch happens to hold.
resolve_date_order
def resolve_date_order(values: Iterable[Any] | None = None) ‑> bool:The field order every date of birth is read with, resolved once.
Resolved from the best evidence available on the first call that needs it, and then reused unchanged for the life of the process. Deriving it afresh per column made it a fact about a batch: two exports with different contents would resolve differently and key one patient two ways, and these IDs are cache keys joined across runs and published in exports.
In order:
patient_id_date_order, when set today-firstormonth-first.- values, when they prove their own order — a field above 12 can only be one thing. A file says what it is; the host only says where it sits, so a US-formatted export landing on a UK pod is read as written.
- The pod's
LC_TIME, when one is genuinely set, and otherwise its timezone, through CLDR. - Month-first,
pd.to_datetime's own default, with a warning.
Because the answer is kept, the first call carrying decisive data settles
it and a later column disagreeing does not move it. check_date_order
reports that disagreement rather than acting on it.
Arguments
values: A column of dates of birth, when the caller has one. A caller holding a single row has nothing to infer from and omits it.
Returns
True to read day-first, False to read month-first.