Skip to main content

date_order

How a two-field date of birth is read when deriving a patient identity.

04/03/1952 is 3 April read month-first and 4 March read day-first, and the two derive different BitfountPatientIDs for one patient — which splits them across every cache keyed on the value. Something has to settle the order, and it has to settle it the same way on every run, because those IDs are cache keys joined across runs and appear in published exports.

This lives here, rather than beside either of its two callers, because both of them need it: bitfount.steps.patient_identity is the entry point every name-and-DOB route goes through, and bitfount.federated.algorithms.ophthalmology.dataframe_generation_extensions holds the derivation itself and imports from neither direction cleanly — the first already imports the second, so putting this in the first would be circular.

Module​

Functions​

check_date_order​

def check_date_order(values: Iterable[Any], *, source: str) ‑> None:

Report a column whose own values disprove the resolved field order.

The column does not get to change the order — that is what would make the order a fact about a batch — but a file written the other way round keys every ambiguous patient in it to the wrong date, so it must not be silent.

Arguments

  • values: The column's cells, as supplied.
  • source: What to name in the log line. Never a cell value: a date of birth is PHI.

infer_date_order​

def infer_date_order(values: Iterable[Any]) ‑> bool | None:

Read the field order off a column of dates of birth.

A value settles its own order only when the two readings agree — which means the parser had no choice, because one of the fields cannot be a month. Which field it put the day in is then the answer: 25/12/1950 resolves to the 25th, and 25 is the first field, so the column is day-first. 12/25/1950 resolves to the same day from the second field, so that column is month-first. Everything else — an ambiguous two-field date, an ISO date, a DICOM DA, a blank — carries no evidence either way.

Arguments

  • values: The column's cells, as supplied.

Returns True for day-first, False for month-first, or None when the column proves neither. A column proving both also returns None: two orders in one export is a broken source, not a verdict, and guessing between them would key some of its patients wrongly either way.

reset_date_order​

def reset_date_order() ‑> None:

Forget the resolved field order, so the next call derives it again.

Only tests need this. The order is deliberately resolved once for the process: it decides BitfountPatientIDs, which are cache keys joined across runs, so it must not vary with what a given batch happens to hold.

resolve_date_order​

def resolve_date_order(values: Iterable[Any] | None = None) ‑> bool:

The field order every date of birth is read with, resolved once.

Resolved from the best evidence available on the first call that needs it, and then reused unchanged for the life of the process. Deriving it afresh per column made it a fact about a batch: two exports with different contents would resolve differently and key one patient two ways, and these IDs are cache keys joined across runs and published in exports.

In order:

  1. patient_id_date_order, when set to day-first or month-first.
  2. values, when they prove their own order — a field above 12 can only be one thing. A file says what it is; the host only says where it sits, so a US-formatted export landing on a UK pod is read as written.
  3. The pod's LC_TIME, when one is genuinely set, and otherwise its timezone, through CLDR.
  4. Month-first, pd.to_datetime's own default, with a warning.

Because the answer is kept, the first call carrying decisive data settles it and a later column disagreeing does not move it. check_date_order reports that disagreement rather than acting on it.

Arguments

  • values: A column of dates of birth, when the caller has one. A caller holding a single row has nothing to infer from and omits it.

Returns True to read day-first, False to read month-first.