timestamps
Reading and normalising the scan_metadata.scan_datetime column.
The column holds a scan's acquisition time, and it is the only clinical
recency signal the cache has: file_metadata.modified_at is when the bytes
last changed on this disk, which a bulk copy of an archive resets wholesale.
Anything that orders scans by when the patient was seen has to read this
column, so it has to be comparable.
It was not. The runtime's shim took the first of three vendor keys and
stringified it with no parse and no format check
(bitfount.runtimes.scan_metadata.functions.normalize_to_scan_metadata), so one
column held at least three mutually incomparable spellings:
- a private-eye value stringified from a
datetime—2026-09-15 12:00:00, space-separated rather than ISO'sT; - DICOM
Acquisition DateTime, a DT —20260915120000.123456, optionally suffixed with a±HHMMoffset, and valid at any truncation fromYYYYup; - DICOM
Study Date, a DA —20260915, carrying no time at all.
Sorting those as text interleaves them, and comparing them as times needs the format detected first. This module is that detection, in one place, so the writer, the v2 migration and every reader agree.
Two halves, and both are needed. normalise_scan_datetime is what the writer
and the migration store, so the column converges on ISO-8601 UTC and can be
compared as text. parse_scan_datetime is what a reader uses, because a cache
written before the migration ran — or by an older SDK — still holds the old
spellings, and because scan_metadata re-parses a file only when its content
hash changes, so nothing rewrites an unchanged file's row on its own.
Lives under cache/types rather than beside the runtime that writes the column
because the migration needs it and the cache layer must not import the runtimes
layer. It is a plain module, not a vN package, so the record registry's
version discovery skips it.
Module
Functions
normalise_scan_datetime
def normalise_scan_datetime(value: object) ‑> str | None:Render a scan_datetime cell as the one spelling the column stores.
ISO-8601 in UTC, which datetime.isoformat() emits with either zero or
exactly six fractional digits and always a +00:00 suffix — so the column
sorts lexicographically in timestamp order, which is what lets a recency
query be a plain ORDER BY.
Idempotent: normalising an already-normalised value returns it unchanged, so re-running the v2 migration over migrated rows is a no-op.
Arguments
value: The cell, in any spellingparse_scan_datetimeaccepts.
Returns
The ISO-8601 UTC string, or None when the value is unreadable — which
is stored as NULL rather than as the junk it came from.
parse_scan_datetime
def parse_scan_datetime(value: object) ‑> datetime.datetime | None:Read a scan_datetime cell as an instant, whatever spelling it holds.
Recognises ISO-8601 (with or without an offset, T- or space-separated,
date-only), DICOM DT at any truncation, and DICOM DA. A value carrying no
time of day resolves to the start of its day — the conservative end, since
a scan is never earlier than that.
Never raises, and never guesses: a value it cannot read is None, so the
caller falls back to another key rather than sorting a corrupted header to
one end of the order and leaving it there.
Arguments
value: The cell, which may already be adatetimeordate.
Returns
The instant in UTC, or None when the value is absent or unreadable.