Skip to main content

timestamps

Reading and normalising the scan_metadata.scan_datetime column.

The column holds a scan's acquisition time, and it is the only clinical recency signal the cache has: file_metadata.modified_at is when the bytes last changed on this disk, which a bulk copy of an archive resets wholesale. Anything that orders scans by when the patient was seen has to read this column, so it has to be comparable.

It was not. The runtime's shim took the first of three vendor keys and stringified it with no parse and no format check (bitfount.runtimes.scan_metadata.functions.normalize_to_scan_metadata), so one column held at least three mutually incomparable spellings:

  • a private-eye value stringified from a datetime — 2026-09-15 12:00:00, space-separated rather than ISO's T;
  • DICOM Acquisition DateTime, a DT — 20260915120000.123456, optionally suffixed with a ±HHMM offset, and valid at any truncation from YYYY up;
  • DICOM Study Date, a DA — 20260915, carrying no time at all.

Sorting those as text interleaves them, and comparing them as times needs the format detected first. This module is that detection, in one place, so the writer, the v2 migration and every reader agree.

Two halves, and both are needed. normalise_scan_datetime is what the writer and the migration store, so the column converges on ISO-8601 UTC and can be compared as text. parse_scan_datetime is what a reader uses, because a cache written before the migration ran — or by an older SDK — still holds the old spellings, and because scan_metadata re-parses a file only when its content hash changes, so nothing rewrites an unchanged file's row on its own.

Lives under cache/types rather than beside the runtime that writes the column because the migration needs it and the cache layer must not import the runtimes layer. It is a plain module, not a vN package, so the record registry's version discovery skips it.

Module​

Functions​

normalise_scan_datetime​

def normalise_scan_datetime(value: object) ‑> str | None:

Render a scan_datetime cell as the one spelling the column stores.

ISO-8601 in UTC, which datetime.isoformat() emits with either zero or exactly six fractional digits and always a +00:00 suffix — so the column sorts lexicographically in timestamp order, which is what lets a recency query be a plain ORDER BY.

Idempotent: normalising an already-normalised value returns it unchanged, so re-running the v2 migration over migrated rows is a no-op.

Arguments

  • value: The cell, in any spelling parse_scan_datetime accepts.

Returns The ISO-8601 UTC string, or None when the value is unreadable — which is stored as NULL rather than as the junk it came from.

parse_scan_datetime​

def parse_scan_datetime(value: object) ‑> datetime.datetime | None:

Read a scan_datetime cell as an instant, whatever spelling it holds.

Recognises ISO-8601 (with or without an offset, T- or space-separated, date-only), DICOM DT at any truncation, and DICOM DA. A value carrying no time of day resolves to the start of its day — the conservative end, since a scan is never earlier than that.

Never raises, and never guesses: a value it cannot read is None, so the caller falls back to another key rather than sorting a corrupted header to one end of the order and leaving it there.

Arguments

  • value: The cell, which may already be a datetime or date.

Returns The instant in UTC, or None when the value is absent or unreadable.