migrations
Migration for scan_metadata v2 -> v3: drop rows whose timestamp may be wrong.
A data migration, not a schema one — v3's ORM is v2's, which is v1's.
parse_scan_datetime used to try datetime.fromisoformat before _parse_dicom
and fall back only when the ISO attempt raised. It does not raise on a DICOM DT
carrying a UTC offset: fromisoformat accepts any single character as the
date/time separator, so in 20260915120000+0000 the 1 of the hour is taken as
the separator and every following digit shifts left by one, giving a valid but
wrong instant with no exception to notice. A truncated DT fails differently —
in 20260915+0100 the + itself becomes the separator and the offset is
discarded — and a DT with no offset always raised, so it was never affected.
Both the writer and v2's migration store what that parse returned, so a wrong instant is now sitting in the column spelled exactly like a right one.
Why this hop deletes instead of repairing. The original vendor string is
gone: v2 rewrote it, has no downgrade, and is idempotent on its own output. A
correct value can only come from re-reading the source file's header, which a
migration cannot do — it has the database and nothing else. Deleting the
file's rows is what asks the runtime to do that re-read: _already_extracted
treats a file with no stored rows as never extracted, which is the state the
system already produces for a file it has not seen. The alternative — keeping
the rows, clearing source_file_hash and backdating processed_at — invents a
state nothing else in the codebase writes, and it would not even be safer in
the meantime: scan_reduce sorts a null scan_datetime as _OLDEST, so a
kept-but-cleared row loses to its siblings rather than being treated as unknown.
Which rows are suspect, and why the rest are provably safe. The mis-parse
consumed one digit and re-read the remainder as HH, MM and a trailing digit
it discarded, so every value it produced has a zero second and a zero
microsecond — verified exhaustively across the hour, minute, second and offset
ranges, at every DT truncation. A stored value carrying a non-zero second or a
fractional second therefore cannot have come from it and is kept. The converse
does not hold: a scan genuinely acquired on the minute is re-read too, along
with every Study Date-derived midnight. That is the cost of the original
string no longer existing, and it errs towards re-reading a file whose
timestamp was fine rather than keeping one that is wrong.
Why the whole file's rows go, not just the suspect one.
get_stored_extraction_state reads whichever row for a file comes back first
and takes its hash and extraction time as the file's, so leaving any row behind
would make the skip decision depend on row order. The runtime rewrites a
file's rows as a set in any case — delete_scans_for_file before every
re-extraction — so a partial deletion has no meaning to it.
No downgrade. There is nothing to restore: the values this hop discards
were unrecoverable before it ran, and the rows rebuild themselves from the
source files on the next indexing pass. A cache rolled back to an older SDK
sees files it has not extracted yet, which is a state it already handles.
Module
Functions
upgrade
def upgrade(op: Operations) ‑> None:Delete the rows of every file holding a timestamp that may be wrong.
Idempotent: the rows it would delete are gone after the first pass, and a file it left alone holds nothing suspect, so replaying the hop by hand is safe.
Arguments
op: The bound Alembic operations handle; its connection is the one the caller's per-hop transaction runs in, so the deletion commits or rolls back with the version bump.