Skip to main content

migrations

Migration for scan_metadata v2 -> v3: drop rows whose timestamp may be wrong.

A data migration, not a schema one — v3's ORM is v2's, which is v1's.

parse_scan_datetime used to try datetime.fromisoformat before _parse_dicom and fall back only when the ISO attempt raised. It does not raise on a DICOM DT carrying a UTC offset: fromisoformat accepts any single character as the date/time separator, so in 20260915120000+0000 the 1 of the hour is taken as the separator and every following digit shifts left by one, giving a valid but wrong instant with no exception to notice. A truncated DT fails differently — in 20260915+0100 the + itself becomes the separator and the offset is discarded — and a DT with no offset always raised, so it was never affected.

Both the writer and v2's migration store what that parse returned, so a wrong instant is now sitting in the column spelled exactly like a right one.

Why this hop deletes instead of repairing. The original vendor string is gone: v2 rewrote it, has no downgrade, and is idempotent on its own output. A correct value can only come from re-reading the source file's header, which a migration cannot do — it has the database and nothing else. Deleting the file's rows is what asks the runtime to do that re-read: _already_extracted treats a file with no stored rows as never extracted, which is the state the system already produces for a file it has not seen. The alternative — keeping the rows, clearing source_file_hash and backdating processed_at — invents a state nothing else in the codebase writes, and it would not even be safer in the meantime: scan_reduce sorts a null scan_datetime as _OLDEST, so a kept-but-cleared row loses to its siblings rather than being treated as unknown.

Which rows are suspect, and why the rest are provably safe. The mis-parse consumed one digit and re-read the remainder as HH, MM and a trailing digit it discarded, so every value it produced has a zero second and a zero microsecond — verified exhaustively across the hour, minute, second and offset ranges, at every DT truncation. A stored value carrying a non-zero second or a fractional second therefore cannot have come from it and is kept. The converse does not hold: a scan genuinely acquired on the minute is re-read too, along with every Study Date-derived midnight. That is the cost of the original string no longer existing, and it errs towards re-reading a file whose timestamp was fine rather than keeping one that is wrong.

Why the whole file's rows go, not just the suspect one. get_stored_extraction_state reads whichever row for a file comes back first and takes its hash and extraction time as the file's, so leaving any row behind would make the skip decision depend on row order. The runtime rewrites a file's rows as a set in any case — delete_scans_for_file before every re-extraction — so a partial deletion has no meaning to it.

No downgrade. There is nothing to restore: the values this hop discards were unrecoverable before it ran, and the rows rebuild themselves from the source files on the next indexing pass. A cache rolled back to an older SDK sees files it has not extracted yet, which is a state it already handles.

Module​

Functions​

upgrade​

def upgrade(op: Operations) ‑> None:

Delete the rows of every file holding a timestamp that may be wrong.

Idempotent: the rows it would delete are gone after the first pass, and a file it left alone holds nothing suspect, so replaying the hop by hand is safe.

Arguments

  • op: The bound Alembic operations handle; its connection is the one the caller's per-hop transaction runs in, so the deletion commits or rolls back with the version bump.