indexing
When a reader of the file inventory has to wait for one to exist.
Every background DAG reads the file_metadata inventory a file_metadata
deployment run builds, and asking "is that run in flight?" is not the question a
reader wants answered. The question is whether an inventory exists to read at
all, and those two differ for the case that happens nightly.
A reader that starts mid-index banks whatever it found and finishes
COMPLETED — under-coverage wearing the costume of success — but only while
the inventory is being built from nothing. A re-index leaves the previous one
whole: bulk_insert_file_metadata merges rows in place rather than replacing
them, and prune_missing_under_root runs only after a successful walk and
deletes only rows whose file is gone from disk. So a reader overlapping a
re-index sees every file the last walk found, at worst with stale metadata for
the ones not yet revisited — and the files it cannot see are the ones that
arrived since, which waiting would not have shown it either.
The distinction is load-bearing rather than academic: the orchestrator's
metadata_refresh_cron defaults to 03:00 daily with metadata_staleness_days=1,
so essentially every datasource re-indexes every night, three hours into a
12-hour rerun window. Deferring on that starves whichever lineages did not get
in first.
A pass that stops because its datasource root became unreachable ends
partial and skips the prune, so a partial re-index also leaves the previous
inventory whole. A partial first index does not: its rows are what the
cut-short pass reached, and no complete inventory has ever existed. Readers
wait on it until a full pass completes, whether its follow-up is scheduled,
running, or has failed — the rows being present says nothing about whether
they are all of them.
Module
Functions
first_index_left_partial
def first_index_left_partial(cache: CacheProtocol, task_hash: str) ‑> bool:Whether the only inventory is what a cut-short first index reached.
first_index_incomplete, which scan's stand-down shares. Never raises; an
unreadable cache reads as False.
Arguments
cache: The pod's background cache.task_hash: The(pod, datasource)task hash.
is_first_index_in_flight
async def is_first_index_in_flight( client: PrefectClient, cache: CacheProtocol, task_hash: str,) ‑> bool:Return whether task_hash is being indexed with nothing yet to read.
The predicate every background reader of the inventory asks before starting:
True means there is no complete inventory and one is being built, so a
reader would process a partial dataset and finish looking successful.
Arguments
client: An openPrefectClient.cache: The pod's background cache, holding the inventory.task_hash: The(pod, datasource)task hash whose inventory is read.
Returns
True if a first index is in flight with nothing yet written, or a first
index was cut short and no full pass has completed since, and the reader
should wait.