Skip to main content

indexing

When a reader of the file inventory has to wait for one to exist.

Every background DAG reads the file_metadata inventory a file_metadata deployment run builds, and asking "is that run in flight?" is not the question a reader wants answered. The question is whether an inventory exists to read at all, and those two differ for the case that happens nightly.

A reader that starts mid-index banks whatever it found and finishes COMPLETED — under-coverage wearing the costume of success — but only while the inventory is being built from nothing. A re-index leaves the previous one whole: bulk_insert_file_metadata merges rows in place rather than replacing them, and prune_missing_under_root runs only after a successful walk and deletes only rows whose file is gone from disk. So a reader overlapping a re-index sees every file the last walk found, at worst with stale metadata for the ones not yet revisited — and the files it cannot see are the ones that arrived since, which waiting would not have shown it either.

The distinction is load-bearing rather than academic: the orchestrator's metadata_refresh_cron defaults to 03:00 daily with metadata_staleness_days=1, so essentially every datasource re-indexes every night, three hours into a 12-hour rerun window. Deferring on that starves whichever lineages did not get in first.

A pass that stops because its datasource root became unreachable ends partial and skips the prune, so a partial re-index also leaves the previous inventory whole. A partial first index does not: its rows are what the cut-short pass reached, and no complete inventory has ever existed. Readers wait on it until a full pass completes, whether its follow-up is scheduled, running, or has failed — the rows being present says nothing about whether they are all of them.

Module​

Functions​

first_index_left_partial​

def first_index_left_partial(cache: CacheProtocol, task_hash: str) ‑> bool:

Whether the only inventory is what a cut-short first index reached.

first_index_incomplete, which scan's stand-down shares. Never raises; an unreadable cache reads as False.

Arguments

  • cache: The pod's background cache.
  • task_hash: The (pod, datasource) task hash.

is_first_index_in_flight​

async def is_first_index_in_flight(    client: PrefectClient, cache: CacheProtocol, task_hash: str,) ‑> bool:

Return whether task_hash is being indexed with nothing yet to read.

The predicate every background reader of the inventory asks before starting: True means there is no complete inventory and one is being built, so a reader would process a partial dataset and finish looking successful.

Arguments

  • client: An open PrefectClient.
  • cache: The pod's background cache, holding the inventory.
  • task_hash: The (pod, datasource) task hash whose inventory is read.

Returns True if a first index is in flight with nothing yet written, or a first index was cut short and no full pass has completed since, and the reader should wait.