Skip to main content

coverage

Sampling how much of a datasource is indexed, for the published pointer.

patient_eligibility flips trials_published_data_pointer over whatever the run computed, and a run launched while file_metadata is still walking computes over a fraction of the datasource. The published cohort is then correct but partial, and — without this — indistinguishable from a genuinely small one. Publishing a coverage figure beside it is what lets a dataset be linked to a project immediately rather than after an operator judges the walk far enough along.

Why the count is taken fresh rather than from the run's own inputs. One FileMetadataContext serves a whole run and memoises its selection on first access, so every step reading $file_metadata.cache sees the identical list. A numerator and denominator both drawn from it are equal by construction and the ratio is a constant 1.0. The inventory has grown underneath the run since that list was resolved — file_metadata commits bounded chunks as it walks — so re-reading it at publish time is the only way the figure carries information.

Why the in-flight flag is not optional. A ratio alone reads as complete whenever a run catches up with the inventory, which includes a walk paused between chunk commits, and a walk that has stalled or crashed. The flag is what separates "this is the whole datasource" from "the denominator is still moving".

This module lives here rather than under bitfount.steps because it needs the runtimes layer (the inventory selection rule and the Prefect liveness query), and no step imports that layer. Steps see only bitfount.steps.protocols.scan_coverage.ScanCoverageProbeProtocol.

Classes​

InventoryScanCoverageProbe​

class InventoryScanCoverageProbe(    cache: CacheProtocol,    task_hash: str,    datasource: BaseSource | None = None,    wave_plan: Any = None,):

Reports how much of one datasource the inventory currently holds.

Implements ScanCoverageProbeProtocol. Built once per run in bitfount.flows.dag.setup.build_background_run_context and injected as a runtime parameter, so the step that publishes never imports the runtimes layer.

Deliberately not a task_hash_resource: it is a plain runtime parameter, which keeps it out of the Merkle hash (bitfount.flows.dag.hashing) and so out of every partition key. Declaring it as a resource would re-partition every step that takes it and force a pod-wide recompute for a reporting figure.

Arguments

  • cache: The pod's background cache, holding the inventory.
  • task_hash: The (pod, datasource) task hash.
  • datasource: The datasource whose files are being counted.
  • wave_plan: The run's wave plan, or None when it is not waved.

Methods​


sample​

def sample(self) ‑> ScanCoverage:

Return the inventory size and indexing state as of now.

Never raises. This feeds a reporting figure written beside a published pointer; a probe that could fail the publish would trade a cohort for an annotation.

Returns A ScanCoverage reading, with None in either field that could not be determined.