Skip to main content

task

Prefect task for the scan_reduce step.

The scope step of the v9 background DAG: it runs once, before waves are planned, and narrows the run to the scans worth paying to infer. Every later step — and model_inference above all, which is scoped by file name rather than by a filenames input — sees the reduced selection.

Reads the stored scan_metadata rows rather than the datasource. The scan runtime has already parsed every header into scan_datetime, laterality and protocol, so deciding which files to skip costs a table scan rather than a second pass over the files we are trying not to open.

Module​

Functions​

scan_reduce_task​

def scan_reduce_task(    cache: CacheProtocol,    config: ScanReduceConfig,    filenames: list[str],    task_hash: str | None = None,    run_id: str | None = None,) ‑> ScanReduceResult:

Reduce filenames to the scans worth inferring.

Arguments

  • cache: The pod's background cache, holding the scan rows.
  • config: The reduction to apply.
  • filenames: The run's current selection, wired from $file_metadata.cache.
  • task_hash: Provenance only, for the telemetry event.
  • run_id: Provenance only, for the telemetry event.

Returns ScanReduceResult with the selected filenames, sorted.