task
Prefect task for the scan_reduce step.
The scope step of the v9 background DAG: it runs once, before waves are
planned, and narrows the run to the scans worth paying to infer. Every later
step — and model_inference above all, which is scoped by file name rather
than by a filenames input — sees the reduced selection.
Reads the stored scan_metadata rows rather than the datasource. The scan
runtime has already parsed every header into scan_datetime, laterality and
protocol, so deciding which files to skip costs a table scan rather than a
second pass over the files we are trying not to open.
Module
Functions
scan_reduce_task
def scan_reduce_task( cache: CacheProtocol, config: ScanReduceConfig, filenames: list[str], task_hash: str | None = None, run_id: str | None = None,) ‑> ScanReduceResult:Reduce filenames to the scans worth inferring.
Arguments
cache: The pod's background cache, holding the scan rows.config: The reduction to apply.filenames: The run's current selection, wired from$file_metadata.cache.task_hash: Provenance only, for the telemetry event.run_id: Provenance only, for the telemetry event.
Returns
ScanReduceResult with the selected filenames, sorted.