Skip to main content

serving

Serving the background DAG deployment from inside the pod process.

A background DAG run has to execute somewhere that is not the pod's event loop (see flows.dag.deployment), and the machinery that does that is a Prefect Runner: it polls for scheduled runs of the deployments it knows about and launches each in its own subprocess.

One deployment per project, created lazily. A pod does not know at start-up which projects it will be asked to run — a dataset-project link arrives as a message — so deployments are registered as work appears, which Runner.add_deployment supports against an already-running runner. Per project rather than one for the pod because that is where the concurrency limit has to live: two datasources linked to the same project are one project's work and queue behind each other, while two projects do not.

The runner runs on a daemon thread with its own event loop, mirroring run_pod.py's serve() threads and the pod control server. Nothing here may touch the pod's loop, and nothing on the pod's loop may block on the runner beyond the bounded waits below: the whole point of the arrangement is that a busy DAG cannot stop the pod's heartbeat.