Skip to main content

child_datasource

Building the one datasource a background DAG child process needs.

A background DAG run executes in a subprocess that holds no pod, so it has to materialise its datasource itself. It builds only the datasource its run is for, which is the whole point of this module rather than reusing setup_pod_from_config: that constructs every datasource in the config (runners/pod_runner.py), and a pod's config routinely names datasets that are offline, retired, or on a network share that is no longer mounted. Touching one of those to run a DAG against a different one is how a child hangs before it starts.

Pod-first. The pod serves the configuration it is running the datasource with over its child-config endpoint, so the child builds exactly what an in-process run would have used — including a datasource the Hub holds no copy of, which is every datasource of a pod whose YAML or desktop app is authoritative. DatasourceConfigReloader keeps that copy reconciled with the Hub, so reading the Hub as well would only add a request per run. The Hub is consulted only when the pod served nothing, as a pod older than that field does.

Module​

Functions​

build_child_datasource_manager​

def build_child_datasource_manager(    hub: BitfountHub, pod_name: str, pod_key_path: Path,) ‑> DatasourceManager:

Build the DatasourceManager a child assembles its one datasource with.

The same pair of objects the pod builds for itself, from the same inputs: the pod's RSA keys off disk and the access manager's public key from the Hub. Rebuilt rather than shipped because neither a key object nor a Hub client survives a JSON deployment parameter.

pod_dp is None and no secrets are passed: both are only consulted when a schema is generated, and a child loads schemas (update_schema=False) rather than generating them — it must never overwrite the schema the pod published.

Arguments

  • hub: This child's authenticated Hub client.
  • pod_name: The pod the run belongs to.
  • pod_key_path: Path to the pod's pod_rsa.pem; its directory is the pod's key directory.

Returns A manager wired to hub, ready to build one datasource.

build_datasource_container​

def build_datasource_container(    config: DatasourceConfig, manager: DatasourceManager,) ‑> DatasourceContainer:

Materialise the single datasource config describes, with its schema.

Goes through DatasourceManager rather than reimplementing the assembly so the child builds a datasource exactly as the pod does — the data-config migration, the data-cache wiring and the predefined-schema handling are all behaviour a DAG depends on, and a second implementation of them would drift.

The schema is loaded, never regenerated: the caller's manager is built with update_schema off, so a child cannot overwrite the schema the pod published.

Arguments

  • config: The datasource's configuration.
  • manager: A DatasourceManager built against this child's Hub client.

Returns The DatasourceContainer for this one datasource.

fetch_hub_datasource_config​

def fetch_hub_datasource_config(    hub: BitfountHub, username: str, datasource_name: str,) ‑> DatasourceConfig | None:

Return the Hub's configuration for one datasource.

One request for the one datasource, rather than get_pod_datasets_configurations, which fetches the list and then a configuration per dataset — N+1 calls to end up discarding all but one.

Arguments

  • hub: An authenticated Hub client.
  • username: The pod's owner.
  • datasource_name: Name of the datasource within the pod.

Returns The configuration, or None when the Hub holds none, when the request fails, and when what comes back cannot be read.

resolve_datasource_config​

def resolve_datasource_config(    hub: BitfountHub,    username: str,    datasource_name: str,    pod_config: DatasourceConfig | None = None,) ‑> DatasourceConfig:

Return the configuration to build datasource_name from.

Arguments

  • hub: An authenticated Hub client.
  • username: The pod's owner.
  • datasource_name: Name of the datasource within the pod.
  • pod_config: The configuration the pod served for this run, used in preference to the Hub's. Ignored when it names another datasource.

Returns The datasource's configuration.

Raises

  • DatasourceNotFoundError: If neither the pod nor the Hub has it.

Classes​

DatasourceNotFoundError​

class DatasourceNotFoundError(*args, **kwargs):

Raised when the named datasource has no configuration anywhere.

A ValueError for the same reason runtimes.deployments.ConfigError is: it always means the run was asked for a datasource that does not exist, so the flow run must fail loudly rather than complete having done nothing.