child_datasource
Building the one datasource a background DAG child process needs.
A background DAG run executes in a subprocess that holds no pod, so it has to
materialise its datasource itself. It builds only the datasource its run
is for, which is the whole point of this module rather than reusing
setup_pod_from_config: that constructs every datasource in the config
(runners/pod_runner.py), and a pod's config routinely names datasets that are
offline, retired, or on a network share that is no longer mounted. Touching
one of those to run a DAG against a different one is how a child hangs before
it starts.
Pod-first. The pod serves the configuration it is running the datasource
with over its child-config endpoint, so the child builds exactly what an
in-process run would have used — including a datasource the Hub holds no copy
of, which is every datasource of a pod whose YAML or desktop app is
authoritative. DatasourceConfigReloader keeps that copy reconciled with the
Hub, so reading the Hub as well would only add a request per run. The Hub is
consulted only when the pod served nothing, as a pod older than that field
does.
Module
Functions
build_child_datasource_manager
def build_child_datasource_manager( hub: BitfountHub, pod_name: str, pod_key_path: Path,) ‑> DatasourceManager:Build the DatasourceManager a child assembles its one datasource with.
The same pair of objects the pod builds for itself, from the same inputs: the pod's RSA keys off disk and the access manager's public key from the Hub. Rebuilt rather than shipped because neither a key object nor a Hub client survives a JSON deployment parameter.
pod_dp is None and no secrets are passed: both are only consulted when
a schema is generated, and a child loads schemas (update_schema=False)
rather than generating them — it must never overwrite the schema the pod
published.
Arguments
hub: This child's authenticated Hub client.pod_name: The pod the run belongs to.pod_key_path: Path to the pod'spod_rsa.pem; its directory is the pod's key directory.
Returns A manager wired to hub, ready to build one datasource.
build_datasource_container
def build_datasource_container( config: DatasourceConfig, manager: DatasourceManager,) ‑> DatasourceContainer:Materialise the single datasource config describes, with its schema.
Goes through DatasourceManager rather than reimplementing the assembly so
the child builds a datasource exactly as the pod does — the data-config
migration, the data-cache wiring and the predefined-schema handling are all
behaviour a DAG depends on, and a second implementation of them would drift.
The schema is loaded, never regenerated: the caller's manager is built with
update_schema off, so a child cannot overwrite the schema the pod
published.
Arguments
config: The datasource's configuration.manager: ADatasourceManagerbuilt against this child's Hub client.
Returns
The DatasourceContainer for this one datasource.
fetch_hub_datasource_config
def fetch_hub_datasource_config( hub: BitfountHub, username: str, datasource_name: str,) ‑> DatasourceConfig | None:Return the Hub's configuration for one datasource.
One request for the one datasource, rather than
get_pod_datasets_configurations, which fetches the list and then a
configuration per dataset — N+1 calls to end up discarding all but one.
Arguments
hub: An authenticated Hub client.username: The pod's owner.datasource_name: Name of the datasource within the pod.
Returns
The configuration, or None when the Hub holds none, when the request
fails, and when what comes back cannot be read.
resolve_datasource_config
def resolve_datasource_config( hub: BitfountHub, username: str, datasource_name: str, pod_config: DatasourceConfig | None = None,) ‑> DatasourceConfig:Return the configuration to build datasource_name from.
Arguments
hub: An authenticated Hub client.username: The pod's owner.datasource_name: Name of the datasource within the pod.pod_config: The configuration the pod served for this run, used in preference to the Hub's. Ignored when it names another datasource.
Returns The datasource's configuration.
Raises
DatasourceNotFoundError: If neither the pod nor the Hub has it.
Classes
DatasourceNotFoundError
class DatasourceNotFoundError(*args, **kwargs):Raised when the named datasource has no configuration anywhere.
A ValueError for the same reason runtimes.deployments.ConfigError is:
it always means the run was asked for a datasource that does not exist, so
the flow run must fail loudly rather than complete having done nothing.