Skip to main content

prioritisation

Choose which datasources to index for file metadata, and in what order.

Pure: every input is passed in, so the pod, the nightly refresh and the headless run_pod path share one rule and it can be tested without a Hub or a cache.

A datasource is skipped only when every project the Hub reports it linked to is archived. Every other datasource is indexed, ordered by the pod's last health check (healthy, then unknown, then unhealthy). Within each band, datasources never indexed come first; then the Hub's configurationLastUpdated, newest first; then the least recently indexed. Health orders and never skips: an unreachable source fails fast when its turn comes, whereas a transient probe failure that skipped a datasource would leave it unindexed until the next kickoff.

A Hub input of None means that call failed. Its filter is then not applied, never read as "nothing linked" or "nothing recent": a Hub outage must not stop indexing.

Module​

Functions​

order_candidates​

def order_candidates(    candidates: Sequence[MetadataCandidate],) ‑> list[MetadataCandidate]:

Order candidates for indexing.

By health band; within it, never-indexed first, then newest config_last_updated, then least recently indexed. Stable, so candidates that tie keep their input order. A candidate with no config_last_updated sorts after those that have one.

Arguments

  • candidates: The candidates to order.

Returns A new list in submission order.

prioritise_datasources​

def prioritise_datasources(    names: Sequence[str],    hub_datasets: Sequence[_DatasetListEntryJSON] | None,    project_view: HubProjectView | None,    health: Iterable[Mapping[str, str]] | None,    last_indexed: Mapping[str, datetime] | None = None,    owner: str | None = None,) ‑> MetadataKickoffPlan:

Decide which of names to index, and in what order.

Arguments

  • names: The pod's configured datasource names, in config order. Ties keep this order.
  • hub_datasets: BitfountHub.get_pod_dataset_list(), or None if it failed.
  • project_view: Pod._hub_project_view(), or None if it failed.
  • health: The pod's last _collect_datasource_health() result, or None if no check has completed.
  • last_indexed: When each datasource's last file_metadata run completed, by name. A name absent from it was never indexed.
  • owner: The pod owner's username. A project link counts only when the linked dataset is this owner's: a project can link another user's dataset of the same name. None matches on name alone.

Returns The plan: the datasources to index in order, and those skipped.

Classes​

MetadataCandidate​

class MetadataCandidate(    name: str,    health: HealthStatus,    error_code: str | None,    config_last_updated: datetime | None,    archived_only: bool | None,    last_indexed_at: datetime | None = None,):

One datasource considered for indexing.

Arguments

  • name: The pod's datasource name.
  • health: Status from the pod's last accessibility check; unknown when no check has completed for it.
  • error_code: The AccessibilityErrorCode behind an unhealthy status.
  • config_last_updated: The Hub's configurationLastUpdated, or None when the Hub did not report one.
  • archived_only: Whether every project the datasource is linked to is archived. None when the Hub's project links are not known.
  • last_indexed_at: When its last file_metadata run completed, or None if none ever has.

Variables​

  • static archived_only : Optional[bool]
  • static error_code : str | None
  • static health : Literal['healthy', 'unknown', 'unhealthy']
  • static name : str

MetadataKickoffPlan​

class MetadataKickoffPlan(    ordered: list[MetadataCandidate],    skipped_archived_only: list[MetadataCandidate],    hub_filters_applied: list[str],    hub_filters_failed: list[str],):

The outcome of one indexing decision.

Arguments

  • ordered: Datasources to index, in submission order.
  • skipped_archived_only: Datasources left out because every project they are linked to is archived.
  • hub_filters_applied: Hub filters whose input was available.
  • hub_filters_failed: Hub filters whose call failed and were not applied.

Variables​

  • static hub_filters_applied : list[str]
  • static hub_filters_failed : list[str]
  • candidates : list[MetadataCandidate] - Every datasource considered, indexed or skipped.
  • ordered_names : list[str] - Names of the datasources to index, in submission order.
  • skipped_names : list[str] - Names of the datasources left out of indexing.

Methods​


metadata​

def metadata(    self,) ‑> dict[str, DatasourceMetadata]:

Cache records for every datasource considered, by name.

restricted_to​

def restricted_to(    self, names: Iterable[str],) ‑> MetadataKickoffPlan:

This plan limited to names, keeping its order.

For a reload, which indexes only the datasources it added.

Arguments

  • names: The datasource names to keep.

summary​

def summary(self, pod_name: str, trigger: KickoffTrigger) ‑> str:

A one-line description of this decision for the pod log.

Arguments

  • pod_name: Pod the decision was made for.
  • trigger: What caused the decision.

to_event​

def to_event(    self, pod_name: str, trigger: KickoffTrigger,) ‑> MetadataKickoffEvent:

The Datadog event describing this decision.

Arguments

  • pod_name: Pod the decision was made for.
  • trigger: What caused the decision.