Skip to main content

gpu

GPU usage telemetry for model_inference.

GpuUsageSampler probes once for a GPU, samples its compute and memory utilisation on a background thread while the step works, and summarises the window as a GpuUsageSummary. emit_gpu_usage_event sends that summary as one GpuUsageEvent, so the Datadog volume is one log event per step invocation regardless of how long the step runs.

Utilisation and device memory come from NVML (nvidia-ml-py), which reads the NVIDIA driver directly and so works even when torch is a CPU-only build. NVML is not available on macOS; there, presence and memory come from torch's MPS backend and utilisation is left empty.

The sampler never initialises CUDA itself. Every torch.cuda memory read is guarded by torch.cuda.is_initialized(), since creating a context costs hundreds of megabytes of device memory in a run that would otherwise have stayed on the CPU.

DAG steps that sample run serially (every shipped template sets parallel=False), so device-wide NVML figures and process-wide torch peak statistics are attributed to the one step that is sampling.

Module​

Functions​

emit_gpu_usage_event​

def emit_gpu_usage_event(    sampler: GpuUsageSampler,    *,    step_name: str,    model_ref: str | None = None,    model_class: str | None = None,    task_hash: str | None = None,    run_id: str | None = None,    exc: BaseException | None = None,) ‑> None:

Send a sampler's summary as one GpuUsageEvent.

Nothing is sent for a disabled sampler. Never raises.

Arguments

  • sampler: The sampler to report; stopped first if still running.
  • step_name: The step that was sampled.
  • model_ref: Which model ran.
  • model_class: Class name of the loaded model.
  • task_hash: The step's partition.
  • run_id: The flow run.
  • exc: The exception that ended the step, or None if it succeeded.

Classes​

GpuUsageSampler​

class GpuUsageSampler(interval_seconds: float = 1.0):

Samples GPU utilisation over a window of work.

Use as a context manager, or call start and stop directly when the window does not match a block; both are idempotent. Never raises: every probe and sample failure is logged at debug level and leaves the affected fields empty.

A no-op when settings.enable_gpu_telemetry is False.

Arguments

  • interval_seconds: Time between samples.

Variables​

  • duration_seconds : float | None - Length of the sampled window, once stopped.
  • enabled : bool - Whether this sampler is collecting, as decided at start.

Methods​


start​

def start(self) ‑> None:

Probe for a GPU and begin sampling.

stop​

def stop(self) ‑> None:

Stop sampling and read torch's own figures for the window.

summary​

def summary(self) ‑> GpuUsageSummary:

Summarise what was observed so far.

Returns The summary; all fields empty when the sampler is disabled.

GpuUsageSummary​

class GpuUsageSummary(    gpu_present: bool | None = None,    gpu_count: int | None = None,    gpu_name: str | None = None,    gpu_memory_total_bytes: int | None = None,    nvml_driver_version: str | None = None,    gpu_probe_error: str | None = None,    torch_cuda_build: str | None = None,    torch_cuda_available: bool | None = None,    torch_device: str | None = None,    torch_used_gpu: bool | None = None,    torch_gpu_memory_peak_bytes: int | None = None,    gpu_utilisation_mean_pct: float | None = None,    gpu_utilisation_max_pct: float | None = None,    gpu_memory_used_mean_bytes: float | None = None,    gpu_memory_used_max_bytes: int | None = None,    sample_count: int = 0,):

What a GpuUsageSampler observed over its window.

Field meanings match the same-named fields on GpuUsageEvent. A field the sampler could not determine is None.

Variables​

  • static gpu_count : int | None
  • static gpu_memory_total_bytes : int | None
  • static gpu_memory_used_max_bytes : int | None
  • static gpu_memory_used_mean_bytes : float | None
  • static gpu_name : str | None
  • static gpu_present : bool | None
  • static gpu_probe_error : str | None
  • static gpu_utilisation_max_pct : float | None
  • static gpu_utilisation_mean_pct : float | None
  • static nvml_driver_version : str | None
  • static sample_count : int
  • static torch_cuda_available : bool | None
  • static torch_cuda_build : str | None
  • static torch_device : str | None
  • static torch_gpu_memory_peak_bytes : int | None
  • static torch_used_gpu : bool | None