gpu
GPU usage telemetry for model_inference.
GpuUsageSampler probes once for a GPU, samples its compute and memory
utilisation on a background thread while the step works, and summarises the
window as a GpuUsageSummary. emit_gpu_usage_event sends that summary as one
GpuUsageEvent, so the Datadog volume is one log event per step invocation
regardless of how long the step runs.
Utilisation and device memory come from NVML (nvidia-ml-py), which reads the
NVIDIA driver directly and so works even when torch is a CPU-only build. NVML
is not available on macOS; there, presence and memory come from torch's MPS
backend and utilisation is left empty.
The sampler never initialises CUDA itself. Every torch.cuda memory read is
guarded by torch.cuda.is_initialized(), since creating a context costs
hundreds of megabytes of device memory in a run that would otherwise have
stayed on the CPU.
DAG steps that sample run serially (every shipped template sets
parallel=False), so device-wide NVML figures and process-wide torch peak
statistics are attributed to the one step that is sampling.
Module
Functions
emit_gpu_usage_event
def emit_gpu_usage_event( sampler: GpuUsageSampler, *, step_name: str, model_ref: str | None = None, model_class: str | None = None, task_hash: str | None = None, run_id: str | None = None, exc: BaseException | None = None,) ‑> None:Send a sampler's summary as one GpuUsageEvent.
Nothing is sent for a disabled sampler. Never raises.
Arguments
sampler: The sampler to report; stopped first if still running.step_name: The step that was sampled.model_ref: Which model ran.model_class: Class name of the loaded model.task_hash: The step's partition.run_id: The flow run.exc: The exception that ended the step, orNoneif it succeeded.
Classes
GpuUsageSampler
class GpuUsageSampler(interval_seconds: float = 1.0):Samples GPU utilisation over a window of work.
Use as a context manager, or call start and stop directly when the
window does not match a block; both are idempotent. Never raises: every
probe and sample failure is logged at debug level and leaves the affected
fields empty.
A no-op when settings.enable_gpu_telemetry is False.
Arguments
interval_seconds: Time between samples.
Variables
duration_seconds : float | None- Length of the sampled window, once stopped.
enabled : bool- Whether this sampler is collecting, as decided atstart.
Methods
start
def start(self) ‑> None:Probe for a GPU and begin sampling.
stop
def stop(self) ‑> None:Stop sampling and read torch's own figures for the window.
summary
def summary(self) ‑> GpuUsageSummary:Summarise what was observed so far.
Returns The summary; all fields empty when the sampler is disabled.
GpuUsageSummary
class GpuUsageSummary( gpu_present: bool | None = None, gpu_count: int | None = None, gpu_name: str | None = None, gpu_memory_total_bytes: int | None = None, nvml_driver_version: str | None = None, gpu_probe_error: str | None = None, torch_cuda_build: str | None = None, torch_cuda_available: bool | None = None, torch_device: str | None = None, torch_used_gpu: bool | None = None, torch_gpu_memory_peak_bytes: int | None = None, gpu_utilisation_mean_pct: float | None = None, gpu_utilisation_max_pct: float | None = None, gpu_memory_used_mean_bytes: float | None = None, gpu_memory_used_max_bytes: int | None = None, sample_count: int = 0,):What a GpuUsageSampler observed over its window.
Field meanings match the same-named fields on GpuUsageEvent. A field the
sampler could not determine is None.
Variables
- static
gpu_count : int | None
- static
gpu_memory_total_bytes : int | None
- static
gpu_memory_used_max_bytes : int | None
- static
gpu_memory_used_mean_bytes : float | None
- static
gpu_name : str | None
- static
gpu_present : bool | None
- static
gpu_probe_error : str | None
- static
gpu_utilisation_max_pct : float | None
- static
gpu_utilisation_mean_pct : float | None
- static
nvml_driver_version : str | None
- static
sample_count : int
- static
torch_cuda_available : bool | None
- static
torch_cuda_build : str | None
- static
torch_device : str | None
- static
torch_gpu_memory_peak_bytes : int | None
- static
torch_used_gpu : bool | None