Skip to main content

indexing

Per-run counters for the background metadata runtimes.

A cold indexing pass takes hours, and the interesting question is always which part of it the time went to: walking, computing content keys, parsing vendor headers, or writing chunks to the cache. IndexingStats accumulates those figures for one run and emits them as telemetry — one IndexingProgressEvent per persisted chunk, and one IndexingSummaryEvent when the run ends.

Classes​

IndexingStats​

class IndexingStats(    run_type: str,    task_hash: str,    dataset_identifier: str | None = None,    files_seen: int = 0,    files_keyed: int = 0,    files_reused: int = 0,    files_unchanged: int = 0,    bytes_read: int = 0,    rows_written: int = 0,    key_seconds: float = 0.0,    parse_seconds: float = 0.0,    flush_seconds: float = 0.0,    io_concurrency: int | None = None,    transient_failures: int = 0,    root_probes: int = 0,    root_probe_seconds: float = 0.0,):

Mutable counters for one metadata runtime run.

The run's flow body threads this into the tasks that do the work; each reports what it did (record_key, record_reuse, record_unchanged, record_parse, record_flush) and emit_summary closes the run out. Every emission is swallowed on failure — telemetry must not break an indexing pass.

Variables​

  • static bytes_read : int
  • static dataset_identifier : str | None
  • static files_keyed : int
  • static files_reused : int
  • static files_seen : int
  • static files_unchanged : int
  • static flush_seconds : float
  • static io_concurrency : int | None
  • static key_seconds : float
  • static parse_seconds : float
  • static root_probe_seconds : float
  • static root_probes : int
  • static rows_written : int
  • static run_type : str
  • static task_hash : str
  • static transient_failures : int
  • elapsed_seconds : float - Wall-clock seconds since these counters were created.
  • files_per_second : float - Files seen per second so far, or 0.0 before any time has passed.
  • process_read_bytes : int | None - Bytes this process has read since these counters were created.

    Unlike bytes_read, which sums the sizes of the files handled, this is what the OS counted. It is process-wide, so runs sharing a process mix, and on Windows it includes network I/O. None where the OS has no counter (macOS).

Methods​


emit_summary​

def emit_summary(self, status: str, error: BaseException | None = None) ‑> None:

Emit the run's totals.

Sent at ERROR level for a failed run, carrying error's type and stack frames, so a failed pass does not read as a healthy one.

Arguments

  • status: The run's bookkeeping status ("success", "skipped", "partial" or "error").
  • error: The exception that failed the run, if any.

record_flush​

def record_flush(self, *, rows: int, seconds: float) ‑> None:

Record one persisted chunk and emit a progress event.

Arguments

  • rows: Rows the flush wrote.
  • seconds: Time the flush took.

record_key​

def record_key(self, *, size_bytes: int, seconds: float) ‑> None:

Record a file whose content key was computed.

Arguments

  • size_bytes: Bytes read to compute the key.
  • seconds: Time the key computation took.

record_parse​

def record_parse(self, seconds: float, *, size_bytes: int = 0) ‑> None:

Record one file whose vendor headers were parsed.

file_metadata counts a file via record_key / record_reuse and scan_metadata counts it here, so files_seen is one file per run either way.

Arguments

  • seconds: Time the parse took.
  • size_bytes: Size of the parsed file, counted towards bytes_read — the parse reads most of the file in practice.

record_reuse​

def record_reuse(self) ‑> None:

Record a file whose stored content key was reused unchanged.

record_unchanged​

def record_unchanged(self) ‑> None:

Record a file skipped because its stored row still describes it.

No content key is computed and no row is written for it.

record_unkeyed​

def record_unkeyed(self) ‑> None:

Record a file indexed without computing a content key.

Hashing is optional (settings.file_hashing_enabled) and off by default. With it off no key is computed and no bytes are read, so the file counts as seen but as neither keyed nor reused — counting it as keyed would report the whole dataset as bytes read by a pass that read none of it.