indexing
Per-run counters for the background metadata runtimes.
A cold indexing pass takes hours, and the interesting question is always which
part of it the time went to: walking, computing content keys, parsing vendor
headers, or writing chunks to the cache. IndexingStats accumulates those
figures for one run and emits them as telemetry — one IndexingProgressEvent
per persisted chunk, and one IndexingSummaryEvent when the run ends.
Classes
IndexingStats
class IndexingStats( run_type: str, task_hash: str, dataset_identifier: str | None = None, files_seen: int = 0, files_keyed: int = 0, files_reused: int = 0, files_unchanged: int = 0, bytes_read: int = 0, rows_written: int = 0, key_seconds: float = 0.0, parse_seconds: float = 0.0, flush_seconds: float = 0.0, io_concurrency: int | None = None, transient_failures: int = 0, root_probes: int = 0, root_probe_seconds: float = 0.0,):Mutable counters for one metadata runtime run.
The run's flow body threads this into the tasks that do the work; each
reports what it did (record_key, record_reuse, record_unchanged,
record_parse, record_flush) and emit_summary closes the run out.
Every emission is swallowed on failure — telemetry must not break an
indexing pass.
Variables
- static
bytes_read : int
- static
dataset_identifier : str | None
- static
files_keyed : int
- static
files_reused : int
- static
files_seen : int
- static
files_unchanged : int
- static
flush_seconds : float
- static
io_concurrency : int | None
- static
key_seconds : float
- static
parse_seconds : float
- static
root_probe_seconds : float
- static
root_probes : int
- static
rows_written : int
- static
run_type : str
- static
task_hash : str
- static
transient_failures : int
elapsed_seconds : float- Wall-clock seconds since these counters were created.
files_per_second : float- Files seen per second so far, or0.0before any time has passed.
-
process_read_bytes : int | None- Bytes this process has read since these counters were created.Unlike
bytes_read, which sums the sizes of the files handled, this is what the OS counted. It is process-wide, so runs sharing a process mix, and on Windows it includes network I/O.Nonewhere the OS has no counter (macOS).
Methods
emit_summary
def emit_summary(self, status: str, error: BaseException | None = None) ‑> None:Emit the run's totals.
Sent at ERROR level for a failed run, carrying error's type and stack
frames, so a failed pass does not read as a healthy one.
Arguments
status: The run's bookkeeping status ("success","skipped","partial"or"error").error: The exception that failed the run, if any.
record_flush
def record_flush(self, *, rows: int, seconds: float) ‑> None:Record one persisted chunk and emit a progress event.
Arguments
rows: Rows the flush wrote.seconds: Time the flush took.
record_key
def record_key(self, *, size_bytes: int, seconds: float) ‑> None:Record a file whose content key was computed.
Arguments
size_bytes: Bytes read to compute the key.seconds: Time the key computation took.
record_parse
def record_parse(self, seconds: float, *, size_bytes: int = 0) ‑> None:Record one file whose vendor headers were parsed.
file_metadata counts a file via record_key / record_reuse and
scan_metadata counts it here, so files_seen is one file per run
either way.
Arguments
seconds: Time the parse took.size_bytes: Size of the parsed file, counted towardsbytes_read— the parse reads most of the file in practice.
record_reuse
def record_reuse(self) ‑> None:Record a file whose stored content key was reused unchanged.
record_unchanged
def record_unchanged(self) ‑> None:Record a file skipped because its stored row still describes it.
No content key is computed and no row is written for it.
record_unkeyed
def record_unkeyed(self) ‑> None:Record a file indexed without computing a content key.
Hashing is optional (settings.file_hashing_enabled) and off by
default. With it off no key is computed and no bytes are read, so the
file counts as seen but as neither keyed nor reused — counting it as
keyed would report the whole dataset as bytes read by a pass that read
none of it.