Skip to main content

store

Store for the patients identity table (v2).

Module​

Functions​

bulk_upsert_patients​

def bulk_upsert_patients(cache: CacheProtocol, records: list[PatientRecord]) ‑> None:

Upsert patient identity rows on bitfount_patient_id, coalescing PHI.

Identity is written once per patient (the patients table is not partitioned by task_hash), so a later run must not downgrade what an earlier run already knew. A run without EHR context (EHR outage, a partial batch, or a no-EHR template) produces a PatientRecord with ehr_patient_id=None / mrns=[] and a name that may fall back to the bare patient ID; merging that verbatim would wipe previously-stored PHI that the patient-data API then serves. So each identity/PHI field is updated only when the incoming run actually supplies it, otherwise the stored value is kept. Provenance (task_hash/run_id/processed_at) always reflects the latest write.

The name is the one field where "supplied" is not enough to decide. The table is unpartitioned, so the runs writing this row include runs over other datasources and runs belonging to other projects, and two of them can supply different names for one patient — each perfectly real, one derived from a better source than the other. Coalescing on presence made that a race the last writer won. v2 decides it on name_provenance instead, via _name_replaces_stored; a name that loses leaves name, name_raw and name_provenance all as they were, so the three always describe each other.

Arguments

  • cache: The cache backend.
  • records: Patient identity records to persist.

get_patient​

def get_patient(    cache: CacheProtocol, bitfount_patient_id: str,) ‑> PatientRecord | None:

Return the patient identity row for bitfount_patient_id, or None.

Arguments

  • cache: The cache backend.
  • bitfount_patient_id: The Bitfount patient ID.

Returns The PatientRecord, or None if absent.

list_patients​

def list_patients(    cache: CacheProtocol, ids: list[str] | None = None,) ‑> list[PatientRecord]:

List patient identity rows, optionally restricted to ids.

An ids list is queried in batches. Each ID becomes a bound parameter and SQLite caps those per statement (32766 on a current build, 999 on an older one), so a site with more patients than the cap would otherwise fail the whole listing with too many SQL variables.

Arguments

  • cache: The cache backend.
  • ids: Optional list of patient IDs to restrict to. None lists every row; an empty list lists none.

Returns The matching PatientRecords.

search_patients​

def search_patients(    cache: CacheProtocol,    *,    search: str | None = None,    eligible_for: Mapping[str, str] | None = None,    scoped_to: Mapping[str, str] | None = None,    limit: int,    offset: int,) ‑> tuple[list[PatientRecord], int]:

Return a page of patients plus the total matching count.

Filtering, ordering and windowing all happen in SQL — name and ehr_patient_id are plaintext columns (see the schema docstring), so a case-insensitive substring search is a LIKE and the result is ordered by name for stable pagination.

The eligible_for project filter and the scoped_to identity scope are each applied as a subquery (against patient_eligibility and patient_level_eligibility respectively, built by those types' stores, so this query stays a single statement) rather than by resolving either to a candidate id-set and binding it as an IN (?, ?, ...) — an id-list overflows SQLite's bound-parameter limit once the set is larger than it. eligible_for uses an IN (SELECT ...) form, which also lets the eligibility index drive the filter (see eligible_for_filter); scoped_to uses a correlated EXISTS (see in_partitions_exists), and matters most here because its set is ~every current patient.

A search long enough to be indexed is narrowed by the FTS5 trigram index first (SEARCH_INDEX_TABLE), joined in as a CTE so the candidate rows come off the index instead of a scan of every patient. The LIKE is still applied on top of those candidates rather than replaced by the index match, for two reasons:

  • The index is not semantically identical to the LIKE. Trigram folds case for non-ASCII, where lower()/_ascii_lower fold ASCII only, so a bare index match returns a superset — josé would newly match JOSÉ NÚÑEZ. Re-applying the LIKE to the (small) candidate set keeps the result exactly what it was before the index existed.
  • It makes the index optional. _search_index_hits returns None for a short term or a cache without the index, the join is skipped, and the same LIKE serves the query unaided.

Ordering is served by ix_patients_name_order rather than a temp B-tree sort of the whole matching set. That index is only safe here because search is index-driven: on the LIKE-only path a term matching nothing makes the planner walk the whole name index without short-circuiting. The one term length that still takes that path — below _MIN_INDEXED_TERM_LENGTH, so no FTS index can serve it — is covered by not running the page statement at all once the count has proved the page would be empty.

Arguments

  • cache: The cache backend.
  • search: Case-insensitive substring matched against name or ehr_patient_id; None/blank matches everyone. Case folding is ASCII-only on both sides to stay consistent with SQLite's lower() (see _ascii_lower).
  • eligible_for: Optional project_id → authoritative task_hash map; restricts to patients eligible for any of those projects (union). None means no project restriction; an empty map matches nobody.
  • scoped_to: Optional project_id → authoritative task_hash map for the published patient_level_eligibility partitions; restricts the listing to patients present in them (the identity table's partitioned view), so a patient whose ID was derived under a superseded rule is excluded. None means no identity scope (serve every row — a cache with no partition recorded).
  • limit: Max rows to return (the page size).
  • offset: Rows to skip (page index × page size).

Returns A (records, total) tuple where records is the requested page and total is the count of all matching rows before windowing.