Back to projects

Platforms & Data

Oneka: master patient index

A Master Patient Index that resolves one canonical identity per patient across healthcare systems that each assign their own IDs.

TypeScriptNode.js 20PostgreSQL (RLS)RabbitMQElasticsearch 8RedisZodVitest + fast-checkFHIR R4HL7 v2 / MLLPIHE PIX/PDQDocker

Problem

Hospitals, clinics, insurance platforms, and specialty EMRs each assign their own patient identifiers. The same person accumulates separate records in separate systems with no link between them, so no system holds the full picture of a patient.

Oneka sits between those systems. When a record arrives from any connected source, it normalizes the demographics, scores that record against every known patient, and either links it to an existing identity or creates a new one. It maintains a single canonical golden record per patient that every connected system can reference, and propagates changes back out.

How it works

Every incoming record runs through a fixed pipeline: Validate, Normalize, Match, Route, Cross-reference, Notify. Validation rejects records missing name, date of birth, or gender. Normalization standardizes names to uppercase, phone numbers to E.164, and dates to ISO. The matcher produces a score in [0, 1] and routing acts on it: 0.85 and above auto-links to an existing golden record, 0.70 to 0.85 creates a new record flagged as a potential duplicate for review, and below 0.70 creates a new patient.

The probabilistic matcher compares six demographic fields, each with its own algorithm and weight: name via Jaro-Winkler (0.30), date of birth exact (0.20), government ID exact (0.15), address via normalized Levenshtein (0.15), phone exact on E.164 (0.10), and gender exact (0.10). Jaro-Winkler and Levenshtein are implemented from scratch with no external dependency. The score is a weighted average clamped to [0, 1], with adjustments for transposed names and phonetic similarity.

The service layer persists to PostgreSQL, publishes patient lifecycle events to a RabbitMQ topic exchange, and indexes into Elasticsearch. A connector layer adapts external systems, each sharing a circuit breaker, exponential-backoff retries, and a quarantine store. The API surface speaks REST v1 alongside FHIR R4, HL7 v2 over MLLP, and IHE PIX/PDQ.

API layer (REST v1 / FHIR R4 / HL7v2 MLLP / IHE PIX-PDQ)
   -> Service layer (Registration / Matching / Golden Record / Merge / Search / Audit)
   -> PostgreSQL (records, RLS, cross-refs, block keys)
   -> RabbitMQ (patient events) -> Connector layer (OpenEHR / OpenIMIS / custom)
   -> Elasticsearch (patient search)

Hard parts

  • Candidate generation without a full scan: a blocking layer reduces each record to a small namespaced key set (government-id key, surname-soundex key with and without birth year, full date of birth), and only records sharing at least one key become scoring candidates, turning O(n) per registration into a key-join.
  • Recall preservation under blocking: the key set is chosen so any pair the matcher would score at or above the review threshold shares at least one key. A year-free surname-sound key keeps records with a mistyped birth year co-blocked, so restricting scoring to key-sharing candidates drops no real match.
  • Database-level tenant isolation: PostgreSQL row-level security scopes every query to the current tenant via a session variable, and RLS is FORCED so even the table-owner role cannot bypass it; the block-key policy carries both USING and WITH CHECK so inserts under forced RLS succeed.
  • Golden record survivorship: values are chosen field by field by highest source trust, then most recent update, then most complete record, re-evaluated whenever a contributing source changes; merges are reversible with full pre-merge state in the audit trail.
  • Name comparison that treats an absent given name as unknown, so a middle name recorded on only one record does not penalize the score.

Results

The repository reports no runtime performance metrics, so none are claimed here. The verifiable outcomes are structural.

  • Blocking changes candidate generation from a per-record full scan to a bounded key-join, for both single registration and batch deduplication.
  • Cross-tenant data access is blocked at the database layer through forced row-level security, independent of application code.
  • String similarity is implemented in-project with zero external matching dependencies.
  • The matcher is covered by property-based tests (fast-check) plus unit and integration tests, including a PostgreSQL match-engine integration suite.
  • Both FHIR Implementation Guides compile with zero errors and zero warnings.

Artifacts

Private repository. A TypeScript MPI engine with REST v1, FHIR R4, HL7 v2/MLLP, and IHE PIX/PDQ surfaces; a PostgreSQL schema with reversible migrations; a RabbitMQ event pipeline; Elasticsearch search; a connector SDK with circuit breaker, retry, and quarantine; a Next.js dashboard; FSH-authored FHIR Implementation Guides for African and Nigerian patient identity; and Docker-based demo and distribution packaging.

Back to projects