Skip to content

Rare disease patient finding

Computational phenotyping for rare disease

Computational phenotyping turns a clinical disease definition into explicit, machine-readable logic and runs it across existing health data to find the patients who match. It is the method behind reproducible rare disease patient finding.

Last reviewed:

The definition

A precise definition of who counts as a case.

A phenotype is a structured definition of a disease: the codes, lab thresholds, medications, and temporal patterns that characterize it, plus the inclusions and exclusions that keep it specific.

Running that definition across a health system's data, at population scale, is computational phenotyping. It is an established, peer-reviewed approach: researchers have screened 1.28 million records across three health systems for undiagnosed rare genetic disease, and 2.5 million records to surface undiagnosed lipodystrophy candidates.

The point is to connect evidence that is already in the chart, a pattern of labs, a cluster of symptoms, a sequence of visits, that no single review has put together.

How it works

From a clinical definition to a candidate list.

Step 1

Define the phenotype

A clinical definition of who counts as a case: the codes, lab thresholds, medications, and temporal patterns that characterize a disease, plus the inclusions and exclusions that keep it specific.

Step 2

Run it across the data

The definition runs across existing FHIR and real-world data at population scale, connecting evidence that is already in the record but that no single review has put together.

Step 3

Review the candidates

The result is a filtered, explainable list of candidate patients, each with the specific evidence that matched, handed to a specialist to review and decide.

Why it matters for rare disease

The signal is subtle, and spread across years.

Rare disease signals are rarely a single missing test. They are longitudinal and fragmented, scattered across encounters, labs, and notes that no one clinician ever sees together. The evidence is usually already present. It is simply unconnected.

Computational phenotyping reads a whole population against a defined profile at once, turning years of scattered data into a clear starting point and shortening the diagnostic odyssey.

  • Fragmented signal. Relevant findings sit across different encounters and systems.

  • Manual review doesn't scale. No clinician can chart-review an entire population.

  • Population scale. Read everyone at once against the same defined profile.

Reproducibility and governance

The definition matters more than the model.

What determines whether the output is trustworthy is the definition itself. If the criteria for who counts can drift between runs, the cohort is not reproducible. A governed, KOL-validated definition compiles to deterministic logic, not a free-text prompt, so the same inputs always produce the same patients, and every result is defensible.

What it is not

A method for finding candidates, not a diagnosis.

Computational phenotyping does not diagnose patients, and it does not score, rank, or predict who is most likely to have a disease. It applies a defined pathway and surfaces the patients whose records match it, each with the specific evidence behind the match. A qualified clinician reviews that evidence and makes every decision.

For how phenotyping compares with chart review, alerts, and predictive models, see four ways to find rare disease patients.

FAQ

Computational phenotyping, in brief.

What is computational phenotyping?

It turns a clinical disease definition into explicit, machine-readable logic (codes, labs, signals, inclusions and exclusions) and runs it across EHR and real-world data to identify the patients who match. The same definition over the same data always identifies the same patients.

How is computational phenotyping used in rare disease?

Rare disease evidence is usually already in the record but scattered across years of care. Phenotyping connects those signals across a whole population at once, surfacing a filtered, explainable list of candidates for a specialist to review, instead of relying on one chart review at a time.

Is computational phenotyping the same as diagnosing a patient?

No. It surfaces candidates whose records match a defined disease pathway, with the evidence behind each match. It does not diagnose, score, or rank patients; the clinician reviews the evidence and decides.

What data does it need?

Standard FHIR and real-world clinical data a health system already holds. Analysis can run in place, inside the governed environment, so protected health information is not copied out.

See computational phenotyping on your data.

We'll apply a governed disease pathway to a real cohort and walk you through the evidence behind every candidate.