Date of Award

January 2026

Document Type

Open Access Thesis

Degree Name

Medical Doctor (MD)

Department

Medicine

First Advisor

David Chartash

Second Advisor

R. Andrew Taylor

Abstract

Background - Large language models (LLMs) can convert unstructured clinical text into structured, computable signals, creating a new class of clinical decision support (CDS) that is tightly coupled to real documentation workflows. This is especially relevant in emergency medicine, where decisions must be made rapidly under uncertainty, often with key information scattered across notes, reports, and prior encounters. Emergency department (ED) chest pain evaluation is a prototypical setting for this problem: it is high-volume, high-stakes, operationally costly, and driven by pathway-based triage and risk stratification, yet care remains heterogeneous across clinicians and teams. As health systems and electronic medical record vendors move toward integrating LLMs into routine clinical infrastructure, there is an urgent need for evidence regarding where along chest pain pathways LLMs add value, how they should be engineered and governed, and what risks emerge as artificial intelligence autonomy increases.

Research Aims 1. Empirically characterize real-world chest pain evaluation pathways and practice variation across physicians and advanced practice providers (APPs) to identify actionable CDS decision points. 2. Engineer and validate a HIPAA-aligned LLM workflow to compute the HEART (History, Electrocardiogram, Age, Risk factors, Troponin) score, a critical element of the chest pain pathway, automatically from clinical documentation. 3. Test the generalizable promise and limits of LLM-enabled diagnostic support in chest pain by evaluating stepwise information-seeking behavior and diagnostic calibration under realistic disease prevalence using a probabilistic model of medical decision-making.

Research Question - How can physicians integrate LLMs into existing ED chest pain pathways to improve the consistency, transparency, and operational fidelity of decision-making, particularly at high-leverage, documentation-heavy decision points, while maintaining safety, appropriate calibration, and stewardship of downstream testing?

Methods - Using retrospective data from Yale New Haven Health, we analyzed 49,713 adult ED chest pain encounters to quantify variation in diagnostic test ordering across physician-only, APP-only, and shared physician and APP care with multivariable regression. We then developed an iterative prompt-optimization framework for LLM-based HEART subscore extraction using a small set of synthetic chest pain encounter notes. Applying this framework, we validated the finalized pipeline on 551 deidentified, real ED encounters processed in a non-training, HIPAA-aligned environment. LLM-derived HEART scores were further compared with clinician scores and emergency physician adjudication and evaluated for discrimination of 60-day major adverse cardiac events (MACE). Finally, we evaluated GPT-4o as a stepwise diagnostic agent in simulated chest pain encounters using real-world data, applying a Bayesian network to impute unmeasured data and to generate mutual-information-optimal query sequences as a normative benchmark for information seeking strategy.

Results - Clinician configuration was independently associated with diagnostic strategy: even after adjustment for patient and visit factors, physician-only, APP-only, and shared-team encounters differed in patterns of advanced imaging utilization. Development of a prompt-optimization framework using synthetic encounter notes substantially improved GPT-4 HEART subscore accuracy, yielding a reproducible workflow for reliable structured extraction from clinical text. Applying this framework to real-world HEART score validation (n=551), GPT-4o-derived scores showed substantial agreement with adjudicating physicians (weighted κ≈0.67), comparable to APP agreement. Disagreement was concentrated in subjective components (history and electrocardiogram) and a tendency toward higher risk assignment. LLM-automated HEART score-based discrimination for 60-day MACE was similar across LLM, APP, and adjudicated scores (area under the receiver operating characteristic curve ~0.68–0.72). In autonomous, stepwise diagnostic simulations, GPT-4o’s accuracy and behavior were sensitive to prompt framing and prevalence context; the model over-called rare emergent diagnoses and showed low alignment with mutual-information-optimal query pathways while requesting more imaging and fewer vitals/labs than clinicians.

Statement of Scientific Impact - This thesis provides a workflow-anchored evidence base for LLM decision support in ED chest pain care. It identifies where practice variation concentrates, demonstrates that a HIPAA-aligned LLM can compute a widely used clinical risk score (the HEART score) from real documentation with clinician-comparable agreement, and shows that more agentic diagnostic use introduces calibration and utilization risks that are highly sensitive to context. This work serves as a concrete, practice-grounded stepping stone—linking technical design decisions to clinical, operational, and governance endpoints—to support efforts to build and evaluate the next generation of LLM-augmented decision support in emergency care.

Comments

This is an Open Access Thesis.

Open Access

This Article is Open Access

Share

COinS