Rocco Sheldon, Morya Wadodkar, Andrew Gan, Rachael Yip, Yimeng Zhang, Jyoti Baharani
Abstract
Outpatient clinic letters are a cornerstone of clinical communication but frequently exceed recommended reading levels, limiting patient comprehension and engagement. Large Language Models (LLMs) and Artificial Intelligence (AI) have been proposed as scalable tools for generating and simplifying these documents, but the evidence regarding their clinical accuracy, safety, and effectiveness remains fragmented. We conducted a systematic review of studies evaluating AI or LLMs for generating, simplifying, or enhancing outpatient clinic letters. Five databases (PubMed, EMBASE, Web of Science, CENTRAL, and CINAHL) were searched from inception to 1 November 2025.
Introduction
Outpatient clinic letters are a cornerstone of clinical communication, serving as a formal record of patient encounters and a key mechanism for information exchange between patients and their primary and secondary care physicians. These letters typically summarise diagnoses, investigations, management and follow-up plans, and play a critical role in continuity and safety of care [1]. These differ from clinic notes, which are generally shorthand internal records of an encounter.
Methods
Protocol registration
This systematic review was prospectively registered in the International Prospective Register of Systematic Reviews (PROSPERO) database as CRD420251181303 (available from https://www.crd.york.ac.uk/PROSPERO/view/CRD420251181303). The protocol was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines (S1 Table) [20].
Results
The literature search across five databases and citation searching produced 1581 results. After removal of 416 duplicates, 1165 titles and abstracts were screened. Of these, 1139 were excluded, leaving 26 studies for full text screening. S4 Table lists all excluded studies with reasons. 19 studies were excluded for reasons outlined in the PRISMA flow diagram (Fig 1). This left seven eligible studies for data extraction. Study characteristics and results are summarised in Table 1.
Discussion
Key findings
This systematic review identified seven studies evaluating the use of AI and LLMs to generate, simplify, and enhance outpatient clinic letters. Overall, the evidence base remains small, heterogeneous, and subject to substantial bias. Nevertheless, several preliminary findings emerge.
Conclusion
In summary, preliminary evidence suggests AI and LLMs used to support the generation and simplification of outpatient clinic letters may lead to potential benefits for clinical accuracy, patient understanding, and efficiency. However, current evidence is insufficient to recommend routine clinical use beyond experimental contexts, and significant gaps remain regarding safety, equity, and implementation.
Citation: Sheldon R, Wadodkar M, Gan A, Yip R, Zhang Y, Baharani J (2026) Large language models and artificial intelligence for generating, simplifying, and enhancing outpatient clinic letters: A systematic review. PLOS Digit Health 5(9): e0001745. https://doi.org/10.1371/journal.pdig.0001745
Editor: Milit Patel, The University of Texas at Austin, UNITED STATES OF AMERICA
Received: January 11, 2026; Accepted: September 6, 2026; Published: September 22, 2026
Copyright: © 2026 Sheldon et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All relevant data are within the manuscript and its Supporting information files. The search strategy is detailed in the methods and appendices to allow the search to be replicated. The completed data extraction sheet is also available in the appendix to allow for repeated exploration of the data.
Funding: The author(s) received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.