Table of Contents
AI patient portal messages can take physicians more time to edit than it would take to write the replies from scratch, Dartmouth researchers found. The Hanover, New Hampshire-based team’s findings challenge the widely held assumption that AI drafting tools automatically save clinician time.
How Researchers Studied AI Patient Portal Messages
The team analyzed 146,000 messages exchanged between 10,105 patients and their primary care physicians on the online portal at Dartmouth Health in Lebanon, New Hampshire. This large dataset gave researchers a substantial real-world sample to evaluate how AI patient portal messages performed compared to responses clinicians wrote themselves.
Errors and Missing Follow-Up Questions
AI-generated drafts frequently introduced errors and extraneous details and often failed to ask relevant follow-up questions, according to the study presented July 7 at the 64th Annual Meeting of the Association for Computational Linguistics in San Diego. These quality gaps directly affected how much editing time clinicians needed to invest before a message was ready to send.
Which AI Models Were Tested for Patient Portal Messages
Researchers tested drafts from Claude, Gemini and ChatGPT, along with three smaller commercial models, Llama, Aloe and Qwen, against a dataset of real clinician-written responses. Comparing multiple models against actual physician replies allowed the team to assess how AI patient portal messages performed across different underlying systems rather than relying on a single AI tool’s output.
AI Can Sound Right Without Reasoning Right
“We find that AI can sound like a doctor but not think like one,” said Sarah Preum, PhD, the study’s co-corresponding author and an assistant professor of computer science at Dartmouth College, in a July 6 Dartmouth news release. This distinction between tone and clinical reasoning captures a core limitation researchers identified across the AI patient portal messages they evaluated.
A Technique That Improved AI Patient Portal Messages
Adapting the models to an individual physician’s communication style improved response accuracy by 33% and cut editing time by 26%, using a technique the researchers developed called TADPOLE. This personalization approach suggests that tailoring AI patient portal messages to a specific clinician’s voice and habits can meaningfully improve their usefulness.
The Limits of Personalization
The gains have limits, however. “If you have to edit 75% of the message, you may be spending more time and energy on making changes than if you were to just write it from scratch,” said co-author Tim Burdick, MD, an associate professor of community and family medicine at Dartmouth’s Geisel School of Medicine and a family medicine physician at Dartmouth Health. His comment underscores that even improved AI patient portal messages may not offer genuine time savings once heavy editing is required.
What This Study Means for Clinical AI Adoption
The findings add to a growing body of research questioning whether generative AI reduces or simply redistributes clinician workload. Rather than eliminating administrative burden, AI patient portal messages in their current form may simply shift the type of work clinicians do, from drafting from scratch to reviewing and correcting AI-generated content.
What Health Systems Should Consider
For health systems weighing whether to deploy AI drafting tools for patient messaging, this research suggests that measuring true time savings requires looking beyond whether a draft exists at all, and instead tracking how much editing that draft actually requires. As techniques like TADPOLE continue to develop, personalization may prove to be a key factor in determining whether AI patient portal messages genuinely reduce physician workload or simply create a different kind of administrative task.
For more healthcare industry updates, insights and news, visit DistilINFO. Click here to subscribe to stay informed.
