Why AI Psychiatrists Can't Just Bundle Diagnosis, Observation, and Note-Taking Together
AI scribes are now routine in psychiatric practice, but they're quietly bundling three fundamentally different tasks into a single clinical note without adequate oversight of how each one fails. A new analysis in The Lancet Psychiatry argues that AI-generated clinical narratives, mental status examinations, and diagnostic reasoning are functionally distinct operations that demand separate evaluation and regulation.
What Are AI Scribes Actually Doing in Psychiatry?
AI scribes promise to reduce administrative burden by automatically generating clinical notes from session transcripts. The appeal is straightforward: if psychiatrists spend less time typing, they can be more present with patients, gathering richer clinical data and providing a therapeutic experience of being truly seen. In practice, however, a single AI scribe output bundles together three separate cognitive tasks that operate in fundamentally different ways.
The first task is narrative synthesis: converting a conversation transcript into a coherent clinical history. The second is observational inference: generating a mental status examination from ambient audio or video. The third is diagnostic reasoning: producing an assessment and plan based on the available information. These sound like they belong together, but they don't. Each one has a different relationship to the underlying data, and each one fails in distinct ways.
Why Does Psychiatry Expose These Differences So Clearly?
Psychiatry is uniquely positioned to surface these distinctions because of three factors: the therapeutic and legal weight of word choice in psychiatric documentation, the unnarrated structure of the mental status examination, and the formulation-driven nature of diagnostic reasoning. In other medical specialties, these differences might remain hidden. In psychiatry, they become impossible to ignore.
Consider word choice. In psychiatry, the exact language used to describe a patient's mental state can have profound legal and therapeutic consequences. A patient described as "depressed" versus "experiencing anhedonia" versus "expressing passive suicidal ideation" are three different clinical pictures, and the AI's choice of words matters enormously. Similarly, the mental status examination is traditionally unnarrated; it's not something the patient tells the clinician, but rather something the clinician observes and interprets. An AI generating this from a transcript is performing a fundamentally different operation than one generating a narrative history.
How Do These Operations Fail Differently?
When an AI scribe bundles these three operations into a single note, it becomes nearly impossible to evaluate each one independently. More troubling, errors can cascade across them. A mistake in narrative synthesis might propagate into the mental status examination, which might then distort the diagnostic reasoning. Clinicians reviewing the final note may not realize that the error originated three steps back.
The research community has begun to grapple with these distinctions through mechanistic interpretability studies, which examine how large language models (LLMs) actually arrive at their outputs. One recent mechanistic interpretability study specifically investigated clinical reasoning variability in medical LLMs, finding that these models show significant inconsistency in how they approach diagnostic reasoning. This variability is not random; it reflects genuine differences in how the model processes different types of clinical information.
Steps for Clinicians Using AI Scribes Today
- Separate Evaluation: Review each component of the AI-generated note independently; don't assume that because the narrative is accurate, the mental status examination and assessment are also correct.
- Verify Observational Claims: Pay particular attention to statements about the patient's mental status that the AI inferred from the transcript; these are the most likely to contain interpretive errors that don't match your direct clinical observation.
- Document Your Reasoning: When you disagree with the AI's assessment or plan, explicitly note your own clinical reasoning in the record; this creates a paper trail that distinguishes your judgment from the AI's output.
- Understand Liability Boundaries: Recognize that you remain clinically and legally responsible for every statement in the note, even those generated by the AI; passive acceptance of AI output without verification is not a defensible practice.
The authors of the Lancet Psychiatry analysis call for the field to build frameworks that reflect the different claims these AI scribes' outputs make. This is not a call to abandon AI scribes entirely, but rather to use them more thoughtfully, with explicit attention to which operations they're performing well and which ones require human oversight.
What Do Regulators and Ethicists Need to Know?
The regulatory and consent implications are substantial. Current informed consent processes for AI scribes typically treat them as a monolithic tool: "Your psychiatrist may use an AI scribe to generate your clinical note." But patients should arguably understand that the AI is performing three distinct operations, each with different accuracy profiles and different implications for their care. A patient might reasonably consent to AI assistance with administrative note-taking while objecting to AI-generated diagnostic reasoning.
Liability frameworks also need refinement. If an AI scribe makes an error in narrative synthesis, the psychiatrist bears responsibility for not catching it. But if the error originates in the AI's observational inference or diagnostic reasoning, the nature of the psychiatrist's oversight duty may be different. The psychiatrist cannot be expected to reverse-engineer the AI's reasoning process; they can only be expected to verify the final output against their own clinical judgment. This distinction matters for malpractice law and regulatory oversight.
The broader challenge is that AI scribes are entering clinical practice faster than the field can develop adequate oversight mechanisms. Psychiatry has an opportunity to get this right by insisting on frameworks that distinguish among these operations, but that window may be closing as adoption accelerates.