All articles

Clinical practice insights

What the research actually says about AI-written psychology reports

Why the reasoning and narrative — not the speed — are the real frontier, especially in medico-legal work

There is a great deal of confident marketing in the AI documentation space right now, and very little of it is grounded in evidence. Most of it sidesteps the only question that matters. It is not can AI draft a report; of course it can. It is whether it can construct the reasoning a report stands on: reasoning accurate, safe and defensible enough to carry your name and your registration.

That matters more in psychology than in almost any other field, and nowhere more than in medico-legal work. Our reports are not administrative summaries. They decide things, in workers compensation, public mental health, education, corrections and defence. Get the reasoning right and the report does its job; get it wrong and the fluent prose around it is worthless.

Here is what the research genuinely shows, what it does not, and where I think the real frontier lies.

Reasoning is the report

A medico-legal report is a chain of reasoning, not a collection of sections

A medico-legal report is not a form with the boxes filled in. It is a single, sustained line of reasoning: you gather the material, weigh competing explanations, address causation and contribution, connect the findings to a diagnosis, and carry that thread unbroken to your opinion. When a solicitor, an insurer or a tribunal reads it, they are not admiring the prose; they are following your logic and looking for the place it breaks.

This is why report writing swallows so many hours. The hard part was never the typing; it is the integration, and building an argument that survives a hostile cross-examiner. A report can be fluent and complete and still collapse the moment its reasoning is tested. In medico-legal work, the reasoning and the narrative are not one part of the job. They are the most important part of it.

The documentation problem is real. The prize for solving this is large, because the workforce is stretched to its limit

The problem AI is aimed at is real. The Australian Government's 2026 Psychology Supply and Demand study counts 41,066 registered psychologists in 2023, working an average of just 29.8 hours a week, against a shortfall set to grow from 534 full-time-equivalent psychologists in 2025 to more than 7,000 by 2038, with nearly one in three psychologists reporting emotional exhaustion. Documentation is a large, mostly invisible driver of that strain, and complex medico-legal reports are its most punishing form. So anything that safely gives clinicians that time back is a real contribution to access and capacity. The emphasis being the word safely.

What the research shows - and what it doesn't

The evidence is a snapshot of the tools that happened to exist at the time The most directly relevant study is Lockwood and colleagues (2025), which put AI-generated psychological reports head to head with human-written ones before licensed psychologist judges.

The findings supporting use of AI:

  1. AI reports were rated more highly for the quality of their recommendations.
  2. AI is strong precisely where reports are structured, action-oriented and forward-looking.

But the humans still won on overall quality, summary quality and willingness to sign off, and where the edge sat is the whole story: AI slipped exactly where reasoning and narrative live, in synthesis, integration and interpretive judgement.

Though this also ignores that the AI reports were not being used in the way that psychologists would use AI, which is as a first draft.

Here is the part the marketing never mentions, and it cuts both ways. This research is a snapshot of the AI products that happened to exist at the moment the studies were run.

Systematic review: Perkins and colleagues' 2024 systematic review finding showed no highly accurate, end-to-end AI documentation assistant anywhere in the peer-reviewed record.

  1. The one directly relevant study ran on a general-purpose United States model, GPT-4, evaluated in a United States setting.
  2. The scribes and assessment tools these reviews assess were, almost without exception, built to do one thing: accelerate drafting and impose structure.

So the evidence tells us something true but narrow.

It tells us that the general-purpose and speed-focused tools of the day were weak at reasoning. It does not, and cannot, tell us what a tool purpose-built for reasoning would do, because at the time of these studies, no such tool had been built or tested.

The gap in the evidence is not a verdict on what is possible. It is a map of what nobody had attempted yet.

Why Virtuosa is working on the frontier

Others optimised for speed and structure. We are going into the reasoning and the narrative.

This is exactly where Virtuosa AI is different. Most AI report-writing and scribe tools have made a reasonable but limited bet: focus on speed and structure. Get the clinician to a readable first draft faster, lay out the sections, organise the assessment data.

That is genuinely useful, and Virtuosa does it too. It cuts report-writing time and cognitive load and gives you a solid scaffold instead of a blank page.

But that is the initial goal, not the longer term one.

What is genuinely new about Virtuosa is that we are going into the part everyone else has stepped around: the reasoning and the narrative. We are building for the hardest, highest-stakes work, complex and medico-legal reports, where the interpretive thread that binds history, findings and opinion is the whole point, and where a tool that only accelerates structure leaves the most demanding work untouched.

That is the part the existing research shows is hardest, and the part it has never actually measured in a purpose-built tool.

And we intend to prove it rather than simply claim it.

We are planning preliminary validation research in collaboration with Oliver Guidetti Lecturer, PhD at the University of Wollongong, to test whether a tool built around reasoning and narrative can genuinely support the interpretive work at the heart of medico-legal reporting, to a standard the field has not yet examined.

That, to me, is an exciting new frontier - to build in this space: identify the exact gap the evidence exposes, build for it deliberately, and then submit the result to independent scrutiny.

What this means for you

Let AI accelerate the drafting; demand a tool that respects the reasoning

The practical takeaway is a division of labour that the evidence supports and that maps neatly onto how a good report is built. Let AI carry the structural, repetitive parts and the readable first draft. Keep every step of the reasoning, causation, competing hypotheses, materiality and the interpretive narrative, firmly in your own hands, because that is what a cross-examination is built to attack and what the research shows AI cannot yet be trusted to do alone. Human-in-the-loop is not a slogan; it is the line between a report that stands up and one that falls over.

But do not accept the quiet implication of the marketing, that speed and structure are all AI can offer a report writer. That was true of the tools the research tested. It does not have to be true of the tools we build next. AI can genuinely help you write reports, and in medico-legal work the reasoning is everything, so the tools worth your time are the ones brave enough to work on the reasoning itself, and honest enough to be tested on it.

Lisa Irving is a clinical psychologist and the founder of Virtuosa AI, an Australian report-writing platform built to support psychologists with high-quality, defensible first drafts of complex reports while keeping the clinician in control.