Dr. Benjamin Yudkoff and Dr. Trishan Panch at Brigham and Women's Hospital: Turning Patient-Reported Outcomes into Actionable Insights

Please note that throughout this blog, we may refer to ketamine, esketamine, and Spravato relatively interchangeably. This is due to the inherent similarities in chemical makeup between ketamine and esketamine, and their similar effects on mental health conditions. Don't hesitate to reach out to Lumin Health staff to ask any questions about treatment at hello@lumin.health or by scheduling a free consultation.

Patient-reported outcomes (PROs) are structured measures of how patients themselves say they are doing, and at this year's PROVE Summer School at Brigham and Women's Hospital, Lumin Health's chief medical officer and chief strategy officer presented a vision for deriving them directly from the clinical conversation using AI – reducing questionnaire burden while keeping validation, consent, and clinician judgment at the center of care.

A brief word first on why we were in the room. Each summer, the PROVE Center at Brigham and Women's Hospital convenes a program on patient-reported outcomes that draws clinicians, trainees, and measurement researchers from across the country and abroad. Dr. Andrea Pusic and her team invited the two of us to close out this year's sessions, and we accepted for a straightforward reason: the questions now being asked about AI and clinical measurement are too consequential for any single organization to answer on its own.

That is what this talk was really about. It was not a product announcement and not a finished result – it was an open conversation with the academic community whose standards we intend to be held to. We went to share what we are building, to hear where the researchers in the room thought we had it wrong, and to begin the kind of working relationship that turns a promising idea into something clinicians and patients can actually trust. What follows is the substance of what we presented, and some of what we heard back.

A Trust Gap in Mental Health Measurement

Dr. Trishan Panch, our chief strategy officer and board chair, and I opened by asking the room two questions.

First: how useful do you believe patient-reported outcomes are for improving the quality of mental health care? 81% said very or extremely useful.

Then: how comfortable would you be sharing your own clinical data with AI models to help improve quality of care? Only 14% said very comfortable.

Sit with that for a moment. The people who have spent their careers building the science of outcome measurement believe deeply in its value – and even they hesitate to hand their own data to an algorithm. That gap between belief and trust is, to my mind, the central design problem for the next decade of mental health care. Any system that wants to measure must first earn the right to listen.

What Are Patient-Reported Outcomes (PROs)?

A patient-reported outcome is exactly what it sounds like: an assessment of health that comes directly from you, without a clinician interpreting or filtering it first. In mental health care, PROs usually take the form of validated questionnaires – the PHQ-9 depression scale is the most familiar example – asking about mood, sleep, energy, concentration, and safety.

When they are collected well, PROs work. In a landmark randomized trial, systematically collecting patient-reported symptoms during cancer treatment was associated with fewer emergency visits and even longer survival. The principle carries into psychiatry: treatment decisions grounded in what patients actually report tend to be better decisions.

This kind of measurement is already woven into ketamine treatment with us today. Patients receiving esketamine (Spravato) or ketamine therapy complete symptom measures over their course of treatment, and our clinical team uses them to guide decisions about pacing, dosing, and next steps.

Why Questionnaire-Based PROs Fall Short

And yet, in day-to-day mental health care, questionnaire-first measurement carries real friction. In our talk, we described three failure modes:

  • Burden. Questionnaires ask patients to do homework outside of their care. The forms can feel tedious and confusing, and they often repeat conversations a patient has already had with their clinician.
  • Missingness. Collection is episodic and incomplete. The patients most likely to complete forms are the most digitally connected and most comfortable with paperwork – which means the most vulnerable voices are often the least represented in the data.
  • Latency. Scores arrive late, or fail to trigger action at all. A form completed days after a difficult week can miss the moment when a change in care would have mattered most.

There is a subtler problem underneath those three. In the dynamic of care, filling out forms can feel like something the patient does for the system, rather than something the system does for the patient. That is backwards, and patients feel it.

The Conversation Already Contains the Evidence

Here is the insight that reframed this problem for me as a psychiatrist: a good clinical conversation already contains nearly everything a PRO is trying to capture.

When I sit with a patient, we talk about mood, sleep, and side effects. But we also talk about whether they are back at work or school, whether joy is returning, whether the burden of treatment feels sustainable, and whether they feel safe. These are the real domains of recovery – expressed in the patient's own words rather than on a five-point scale.

So the question we posed at the Brigham was this: could AI translate that naturally occurring language into valid, structured measurement, without asking patients to do any additional work? Instead of running a parallel process of forms and scoring, measurement would live inside the care already being delivered – and flow back to the people who can act on it.

An AI Measurement Layer: Sense, Measure, Interpret, Act, Learn

The architecture we presented is a loop: sense what is happening in care, turn it into validated measurement, interpret it in clinical context, act on it, and learn from the result. Done well, that loop pays off at three levels:

  • For the clinician: early warning. Seeing, for example, that functioning is improving while mood remains flat, or that treatment burden is rising even though symptoms are responding – patterns good clinicians catch, but not always systematically.
  • For the care team: acting together. Timely signals can trigger review, outreach, and safety workflows while there is still time for them to matter.
  • For the organization: continuous learning. Instead of measuring episodically on the subset of patients who complete forms, a service can understand how all of its patients are doing and keep improving the care itself.

One line from the talk is worth repeating here: a vigilant system is not a surveillance system. It is a quality system that watches for opportunities to improve care. The distinction is real, but it has to be earned – which brings us to validation.

Validation Before Use: Earning the Right to Measure

Let me be direct about where we stand today, because trust depends on it: we do not record clinical consultations. We use AI extensively on the operational side of our organization – scheduling, administrative work, quality systems – but the consultation itself remains a protected space between you and your clinician.

Before an AI-derived measure ever informs care, it has to behave like a clinical-grade measurement instrument. The research program we described, developed with academic collaborators, works in stages: pairing consultation transcripts with completed, validated questionnaires so a model can learn the mapping. Testing prospectively in silent mode, where the model's scores are compared against standard instruments but never shown to clinicians. Measuring agreement, sensitivity to change, and fairness across patient subgroups. And only then, governing any clinical use with clinicians in review at every step. A framework for this kind of conversational measurement was recently published in npj Digital Medicine by our collaborator Prof. Laurent Boyer and colleagues.

A small example shows why the rigor matters. A patient who does not mention sleep is not the same as a patient who denies a sleep problem. A trustworthy system has to know the difference – and when the conversation does not contain enough evidence, it has to say so rather than guess.

The validation bar is not whether the model is technically plausible. It is whether it behaves like a measurement instrument medicine can trust.

What This Means for Patients

None of this changes what treatment looks like at Lumin Health today. Esketamine (Spravato) remains FDA-approved for adults with treatment-resistant depression and for major depressive disorder with suicidal thoughts, delivered under clinical monitoring, and IM ketamine treatment remains an evidence-based, off-label option within our psychiatrist-led protocols.

What this work points toward is care where the measurement burden falls on the system instead of on you. Clinicians who notice changes sooner. Care teams that reach out while it still matters. And measurement that finally hears the people questionnaires miss – patients who are less comfortable with forms, speak English as a second language, or simply have no energy left for homework. If AI in mental health care is going to mean anything, it should mean that listening well becomes the standard, not the exception.

If you are living with depression that has not responded to other treatments and are exploring whether Spravato treatment or ketamine therapy may be a fit, explore our services or meet our clinical team. We would be grateful to walk with you toward relief.

Frequently Asked Questions

What are patient-reported outcomes (PROs) in mental health care?

Patient-reported outcomes are assessments of health that come directly from the patient, usually through validated questionnaires about mood, sleep, functioning, and safety. In mental health care they help clinicians see whether treatment is working from the patient's own perspective.

Does Lumin Health record treatment sessions or consultations?

No. We do not record consultations, and no AI listens to visits today. Any future conversation-based measurement would be introduced only with patient consent, after formal validation, and with clinicians reviewing every output.

How is AI used at Lumin Health today?

AI supports the operational side of the organization – scheduling, administrative work, and quality systems – not the clinical consultation itself. The AI measurement layer described in this article is a research direction being developed and validated with academic collaborators.

Could AI replace mental health questionnaires like the PHQ-9?

Not today. Validated questionnaires remain the standard for measuring outcomes in esketamine and ketamine for depression care, and across psychiatry more broadly. Research is now testing whether AI can derive the same measures from clinical conversations accurately and fairly enough to reduce questionnaire burden over time.

Latest medical review on: July 15th, 2026. Medically reviewed by Instructor in Psychiatry at Harvard Medical School and Lumin Health Co-founder, Chief Medical Officer Dr. Ben Yudkoff.