Ami B. Bhatt, MD, Chief Innovation Officer, American College of Cardiology. Opinions my own.

getty
The way people ask about their own health is changing.
Major consumer AI platforms are launching dedicated health products that let users connect their medical records, lab results and visit summaries, then ask questions grounded in their own data. ChatGPT, in particular, reports that more than 40 million people bring health questions to its platform every day.
Health systems have started answering back, launching their own patient-facing chatbots built on patients' medical records and pitched as the safer, better-grounded option.
The debate that has followed is the one we always have. Are the answers accurate? Are they safe enough to trust? Those are fair questions. They are also, I think, the wrong ones.
The Number That Should Change The Conversation
Earlier this spring, a randomized study published in Nature Medicine tested whether leading large language models could help ordinary people work through common medical scenarios. When tested independently, the models performed remarkably well, identifying the correct underlying condition in roughly 95% of cases.
Then the researchers put real people in front of those same models, nearly 1,300 of them, each asked to reason through a realistic scenario with the model's help. They identified the right condition fewer than 35% of the time. On what to actually do next, they did no better than the participants who were given no AI at all. In other words, the model passed the test, but the patient did not.
The evidence now says scoring well and helping a real person are not the same thing. A tool that scores 95% on its own and produces a 35% result in a person's hands is not a 95% tool. The researchers were specific about where it broke down: The failure was in the transfer between the model and the human. This means we have been measuring the wrong end of that exchange.
Why Accuracy Is Not The Product
In medicine, the answer alone doesn't tell you whether the person across from you understood what was true about their body and what to do about it.
None of this is new. Long before generative AI, the distance between what a clinician knew and what a patient understood was one of the most expensive and least measured failures in the system. A correct diagnosis, poorly explained, could become a missed medication or a skipped follow-up or a readmission that did not have to happen. We have known this for decades.
However, AI has changed the scale of this problem. The gap that used to live inside one exam room at a time now sits in front of tens of millions of people who are asking these tools to explain their health every day.
Consider if someone gets a cholesterol panel back and asks a chatbot what it means. The explanation may be, in the narrow sense, correct. But they may walk away certain that a borderline number is a crisis or that a genuinely worrying one is nothing to act on because the answer was accurate without ever being understood.
The correct output can lead to the wrong outcome.
Three Doors To The Same Room
A patient's first question can now travel through three different doors, and we have not been honest about how differently they are built.
The first is the consumer platform, which connects directly to a person's medical records. It is powerful, convenient and sits outside the framework that governs the rest of healthcare. However, these platforms are not marketed as HIPAA-compliant, according to a Journal of Medical Internet Research study. Instead, the same record a hospital is legally bound to protect becomes governed by a company's terms of service the moment a patient connects to it.
The second door is the health system's own chatbot, grounded in the medical record, operating within that legal framework and built to pull the conversation back within institutional walls.
The third door is the clinician, who still carries full accountability and is winning less and less of the patient's attention.
Three doors, one room. They are not built to the same standard, and they do not carry the same obligations. None of them is currently designed or measured to ensure the patient leaves understanding their own care.
What actually separates them, and what no one is grading, is comprehension.
What Leaders Should Be Asking
I chair the FDA's Digital Health Advisory Committee, which means I spend a good deal of time in the rooms where the standards for these tools are debated. The reflex in those rooms and in the boardrooms, when making deployment calls, is to ask whether a model is accurate enough to trust. That question will not get us where we need to go. It measures the tool, not the handoff.
The better question for anyone deploying a patient-facing system is not how often it is right, but how often the person using it comes away with a correct understanding and a sound next step. Those are different measurements, and only the second one tells you whether care actually improved. A health system that buys a chatbot on its benchmark scores has bought based on the wrong assurance.
This is a design and measurement problem, not a case against the tools. Patients should have access to them. The agency they create is real and overdue. People want to understand their own health, and that instinct deserves to be met rather than managed.
But meeting it well means building and evaluating these tools against comprehension, which determines whether someone is genuinely safer for having used them.
The technology will keep getting better at producing answers. The open question for everyone building, buying and regulating it is whether we will measure what matters. The door a patient walks through should not decide whether they understand their own care.
Comprehension, not accuracy, is the standard these tools should be held to. Until that is what we design for, we will keep grading the test the machine takes and missing the one the patient fails.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

1 hour ago
2













English (US)