Harvard Medical School professor David W. Bates: hospitals need to choose applications carefully, confirm a tool works in their own setting, and keep monitoring its performance after go-live.
Medical AI is spreading fast, from generating clinical notes to reading images to predicting patient risk, and hospitals have ever more to choose from. But which applications are worth the investment? And once one is in place, how do you confirm that care has actually improved?
At the second Asia-Pacific Healthcare Quality Forum in December 2025, David W. Bates - professor of medicine at Harvard Medical School and medical director of clinical and quality analysis at Mass General Brigham - gave an online keynote on what he expects of medical AI and what it takes in practice. He also co-directs the system's Center for AI and Bioinformatics Learning Systems and sits on the global expert panel for the Newsweek and Statista World's Best Hospitals and World's Best Smart Hospitals rankings.
Bates is optimistic about AI improving care, but he stressed that healthcare organisations need to choose applications carefully, confirm that a tool really works in their own setting, and keep watching its performance after go-live. Understanding which harms remain poorly prevented is where the choice of application begins.
Bates opened with his team's study in the New England Journal of Medicine: across 2,809 admissions analysed, 23.6% involved at least one adverse event - close to one in four admissions. An outpatient study found 7% of patients had experienced at least one adverse event, most often medication-related harm.
For hospital-acquired infection, venous thromboembolism and falls, he noted, reasonably effective prevention already exists, so he rates the additional potential of AI as moderate. Identifying an individual patient's risk of medication harm, catching undetected deterioration early, and helping reduce diagnostic error are, in his view, the more promising directions.
On medication safety, AI could combine records and laboratory results to predict the risk of a particular drug for a particular patient, and might help identify interactions. For pressure injuries, Bates offered bed surface moisture as an example of an environmental signal that could help predict risk earlier.
These are opportunities he sees; whether they deliver has to be confirmed through real deployment and evaluation. For a hospital, the question is which care problem still has an obvious gap that AI has a chance of narrowing.
Beyond preventing harm, Bates highlighted another common bind: hospitals accumulate enormous amounts of data that quality teams cannot readily get at and use.
Using the widely deployed Epic record system as an example, he explained that even with reporting and data access tools, constraints such as permissions can make routine quality measurement difficult. He has watched many organisations build separate databases to support quality management as a result.
Hospitals, he argued, should work towards risk-adjusted length of stay, resource use analysed by diagnosis and by provider, and complication rates for individual procedures, moving steadily towards outcome-oriented measurement. He also expects the public to demand quality and safety information increasingly.
Electronic clinical quality measures (eCQMs) are the foundation that keeps such measurement running. Shared standards that turn measure specifications into machine-readable form reduce time-consuming manual abstraction and reporting, improve consistency and lower error.
Even so, Bates cautioned, implementation still has to establish whether workflows and information systems need adjusting, whether different systems can interoperate, and whether a measure needs data from other care settings. Whether quality measurement takes hold depends on those practical conditions.
Among applications already showing positive results, Bates singled out AI scribes, which listen during a consultation and help produce the clinical note.
Research he cited associates their use with lower burnout and a higher sense of professional accomplishment; a study at another organisation also observed improved clinician wellbeing, with clinicians reporting that their documentation had improved. This is, in his view, one of the encouraging early developments in medical AI.
A model that performs well in development, however, can perform markedly worse at another hospital.
A model has to be validated at the hospital where it is deployed before you know how it performs for the patients there.
David W. Bates - professor, Harvard Medical School
He used the Epic sepsis prediction model as an example: external evaluation found an AUC of 0.63, below the 0.76 to 0.83 the developer had reported. By the evaluation results he cited, roughly one in five patients triggered an alert, and only 12% of those alerted developed sepsis. The model, he noted, was developed at three hospitals and then used at many different organisations.
Bates also pointed to the self-driving company Waymo, citing its crash reduction data and its practice of opening its raw data to outside experts, to illustrate what AI might do for safety. It is a cross-sector reference, not evidence that medical AI has achieved the same.
On his work with the Newsweek and Statista rankings, he explained that the smart hospital assessment looks at how a hospital uses technology to rethink care, covering electronic capability, telemedicine, digital imaging, artificial intelligence and robotics.
For organisations preparing to invest in AI, Bates suggested ambient clinical documentation and image interpretation, particularly of X-rays, as first candidates. Other applications should follow from the organisation's own care needs and improvement opportunities.
Making these changes requires parallel investment in people, computing capacity and governance. The performance of an AI tool can decay over time, Bates warned: beyond local validation before deployment, hospitals must keep monitoring after go-live to confirm the tool still supports clinical care effectively.