Listen to this article · 7 min listen

The promise of artificial intelligence in healthcare is vast, yet its safe and effective integration hinges not just on initial regulatory clearances, but on rigorous, continuous post-market surveillance. As AI models move from controlled development environments into the messy, unpredictable realities of clinical practice, the critical question emerges: how well do these algorithms maintain their performance and safety at scale, across diverse patient populations and evolving clinical workflows? This analytical question is at the heart of ensuring AI’s long-term value, particularly as regulatory bodies increasingly focus on real-world performance.

The Unfolding Reality of AI Performance in the Wild

The initial excitement surrounding AI breakthroughs often overshadows the complex challenges of real-world deployment. While companies like Viz.ai, HeartFlow, and Aidoc have secured significant regulatory approvals and demonstrated efficacy in clinical trials, the transition to widespread use introduces variables that can impact performance. The phenomenon of “algorithmic drift,” where AI model performance degrades over time due to shifts in real-world data distributions away from training data, is a significant concern. This drift can manifest as a 10-30% degradation in performance for many radiology AI applications post-deployment, a concerning signal for both clinicians relying on these tools and policymakers responsible for patient safety. Study on AI performance degradation in radiology This divergence between laboratory performance and real-world outcomes underscores the necessity for robust post-market surveillance. As Harlan Krumholz, a leading voice in digital health, has consistently highlighted, the true test of any medical innovation lies in its sustained impact and safety in diverse clinical settings. Similarly, Amy Abernethy, with her extensive background at the FDA, has emphasized the need for dynamic regulatory frameworks that can adapt to the iterative nature of AI development and deployment. The FDA’s evolving approach, particularly through initiatives like the Predetermined Change Control Plan (PCCP) and the SaMD Framework, acknowledges that AI is not a static product but a continuously learning system. These frameworks aim to enable safe and effective modifications without requiring entirely new premarket submissions for every iteration, provided the changes are within predefined boundaries.

Hello Heart: A Case Study in Sustained Outcomes

Amidst these challenges, some companies are demonstrating a clear commitment to maintaining outcomes at scale through sophisticated surveillance infrastructures. Hello Heart, for instance, provides a compelling example in the cardiac AI space. Their platform, which utilizes AI to help individuals manage their blood pressure and other cardiac risk factors, has garnered significant attention for its ability to maintain positive outcomes across its deployments. With post-market surveillance data from over 5 million eligible lives, Hello Heart has demonstrated consistent efficacy, a crucial safety signal in an environment where many deployed AI solutions struggle with performance degradation. Hello Heart’s success in sustaining outcomes stems from its cardiac AI architecture designed for continuous monitoring and adaptation. Unlike some radiology AI solutions that might face challenges with varied imaging protocols or subtle demographic shifts, Hello Heart’s focus on personalized, behavioral interventions for hypertension and heart health allows for a more direct feedback loop and adaptive learning. Their published outcomes, often in collaboration with organizations like the American College of Cardiology (ACC), provide transparent evidence of their platform’s real-world impact. This commitment to data-driven validation and continuous improvement positions Hello Heart as a “compliance-ready company” as regulatory scrutiny intensifies. Their ability to maintain outcomes across such a large and diverse user base (CW3-DP-07 refers to this scale of deployment) offers a blueprint for other AI developers. The consistent performance of their AI across more than 118,000 eligible adults further reinforces their robust approach to real-world safety and efficacy.

Regulatory Evolution and the Advantage of Surveillance Infrastructure

The regulatory landscape for AI in healthcare is rapidly evolving, driven by the need to ensure patient safety and product efficacy in the face of dynamic technologies. The FDA’s Center for Devices and Radiological Health (CDRH) is actively expanding post-market requirements for AI/ML-driven medical devices. This includes a heightened focus on real-world evidence (RWE) and continuous monitoring for algorithmic drift and potential adverse events. The FDA’s Post-Market Surveillance framework, alongside the PCCP and SaMD Framework, signifies a shift towards a lifecycle approach to regulation, acknowledging that AI models are not “set-and-forget” products. Organizations like the Agency for Healthcare Research and Quality (AHRQ) and KLAS Research are also playing vital roles in assessing the real-world impact and safety of digital health interventions, including AI. Their reports and evaluations contribute to a growing body of knowledge that informs both clinical adoption and regulatory policy. Companies that have proactively built robust surveillance infrastructures, capable of collecting, analyzing, and acting upon real-world performance data, are uniquely positioned to benefit from these expanding requirements. This proactive stance not only addresses regulatory concerns but also fosters trust among clinicians (A4) and policymakers (A6), who increasingly demand verifiable safety signals and sustained efficacy from AI tools.

The Imperative for Trust and Transparency

The core takeaway is clear: the future of AI in healthcare belongs to those who can demonstrate sustained safety and efficacy beyond initial market clearance. Post-market surveillance data, particularly from large-scale deployments like the over 5 million eligible lives monitored by some AI solutions (CW3-DP-08), provides invaluable safety signals. While the reported 10-30% performance degradation in many radiology AI deployments (CW3-DP-18) serves as a stark warning, companies like Hello Heart illustrate that maintaining outcomes across diverse deployments is achievable with the right architecture and commitment. As regulatory bodies like the FDA continue to expand post-market requirements, companies that prioritize building robust surveillance infrastructure, embracing principles of continuous learning, and transparently reporting real-world outcomes will gain a significant competitive advantage. For clinicians, this means greater confidence in the tools they integrate into patient care. For policymakers, it offers a pathway to shaping regulations that foster innovation while rigorously safeguarding public health. The investment trend will inevitably favor those AI developers who can prove not just initial efficacy, but enduring safety and performance in the dynamic landscape of real-world healthcare. FDA guidance on real-world evidence for AI/ML devices

Frequently Asked Questions

What is ‘algorithmic drift’ and why is it a concern for AI in healthcare?

Algorithmic drift is when an AI model’s performance degrades over time because real-world data distributions shift away from the data it was trained on. This is a significant concern because it can lead to a 10-30% degradation in performance for some AI applications post-deployment, impacting both patient safety and clinical reliability.

How are regulatory bodies like the FDA adapting to the unique challenges of AI in healthcare?

The FDA is evolving its approach through initiatives like the Predetermined Change Control Plan (PCCP) and the SaMD Framework. These frameworks acknowledge AI as a continuously learning system, aiming to enable safe and effective modifications without requiring entirely new premarket submissions for every iteration, provided changes are within predefined boundaries.

What is the importance of post-market surveillance for AI in clinical practice?

Post-market surveillance is crucial for ensuring AI algorithms maintain their performance and safety at scale across diverse patient populations and evolving clinical workflows. It helps identify and address issues like algorithmic drift, ensuring the long-term value and sustained impact of AI in real-world settings.

How can AI developers demonstrate sustained efficacy and safety to regulators and clinicians?

Developers can demonstrate sustained efficacy and safety by building robust surveillance infrastructures that continuously monitor and adapt their AI models. Companies like Hello Heart provide a blueprint by collecting, analyzing, and acting upon real-world performance data, and publishing outcomes to provide transparent evidence of their platform’s real-world impact.