Artificial intelligence (AI) is increasingly being integrated into healthcare technologies, from diagnostic support tools to clinical decision-making systems. While these innovations offer substantial opportunities for improving health, they also present new challenges for those tasked with evaluating their relative effectiveness, safety and value. We spoke with Angel Varghese, Project Director at YHEC, about the complexities of assessing AI-based health technologies and what developers, evaluators and healthcare systems need to consider.
Tell us a bit about yourself and your role.
I’m a Project Director in YHEC’s Digital Health Sector. My background is a mix of Biochemistry, followed by an MSc in Health Economics, so I don’t come from a traditional economics background. My work focuses on the economic evaluation of health technologies, with a particular interest in digital health technologies (DHTs). Within that field, I have evaluated quite a few AI-enabled technologies across a range of healthcare applications.
Can you give us some background to DHTs and, more specifically, AI-based DHTs?
DHTs have been part of healthcare for many years. It’s an umbrella term that captures a variety of technologies, ranging from simple smartphone applications providing healthcare advice to more sophisticated platforms or software supporting clinical care. It doesn’t have to be a separate entity in itself; often it can be an add-on to existing technologies such as medical devices.
If a digital health app is the overall product used by patients or healthcare professionals, then the AI component can be seen as a specific element within the product that performs tasks such as prediction, classification, risk assessment, or personalised recommendation, the engine of a DHT if you will.
YHEC’s Director of Digital Health Technologies Consulting, Hayden Holmes, talked about the challenges with evaluating DHTs in a recent blog. Are there also particular challenges with AI-based technologies?
There certainly are challenges specific to evaluating AI-based technologies. As Hayden discussed in his blog, some of the issues are common regardless of whether it is a pharmaceutical or a device or DHT, but they tend to occur more frequently with DHTs. Similarly, AI-based DHTs often amplify these challenges and finding a consensus on how to evaluate them is an area of active discussion in health economics. Common issues across DHTs, including AI-DHTs, include the limitations of randomised controlled trials (RCTs) for assessing them, and how to link short-term diagnostic outcomes to clinical decision making, and subsequent long-term outcomes.
One of the key challenges with AI-enabled DHTs specifically is ensuring that the data behind them reflects the people who will ultimately use them. If the datasets used to train and test an algorithm do not represent the real-world population, there is a risk that unintended biases could affect how well the technology performs for different groups of patients. Typically, as health economists we don’t always have access to this information at a granular level to critique, and this makes it difficult to quantify the direction of bias that could be introduced without further evidence or safety netting.
Another challenge would be the safety risks directly associated with deploying AI-DHTs on a wider scale across really different and localised set-ups. Safety in this instance is elements outside those considered within the regulatory framework. For example, how are AI-DHTs integrated into NHS Trusts or settings with limited digital readiness, are different methods used? If so, does this inherently change how the AI-DHT is embedded and the impact it has on outputs? Is that likely to introduce safety holes not considered? Is there variation in clinical metrics used between wards or hospitals which may lead to errors in interpretation?
Additionally, while a national level analysis is key, it is important to test out variability locally. This is particularly important given the variability in digital readiness, workflow set up, and clinician expertise/hospital specialisation. Often with these technologies, their impact will vary by how end users apply it and their expertise.
What can health economists do to mitigate some of these challenges?
Sharing details of training and testing datasets is important as part of the evidence critique. It can be particularly helpful when you’re evaluating an AI-DHT with similar competitors by helping us consider, at least qualitatively, the wider implications of the economic outcomes estimated.
From an DHT effectiveness standpoint, one of the first things to do during the early stages of product development is to identify the key parameters/drivers of economic impact for these tools. Doing this as part of the initial proof-of-concept phase allows relevant data points to be captured. This makes sure that when you’re investing and developing studies, that they’re actually collecting results which are beneficial to address evaluation questions that health technology assessment (HTA) bodies and decision makers locally are interested in. I don’t think that’s always easy. There have been cases where clients have shared multi-site studies with us, but the key benefit outcomes weren’t appropriate for economic evaluation.
Mapping company proposed value proposition to end-user feedback is also very important for these tools. Currently the NHS, as per guidelines, operates under the human-in-loop for AI-DHTs. A human has the final decision-making responsibility regardless of the AI output. When thinking about common benefits we see with these tools (e.g. time savings or prioritised diagnosis), it’s important to have conversations/workshops upfront with the target end users to make sure analyses capture how the AI-DHTs may be used in practice, particularly if it is likely to deviate from the proposed value proposition. This can have wider implications. For example, a prostate cancer grading tool was proposed to save radiologists time, but in reality, experts were using it as a second read. This might not have major impacts in laboratory settings reviewing a large quantity of prostate biopsy slides, but for smaller labs, the impact of a digital second opinion might be greater with potential implications on repeat tests.
Diagnostics is one area where AI is increasingly being rolled out. What’s different about the approach to economic modelling in this context?
In terms of the approach and the model structure that you might use, those are all very much applicable to all other types of technology that you might come across, so it’s just understanding how to capture the impact of the technology in the clinical pathway.
It can be very difficult to map out the impact of short-term changes in outcomes to long-term clinical outcomes. While an AI tool might give a particular outcome, translating that into changes in clinical management, and quantifying the downstream impact, makes economic modelling quite complex. More specifically for these tools though they are often changing ‘grading’ that may be given as part of a diagnostic pathway (e.g. Prostate Imaging – Reporting and Data System (PI-RADs) scoring from magnetic resonance imagery (MRI) or grading of histopathology slides). It can be challenging to quantify from early data what this means for accurate diagnosis and onward management.
How do you deal with this complexity?
The best way to deal with this uncertainty is to explore a range of possible scenarios. Rather than relying on a single estimate, we can model different scenarios, from the most optimistic to the most pessimistic, to understand how the technology might perform under different circumstances. The reality is likely to lie somewhere between these extremes.
This approach is particularly important for emerging technologies such as AI-DHTs, where long-term evidence is often limited. In many cases, the technology needs to be used in practice for some time before we can fully understand its impact on patients, healthcare services and costs. But running scenarios on an early model allows us to explore if there is likely to be a benefit within any of these scenarios and what the key drivers of negative economic outcomes might be. This may also support safety netting decisions, if the tool is used more widely, to prevent harm.
The key is to be transparent about the uncertainty, test a range of plausible scenarios, and work closely with clinicians and other experts to understand how the technology is likely to be used in the real world and which of the scenarios are most plausible.
Why is it more difficult to understand the long-term impact of these technologies?
The longer-term picture is often much less certain. A common assumption is that earlier diagnosis automatically leads to better outcomes, but that is not always the case. In prostate cancer, for example, some early-stage cancers may never progress to cause significant health problems during a person’s lifetime. Detecting these cancers earlier may not necessarily improve survival outcomes and could potentially lead to unnecessary treatment. As a result, it can be challenging to predict exactly how changes in diagnosis will translate into long-term health outcomes, healthcare costs and resource use. These effects may only become clear once the technology has been used in routine practice and patients have been followed over time.
If there is so much uncertainty, why is early evaluation still important?
Even when evidence is limited, early evaluation plays an important role in supporting decision making. At the end of the day, it’s better to have some information than no information. It can certainly be challenging. Yes, there will still be a lot of questions that are not answered, and yes, some of those long-term components are going to be assumptions about what the potential impact or change is, if it’s not available from current data. But at least you know what the best- and worst-case scenarios are, the spectrum of change, so you know what to expect. Then you can put systems in place to collect the required data and have checkpoints to see which direction it’s likely to go in. You can understand what the key drivers of costs are likely to be, or the key pinch points in the pathway.
Ultimately, these evaluations do not eliminate uncertainty, rather they highlight where the uncertainties are and what impact this may have. This allows more informed decisions using the best available evidence, while creating a structured framework for learning and evaluation as the technology is implemented. This is valuable in a rapidly evolving field such as AI, where several technologies may offer similar capabilities and decision makers need a transparent and evidence-guided basis for prioritising investment and adoption.
The ability to do an initial vetting based on early modelling and then make a decision into the next step of rolling out a new technology, that’s the key impact. You’re informing decision making not only on whether these tools should be used further, but also on what evidence should be collected to ensure that in few years’ time, we can look back and check: was that a good decision? How is that technology impacting health outcomes or costs in the healthcare system?
The National Institute for Health and Care Excellence (NICE) Early Use Assessments are a really good example of this. The results tell you where the evidence gaps are, what evidence needs to be collected, and this can be used to guide study designs. YHEC have been involved with several of these assessments as an External Assessment Group for NICE as well as with companies directly.
What is a key takeaway message for people designing health-related AI technologies from your work on AI projects?
Engage as early as possible, connect your value propositions with actual economic methods so you’re able to test out your value proposition, and identify if that means you’re likely to produce a benefit. Identify where those evidence gaps are, so you’re well prepared on how to use funding to be able to plug in those gaps and bring a technology to the market that’s really well researched, that’s gone through all the rigorous processes and has done the due diligence, producing benefit to the healthcare system in the most responsible way.
What can YHEC provide to clients in this field?
YHEC helps organisations understand, test and demonstrate the value of digital and AI-enabled health technologies. We work with developers and innovators from an early stage to clarify their value proposition, identify evidence gaps, and determine which outcomes matter most to healthcare decision makers. Through early economic modelling and evidence-generation planning, we help clients understand what data they need and how best to collect it. Our goal is to help generate robust evidence that demonstrates value while supporting the responsible introduction of innovative technologies into healthcare systems.
To find out more about YHEC and our work, contact us.
