Welcome to YHEC’s publication hub

Our latest research, all in one place. Browse our collection of journal articles, reports and conference proceedings to see how we’re contributing to HEOR research. Remember to: 

  • Filter by service, therapeutic area, or geography to narrow your results.
  • Search directly for keywords or specific titles to find what you need instantly.
Conference proceeding

Integrating Large Language Models Into an Existing Review Process: Promises and Pitfalls

YHEC authors: Mary Edwards, Lavinia Ferrante di Ruffano
Publication date: November 2024
Conference: ISPOR EU, Barcelona
Type of conference proceeding: Poster

Abstract

OBJECTIVES: The recent development and rise in accessibility of large language models (LLMs) has generated excitement around their possibilities for reducing the resource burden of conducting reviews. Following testing, we assessed the cost, accuracy, and accessibility of LLMs to reviewers, and consider what types of reviews LLMs are currently best suited to assist with.

METHODS: We conducted internal testing of a LLM, Claude 3 Opus, via the chat interface. We used the tool to conduct high level data extraction for a targeted review, highly granulated extraction for a systematic review, and risk of bias assessment of RCTs.

RESULTS: The LLM via a chat interface was highly accessible, inexpensive, and saved significant time in conducting high level qualitative data extraction for a pragmatic review. Outputs were standardized and easy to manipulate and integrate into our existing work process. Extracting accurate granular data for a systematic review proved more difficult, with the model failing to interpret complexities of patient flow, struggling to respond accurately to lengthy, detailed prompts, and the subsequent checking, correcting, and formatting outweighing any time saved. The model identified some relevant content for conducting risk of bias assessment with the Cochrane RoB 1 tool, although lacked context, and human judgement was needed for final decision making.

CONCLUSIONS: LLM chat interfaces offer significant time savings for pragmatic reviews, although copyright issues exist in uploading published papers for synthesis. Optimal performance for systematic reviews is unlikely to be achieved without fine tuning a version of the model with archive data. This process is currently costly, commercial confidentiality must be considered, and the skill set required is outside the scope of many review teams. Developers should ensure that any LLM based tools for reviewing can be integrated into clients' existing processes with the use of standardized import and export formats such as CSV or RIS.

Conference proceeding

Investigating Input Correlation in Probabilistic Sensitivity Analysis

YHEC authors: Matthew Taylor, Erin Barker, Harriet Fewster, Emily Gregg
Publication date: November 2024
Conference: ISPOR EU, Barcelona
Type of conference proceeding: Poster

Abstract

OBJECTIVES: Probabilistic sensitivity analysis (PSA) is used to characterize uncertainty in cost-effectiveness models. Inputs in PSA are often varied independently even when they may be correlated. This study investigated the effects of input correlation on PSA outputs.

METHODS: A Markov model was developed using R and Shiny to compare a hypothetical treatment and comparator. Three options were built into the model: no correlation (inputs varied independently); part correlation (correlation within but not between costs, utilities and transition matrices); and full correlation (correlation between all inputs). Inputs which improved the incremental cost-effectiveness ratio (ICER) were positively correlated with each other and negatively correlated with inputs which worsened the ICER, and vice versa. The treatment cost and the number of health states and health state costs were varied in scenario analyses to determine the circumstances in which correlation had the largest impact.

RESULTS: While the ICER was comparable across all correlation options, the likelihood of cost-effectiveness differed substantially from 61% to 93%. In all scenarios, the 'no correlation' option displayed the most certain likelihood (closest to either 0 or 1) of cost-effectiveness, while the least certain was produced by the full correlation option. The greater the complexity of the model (i.e. the greater the number of health states), the more pronounced the difference between correlating or not correlating inputs. Counterintuitively, correlating inputs increases uncertainty because it allows for a greater number of 'extreme' scenarios to be generated, whereas allowing independent generation of large numbers of inputs tends to lead to a 'cancelling out' effect. This effect is most pronounced when the ICER is moderately close to the willingness-to-pay threshold.

CONCLUSIONS: This analysis demonstrates that input correlation can have a substantial impact on the level of certainty in model outputs, and by ignoring this, the model may be over- or under-stating the true level of confidence.

Conference proceeding

Large Language Models for Data Extraction in a Systematic Review: A Case Study

YHEC authors: Mary Edwards, Lavinia Ferrante di Ruffano
Publication date: November 2024
Conference: ISPOR EU, Barcelona
Type of conference proceeding: Poster

Abstract

OBJECTIVES: A typical systematic review includes extraction of highly granulated data in a standardized format, a resource intensive part of the review process. We investigated whether the chat interface to a large language model (Claude 3 Opus) could provide time savings in extracting such data while retaining the accuracy necessary for a systematic review.

METHODS: A data extraction sheet from a completed review of biologic treatments was selected. A set of prompts was designed to obtain details of the methods, interventions, and populations assessed by three of the included studies. Each paper was uploaded individually, and the results were copied into the original data sheet and compared with those produced and checked by two independent human reviewers. Testing of outcome extraction was also conducted.

RESULTS: In order to produce suitably formatted granular data, prompts were detailed and consistently structured. Although the model successfully extracted details of the intervention (including dose, scheduling and duration of treatment) and population (including age, gender, duration of disease, and exon10 variants) assessed in each arm, it struggled to interpret complex patient flow through the studies. Primary outcomes in the ITT population were successfully extracted, but extraction of secondary outcomes, subgroups, and outcomes at different timepoints proved much less reliable.

CONCLUSIONS: While chat interfaces to LLMs may provide some time savings in extracting basic study data, such interfaces do not lend themselves to the detailed prompts required for successful extraction of more complex data. Accessing a LLM outside a chat interface can be costly, and requires a skillset not possessed by the majority of reviewers; organizational investment may therefore be needed to facilitate productive access. Fine-tuning using archive data also raises issues of commercial confidentiality. A market is emerging for companies providing affordable access to a protected model, accessible only to the customer and fine-tuned to their needs.

Conference proceeding

Large Language Models for Data Extraction in a Targeted Review: A Case Study

YHEC authors: Mary Edwards, Lavinia Ferrante di Ruffano
Publication date: November 2024
Conference: ISPOR EU, Barcelona
Type of conference proceeding: Poster

Abstract

OBJECTIVES: Accurate, consistent extraction and presentation of data is a time-consuming process. Large language models (LLMs) accessed via a chat interface require minimal user training and perform tasks without any setup overheads beyond an initial phase of prompt engineering. We assessed the chat interface to Claude 3 Opus for accuracy, consistency, presentation of data, and time savings in the context of high-level extraction for a targeted review.

METHODS: A targeted review was conducted to investigate disparities in patient characteristics in the diagnosis and treatment of one specific indication. We used the chat interface to Claude to extract data from 30 papers, with a human reviewer checking all data points. Study and population details were extracted, plus brief details of any study results or discussion regarding disparities in diagnosis or treatment. Papers were uploaded in pairs to minimize prompts, with the model explicitly tasked with labelling each set of data points with the name of the relevant paper.

RESULTS: Data were consistently extracted in a suitably structured format. Eleven papers required no edits to the data. Five papers required minimal edits, and nine papers contained minor errors or omissions in the data. One paper was extracted correctly but the answers reported by the model also contained additional data drawn from the second paper of the pair. Another pair of papers was extracted by the model and mislabeled, with data for each paper labelled with the file name of the other paper. Following this error, PDFs were uploaded singly.

CONCLUSIONS: Even allowing time for human checking and minor correction of the extracted data, use of the model enabled extraction and checking of 30 papers in a single day. Access to LLMs via a chat interface is typically relatively inexpensive and can offer significant resource savings in the context of suitable reviews.

Conference proceeding

Large Language Models for Risk of Bias Assessment: A Case Study

YHEC authors: Mary Edwards, Lavinia Ferrante di Ruffano
Publication date: November 2024
Conference: ISPOR EU, Barcelona
Type of conference proceeding: Poster

Abstract

OBJECTIVES: Risk of bias assessment (RoBA) of primary studies is a key part of any systematic review. As a repetitive and structured task, RoBA would initially appear to be well suited to automation or AI support. We assessed the chat interface to Claude 3 Opus for accuracy, consistency, presentation of data, and time savings in the context of RoBA of RCTs for a systematic review.

METHODS: Six RCTs were selected from three reviews conducted by our consultancy over the past five years. Following an initial prompt engineering phase using a report of a seventh RCT, the LLM was used to: 1. Conduct fully automated assessment of each paper using Cochrane RoB 1 tool (Method 1), and 2. Supply information only to facilitate joint human / LLM assessment (Method 2) using the same tool. The results were compared to fully human assessment (Method 3).

RESULTS: Method 1 resulted in very brief answers, with little supporting information provided by the model. Asking for supporting information only (Method 2) resulted in better quality and more complete data, although no judgement was made by the LLM. The agreement percentage between the three methods was mixed, ranging from 16.7% to 100% across domains. The lower agreement level was seen on questions relating to treatment allocation, incomplete outcome data and other sources of potential bias. In these instances, the LLM appeared to have misinterpreted the questions, resulting in different answers to the human assessor. However, there were also a few occasions where the LLM picked up information that the human did not.

CONCLUSIONS: Using LLMs for fully automated RoBA is not recommended at this stage, as such models can misinterpret questions and provide limited or incorrect justification for judgments. However, with suitable prompt engineering, and fine tuning using existing RoBA data, the performance of these models may improve with time.

Conference proceeding

Measuring Environmental Outcomes is More Complex Than we Think

YHEC authors: Matthew Taylor
Publication date: November 2024
Conference: ISPOR EU, Barcelona
Type of conference proceeding: Poster

Abstract

OBJECTIVES: Accounting for the environmental impact of healthcare is an important issue. In recent years, there has been a substantial increase in the number of economic evaluations that also report an environmental outcome (usually carbon emissions). However, the true impact of changes in the care pathway is more complex that is currently being reported.

METHODS: A case study comparing two separate surgical devices is used to demonstrate the various consequences of a healthcare decision on the environment. Two devices are compared, with different levels of cost, health outcomes and environmental outcomes.

RESULTS: Patients who receive Device A incur costs of €3,820, compared with €5,410 for Device B. Device A is more effective (121 months of life expectancy compared with 119 for Device B) and is less harmful to the environment (25kg of CO2 emissions, compared with 30kg for Device B). Based on these outcomes, Device A would normally be considered 'dominant', since it has better outcomes on all three metrics. However, because Device A results in two additional months of life expectancy, overall CO2 emissions will increase (i.e. including non-healthcare-related emissions). Based on an estimate of 12.7 tons of emissions per year of life, the emissions associated with two additional months would be 2,117kg, vastly outweighing the short-term reduction associated with Device A. In addition, the money saved by Device A will be reinvested into other healthcare, which will increase healthcare emissions by an estimated 0.235kg.

CONCLUSIONS: Measuring the true environmental impact of medicines is complex, and ignoring key indirect consequences could results in suboptimal (or even counter-productive) decisions. Moving healthcare systems towards a net zero target will require wider thinking that simply choosing between individual therapies.

Conference proceeding

Reappraisal Following the Loss of Medicine Patent in HTA Guidelines

YHEC authors: Matthew Taylor, James Mahon
Publication date: November 2024
Conference: ISPOR EU, Barcelona
Type of conference proceeding: Poster

Abstract

OBJECTIVES: We aimed to explore whether a proactive approach should be used to update HTA guidance when: (i) A medicine that had originally received a negative recommendation becomes unbranded (and has a lower price), (ii) An original medicine was not appraised but is now off patent, and (iii) A comparator treatment in an original appraisal is now unbranded.

METHODS: We undertook a review of existing National Institute for Health and Care Excellence (NICE) appraisals to explore the practical implications of using existing evidence from previous appraisals to assess whether, and how, this type of information could be repurposed to facilitate a rapid assessment of unbranded technologies. We also undertook a detailed assessment of seven appraisals where an intervention within the pathway had since become unbranded.

RESULTS: Of 74 appraisals that had originally received a negative recommendation, 8 of the originator drugs were later recommended because of later appraisals, and 60 led to changed guidance. In 6 cases, recommendations stood, despite the medicine becoming unbranded. Of the 7 detailed assessments, in 6 cases we found that a rapid re-assessment would not possible because: (i) The pathway had changed, (ii) The clinical effectiveness evidence for the originator was considered poor, or (iii) The model had been deemed unreliable due to structural flaws, implausible assumptions or errors that the ERG had not been able to rectify. Only the documentation from 1 of these STAs had the potential to be repurposed to inform a rapid assessment of an unbranded version of the originator.

CONCLUSIONS: In most cases, current approaches to HTA were able to deal with losses of patent, due to the regular appraisal of disease areas when new medicines reach market. In only a small proportion of cases would a proactive approach (i.e. at the point a medicine loses its patent) be useful.

1 11 12 13 14 15 80