| Home > Publications database > Machine learning models for adverse drug reaction prediction. A systematic review and three-level multilevel meta-analysis. |
| Journal Article | DKFZ-2026-02408 |
; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ; ;
2026
Frontiers Media
Lausanne
Abstract: Adverse drug reactions (ADRs) are a leading cause of preventable hospitalization, yet the dominant pharmacovigilance paradigm remains reactive. Machine learning (ML) offers a data-driven alternative, but the evidence base has not been formally meta-analyzed under clinically realistic inclusion criteria and a dependence-respecting statistical framework. This is the first PRISMA 2020-compliant and PROBAST-screened multilevel meta-analysis of broad-spectrum ML-based ADR prediction. We aimed to quantify pooled discrimination of ML models for broad-spectrum ADR prediction and to test the influence of algorithm class, protein-target features, and outcome breadth.Following a Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 compliant, PROSPERO-registered protocol (CRD420250653686), we searched PubMed and Web of Science Core Collection from inception to 30 January 2025. A Scopus and IEEE Xplore top-up search on 2 February 2026 yielded no further records. Eligible studies applied ML to clinically validated data sources for multi-drug ADR prediction. Areas under the receiver operating characteristic curve (AUCs) were logit-transformed and pooled with a study-level DerSimonian-Laird random-effects model and a pre-specified primary three-level multilevel meta-regression (restricted maximum likelihood [REML]) using 30 model-level AUCs nested within 9 contributing studies.Eleven Prediction model Risk Of Bias ASsessment Tool (PROBAST) screened studies were included, of which 9 contributed quantitatively. The pre-specified primary three-level multilevel pooled AUC was 0.841 (95% confidence interval [CI] 0.789-0.883), with a study-level DerSimonian-Laird sensitivity estimate of 0.832 (95% CI 0.788-0.869, I-squared 90.7%). Leave-one-study-out estimates ranged 0.818-0.842. Neural networks did not statistically outperform logistic regression, random forests, or k-nearest neighbor. An apparent protein-target feature penalty (model-level p = 0.0002) was driven by one dominant study and disappeared in multilevel sensitivity analysis (adjusted p = 0.75).ML models showed moderate discrimination under Grading of Recommendations Assessment, Development and Evaluation (GRADE) low certainty. Within this small and methodologically heterogeneous evidence base, the pooled AUC is a descriptive summary and does not establish broadly generalizable performance across data sources, ADR definitions, feature-engineering strategies, or validation settings. The high heterogeneity indicates substantial variation in underlying performance across contexts. Standardized benchmarks, harmonized outcome taxonomies, mandatory external validation, and regulator-aligned prospective evaluation are prerequisites for clinical deployment.[https://www.crd.york.ac.uk/PROSPERO/view/CRD420250653686], identifier [CRD420250653686].
Keyword(s): PROBAST ; adverse drug reactions ; clinical applicability ; machine learning ; meta-analysis ; pharmacovigilance ; systematic review
|
The record appears in these collections: |