Assessment of the Diagnostic Performance and Clinical Impact of AI in Hepatic Steatosis: Systematic Review and Meta-Analysis

Authors
Category Systematic review
JournalJ. Med. Internet Res.
Year 2025
BACKGROUND: The global rise of metabolic associated fatty liver disease (MAFLD) reflects the urgent need for accurate, non-invasive diagnostic approaches. The invasive nature of liver biopsy and the limited sensitivity of ultrasound (US) in detecting early steatosis highlight a critical diagnostic gap. Artificial intelligence (AI) has emerged as a transformative tool, enabling the automated detection and grading of hepatic steatosis (HS) from medical imaging data. OBJECTIVE: This systematic review and meta-analysis endeavors to quantitatively evaluate the diagnostic performance of AI models for HS, to comprehensively explore sources of inter-study heterogeneity, and to provide an in-depth appraisal of their clinical applicability, translational potential, and the major barriers impeding widespread implementation. METHODS: PubMed, Cochrane Library, Embase, Web of Science, and IEEE Xplore databases were searched until September 24, 2025. Studies employing AI for the diagnosis of HS were included if they met the predefined PIRT framework and provided sufficient extractable data. Diagnostic performance indicators, including sensitivity, specificity, and the area under the summary receiver operating characteristic (SROC) curve (AUC), were extracted and quantitatively synthesized. Meta-analyses were conducted using a bivariate random-effects model. The methodological quality and risk of bias of the included studies were evaluated using the QUADAS-2 tool. Heterogeneity was comprehensively assessed through the I² statistic, bivariate boxplots, 95% prediction intervals (PIs), and threshold effect analysis. The clinical applicability of the diagnostic models was further examined using Fagan's nomogram and likelihood ratio tests. RESULTS: 36 eligible studies were identified, of which 33 (comprising 62 cohorts) were included in the quantitative synthesis. Pooled estimates demonstrated excellent diagnostic accuracy of AI models, with a summary sensitivity of 0.95 (95% CI: 0.93-0.96), specificity of 0.93 (95% CI: 0.91-0.94), and an AUC of 0.98 (95% CI: 0.96-0.99). Clinical applicability analysis (LR+ >10, LR- <0.1) supported AI's strong potential for both confirming and excluding HS. However, substantial heterogeneity was observed across studies (I² >75%). According to QUADAS-2, a high risk of bias, particularly in the Patient Selection domain (44.4%), may have contributed to the overestimation of real-world performance. Subgroup analyses showed that deep learning (DL) models significantly outperformed traditional machine learning (ML) approaches (AUC: 0.98 vs. 0.94). Models using US or histopathology as reference standards both achieved high diagnostic accuracy (AUC: 0.98). Retrospective design (AUC: 0.98), transfer learning (TL), and the use of public datasets (AUC: 0.99) were associated with superior performance but also contributed to inter-study heterogeneity. CONCLUSIONS: AI demonstrates remarkable potential for noninvasive screening and assessment of HS, especially in primary care. Nonetheless, clinical translation remains limited by substantial performance variability, predominance of retrospective designs, lack of rigorous external validation, and practical barriers such as data privacy and workflow integration. Future studies should prioritize prospective multicenter trials, standardized development pipelines, and robust external validation to bridge the gap between current evidence and clinical application. The key innovation of this review lies in establishing a unified, modality-agnostic analytical framework that integrates evidence beyond single-modality evaluations.
Epistemonikos ID: 9f87f90a0ee7bfa7882cec44b02ec5ab431133b2
First added on: Dec 01, 2025