Intelligent Medicine
Volume 06 · Issue 02 · 2026
Published: April 28, 2026
Intell Med
Editorial
Open Access
Beyond multiple-choice questions: Rethinking evaluation frameworks for large language models for clinical medicineZehua Jiang, Haichao Chen, Yilan Wu, Yiming Qin, Chenyang Pei, Dian Zeng, Bin Sheng, Tien Yin Wong
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2026.01.001
Abstract
Large language models (LLMs) have demonstrated encouraging performance for medical natural language processing (NLP) tasks, approaching human-equivalent performance in some of the standard benchmarks, positioning them as game-changers in healthcare. However, there remains a persistent gap between high benchmark performance and clinical utility of NLP algorithms due to limitations of existing evaluation paradigms. Existing benchmarks tend to use static, task-specific benchmarks, and as a result, they do not capture the full dimension of complexity, safety, interpretability, and integration in workflow required for the safe deployment in the clinic. The editorial advocates for a shift in evaluation paradigms from narrow score-based metrics to a dynamic, multi-dimensional, and patient-centered system of clinical gatekeeping. The proposed framework integrates a four-phase process, including retrospective benchmarking, pilot testing, multi-center validation, and real-world monitoring, alongside a capability-task-behavior-value progression and a continuous human-in-the-loop feedback mechanism. This comprehensive strategy ensures not only technical robustness but also clinical relevance, ethical accountability, and adaptive improvement, transforming LLMs from experimental tools into reliable clinical partners for safer and more patient-centric healthcare delivery.
Research Article
Open Access
A pulmonary nodule is worth 8 × 8 × 8 words: Computed tomography-based 3D vision transformer predicts early-stage high-grade lung adenocarcinoma of micropapillary and/or solid subtypesYin Zhou, Yiyang Wang, Cheng Li, Shanshan Wang, Yuning Pan, Weiyu Shen, Chengbin Lin, Xinzhong Ruan, Xianwang Ye, Zhenya Zhao, et al.
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.11.001
Abstract
Background
Early-stage high-grade lung invasive adenocarcinoma (IAC) has poor prognosis and is hard to identify using conventional radiological assessment. Current reliance on postoperative histology to identify highgrade subtypes delays risk-adapted surgical planning. Three-dimensional (3D) vision transformers (ViTs) may improve prediction by modeling long-range dependencies in computed tomography (CT) scans. We aimed to develop and validate 3D-ViT and Swin Transformer (SwinT) for preoperative CT-based prediction of early-stage high-grade IAC subtypes (micropapillary/solid), benchmarking against ResNet.
Methods
A multicenter cohort of 1028 patients with surgically confirmed early-stage lung adenocarcinoma was divided into training (n = 806), validation (n = 100), and external test (n = 122) sets. 3D-ViT, SwinT, and ResNet models were trained on CT to classify nodules harboring high-grade histologic patterns. A novel decision-aid tool for IAC surgery was provided. Performance was evaluated using area under the curve (AUC), accuracy, sensitivity, specificity, and precision. Attention mapping was performed to interpret 3D-ViT decision-making.
Results
The 3D-ViT model achieved AUC values of 0.856 (95% CI: 0.845–0.877) (validation) and 0.806 (95% CI: 0.790–0.816) (testing), compared to 0.854 (95% CI: 0.841–0.872) (validation) and 0.760 (95% CI: 0.743–0.776) (testing) for the ResNet baseline. 3D-ViT showed balanced accuracy, sensitivity, specificity, and precision in the validation set. In external testing, 3D-ViT significantly outperformed ResNet in all metrics with P < 0.01. The SwinT-based AlignSen model from decision-aid tool prioritized sensitivity for high-grade IAC (88.0% validation, 93.4% testing), while maintaining specificity (76.0% validation, 54.1% testing) which significantly outperformed ViT and ResNet-based AlignSen models’ specificity (60.0% and 54% validation, 52.5% and 24.6% testing, respectively). Attention maps highlighted 3D nodule heterogeneity and peripheral irregularities.
Conclusion
The 3D-ViT model demonstrated robust accuracy and generalizability in predicting high-grade IAC subtypes using CT images. Integration of preoperative SwinT into clinical workflows may offer a viable alternative to intraoperative pathology subtyping, potentially reducing reliance on frozen sections while optimizing surgical planning.
Open Access
Cross-sectional comparative evaluation of US and China-developed large language models for bilingual coronary heart disease patient educationKaiyuan Liu, Long Cheng, Wanxin Wang, Runda Wu, Chenguang Li, Kang Yao, Junbo Ge
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.11.002
Abstract
Background
Patient education for coronary heart disease (CHD) is increasingly facilitated by large language models (LLMs). However, it remains unclear whether the origin of the models (United States vs. China) affects their performance in providing CHD-related patient education. This study aimed to systematically compare six mainstream LLMs when responding to common CHD-related patient questions presented in English and Chinese.
Methods
Between 1 and 15 February 2025, we posed 30 clinician-validated CHD questions—extracted from outpatient records—to six LLMs: GPT-4o, OpenAI o1, Gemini 1.5 (United States); and DeepSeek-R1, ERNIE Bot 3.5, Doubao (China). Each prompt was asked in English and Chinese. Each prompt was asked in both English and Chinese. Three blinded cardiologists rated every answer for accuracy, comprehensiveness, understandability, and empathy on a 4-point Likert scale. Ratings were analyzed using cumulative-link mixed models (CLMMs) with a logit link function, including fixed effects for Model, Language, and Dimension, as well as their interactions and random intercepts for Question and Rater. Type III likelihood-ratio χ2 tests assessed the main effects and interactions, followed by Holm-adjusted pairwise contrasts. Inter-rater agreement was quantified using Fleiss’ κ.
Results
Three cardiologists independently rated 360 bilingual responses with high inter-rater reliability (Fleiss’ κ = 0.821). In CLMMs, there were significant main effects of Model and Dimension, as well as a Model × Language interaction and Model × Language × Dimension. OpenAI o1 achieved the highest odds of superior ratings versus GPT-4o (OR = 4.45, 95% CI: 3.01–6.57, P < 0.001), followed by DeepSeek-R1 (OR = 1.32, 95% CI: 0.97–1.78, P = 0.038). Language-stratified contrasts showed that Chinese prompts increased comprehensiveness (OR= 1.48, 95% CI: 1.01–2.17, P = 0.045) and empathy (OR = 2.14, 95% CI: 1.47–3.11, P < 0.001) but reduced understandability (OR = 0.64, 95% CI: 0.42–0.98, P = 0.042). Gemini 1.5 excelled in Chinese (OR = 3.55, 95% CI: 2.35–5.38, P < 0.001), whereas DeepSeek-R1 favored English (OR = 0.64, 95% CI: 0.41–0.99, P = 0.046) and Doubao favored Chinese (OR = 1.64, 95% CI: 1.08–2.49, P = 0.020).
Conclusions
Model performance was strongly modulated by prompt language and evaluation dimension. Our benchmark offers practical guidance for clinicians, patients, and health-information providers choosing LLMs for bilingual patient education.
Open Access
Hemodynamic-morphologic machine learning model improves rupture risk stratification of intradural internal carotid artery aneurysms: A retrospective multicenter studyRong Zou, Shijie Zhu, Lifen Gan, Zhiwen Lu, Wei Lu, Jing Cai, Li Li, Lili Jiang, Jianping Xiang, Qinghai Huang
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.12.011
Abstract
Background
The risk of rupture associated with intradural internal carotid artery (ICA) aneurysms warrants considerable attention. We aimed to develop the first machine learning (ML) model that integrates standardized hemodynamic profiling with clinical and morphological data to stratify rupture risk in intradural ICA aneurysms.
Methods
We consecutively enrolled 511 intradural ICA aneurysms that underwent DSA examinations at four hospitals from July 2017 to July 2022. Utilizing the electronic medical record system and computational fluid dynamics of AneuFlow software, we extracted 10 clinical baseline characteristics, 13 morphological, and 12 hemodynamic features for the aneurysms. Subsequently, the risk of aneurysm rupture was stratified by random forest (RF), XGBoost (XGB), LightGBM (LGB), and logistic regression (LR) models. Data from three hospitals were used to develop the internal training cohort (n = 331) and internal validation cohort (n = 83), while data from the fourth hospital contributed to the external validation cohort (n = 97). The models’ performance across the three cohorts was evaluated using area under the curve (AUC), sensitivity, specificity, and the Youden index. Additionally, we determined the feature importance ranking of the ML models.
Results
The RF model achieved the highest AUC of 0.980 (95% CI: 0.969–0.989) in the internal training cohort. The AUC for the RF, XGB, LGB, and LR models in the internal validation cohort was 0.872 (95% CI: 0.792–0.929), 0.874 (95% CI: 0.794–0.931), 0.852 (95% CI: 0.769–0.914), and 0.827 (95% CI: 0.740–0.894), respectively. In the external validation cohort, the AUC for these models was 0.820 (95% CI: 0.729–0.891), 0.772 (95% CI: 0.675–0.851), 0.782 (95% CI: 0.686–0.859), and 0.782 (95% CI: 0.686–0.859), respectively. Moreover, the RF model achieved the greatest Youden index (0.517) in the external validation cohort, indicating superior discrimination ability. Hemodynamics accounted for 57% of the stratification power, with irregular geometry (nonsphericity index > 0.15) and microvascular inflammation markers (minimum wall shear stress < 0.3 Pa) identified as the key drivers.
Conclusion
The ML framework designed for intradural ICA aneurysms demonstrated strong risk stratification capabilities, allowing more timely and personalized clinical diagnosis and treatment.
Open Access
Artificial intelligence-based retinal vascular fractal dimension quantification and related factors: A retrospective studyChuan Zhang, Haotian Wu, Saiguang Ling, Zhou Dong, Li Dong, Ruiheng Zhang, Wenbin Wei, Lei Shao
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.12.002
Abstract
Background
Fractal dimension (Df) quantifies vascular network complexity and spatial filling density, providing a non-invasive biomarker for microvascular pathology assessment. We aimed to establish an artificial intelligence (AI)–powered framework for automated Df quantification and investigate its clinical implications in a general population of the Beijing Eye Study.
Methods
This retrospective study utilized data from the Beijing Eye Study 2011. Fundus images meeting quality criteria were processed with an AI algorithm to segment retinal vasculature and calculate Df. Multiple linear and logistic regression models were used to assess the associations of retinal vascular Df with systemic and ocular parameters.
Results
AI-based quantification of retinal vascular Df was successfully performed in 3,298 participants (95.1% of 3,468 participants), demonstrating high model performance: segmentation accuracy (0.9660), sensitivity (0.8879), specificity (0.9743), and intersection-over-union (IoU) (0.7110). The mean Df was 1.51 ± 0.09 (median: 1.53; interquartile range: 1.50–1.55). In multiple linear regression analysis, reduced Df was significantly associated with older age (standardized regression coefficient (sβ) = −2.346; P < 0.001), higher systolic blood pressure (sβ = −0.341; P < 0.001), smaller hip circumference (sβ = 0.518; P = 0.014), lower equivalent diopter (sβ = 3.589; P < 0.001), thinner retinal nerve fiber layer thickness (sβ = 0.390; P = 0.002), smaller arteriovenous ratio (sβ = 524.590; P < 0.001), and thinner subfoveal choroidal thickness (sβ = 0.071; P < 0.001). In the logistic regression model, the risk prevalence of hypertension increased with the decrease in Df (OR = 0.853, 95% CI: 0.651–0.801).
Conclusion
AI-powered Df analysis may be used as a novel quantitative platform for characterizing retinal microvascular morphological alterations.
Open Access
From radiology findings to artificial intelligence-powered impressions: A retrospective study on the comparative performance of recent large language modelsNanziba Tasneem, Christian B. van der Pol, Ambreen Zahoor, Nitin Juggath, Kyle McGowan, Cynthia Lokker, Ashirbani Saha
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.11.003
Abstract
Background
Large language models (LLMs), a revolutionary breakthrough in artificial intelligence, can be leveraged to automatically generate impressions for radiology reports, which usually require time, effort, and training. Our objective was to evaluate the performance of five recent LLMs (GPT–4, GPT–4o mini, Gemini 1.5–Pro, Gemini 1.5–Flash, and Llama 3.1) for impression generation.
Methods
In this retrospective study, 100 radiology reports were sampled (20 from each of the report-groups 0–400, 400–800, 800–1,200, 1,200–2,000, and 2,000–8,000 based on character count of the Findings section) from the publicly available "BioNLP 2023 report summarization" dataset (collected between 2001–2016, training subset of size 59,320 considered for sampling), sourced from PhysioNet. Then, each of the five LLMs was zero-shot prompted to generate impressions using the findings from the sample. Generated impressions were evaluated: (a) subjectively for coherence, comprehensiveness, conciseness, and medical harmfulness by two radiology fellows and a large reasoning model (LRM) Gemini 2.5–Pro, and (b) objectively using a composite accuracy metric including recall-oriented understudy for gisting evaluation (ROUGE)-1, bilingual evaluation understudy (BLEU) and cosine similarity, against the original human expert-generated impressions. The LLMs were ranked according to the percentage agreement ranking of subjective and composite scores. Statistical tests (Friedman and post-hoc Nemenyi tests) were used to assess inter-model differences.
Results
The top-ranked models were Gemini 1.5–Pro, GPT–4, and Gemini 1.5–Flash. Performance varied across models for both human and LRM raters (Friedman test: Human P <1.82×10−6; LRM P <9.10×10−40). Composite accuracy scores were significantly higher for the top three models (0.69, 0.68, and 0.68) versus others (0.65; Nemenyi P <1.11×10−16). The LRM aligned closely with human raters (2.15% complete disagreement) and identified all human-rated inaccurate impressions.
Conclusion
Gemini 1.5–Pro outperformed GPT–4, in terms of coherence, comprehensiveness, and medical harmfulness, at a lower cost. Human and LRM evaluations were generally consistent, though the LRM was more conservative.
Open Access
A retrospective validation of a federated machine learning framework (Hepa-FedBoost) for improving liver cancer computed tomography diagnosis across heterogeneous hospital networksChengquan Li, Xiaobin Feng, Dong Li, Jiahong Dong
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.08.003
Abstract
Background
The development of accurate artificial intelligence (AI) models for liver cancer diagnosis using contrast-enhanced computed tomography (CT) is often hindered by patient privacy regulations and considerable data variations between hospitals. These variations—in CT scanners, patient populations, and disease prevalence—can reduce the performance of standard collaborative training methods such as federated learning (FL). This retrospective study evaluated whether a novel machine learning framework, Hepa-FedBoost, can overcome these challenges to improve diagnostic accuracy for liver cancer classification across a simulated multi-center network without sharing raw patient images.
Methods
We developed a new benchmark dataset to better represent real-world clinical diversity. This was done by merging 23,583 CT images of 11 abdominal organs from the public OrganCMNIST subset of the MedMNIST v2 dataset with 11,000 liver lesion images from 500 patients with histopathology-confirmed liver cancer. These cancer-related data were obtained retrospectively from the National Hepatobiliary Standard Database of China (initiated by Beijing Tsinghua Changgung Hospital, with data collected since December 2019). The use of data from the public and private sources did not require authorization from the patients or owners. This combined dataset of 34,583 images was used to simulate a network of 12 hospitals with significant imbalances in data quantity and class labels. The Hepa-FedBoost framework, a type of clustered FL model, was trained on this network. The model works by having each simulated hospital share only compact, anonymized data summaries (prototypes) instead of model parameters or raw data. A central server uses these summaries to group hospitals with similar data and guide the training process to improve overall accuracy. The primary endpoint was the macro area under the receiver operating characteristic curve (mAUC). Furthermore, we compared Hepa-FedBoost’s performance against 6 other established FL methods.
Results
On the simulated hospital network, Hepa-FedBoost demonstrated superior diagnostic performance, improving the liver cancer classification mAUC by 4.9 absolute points and overall accuracy by 2.4 absolute points compared with the next best method. It reached a threshold of clinical-grade accuracy in just 17 training rounds, 43% faster than the standard FedAvg approach, while reducing the total data transmission required by approximately half (to 800 MB).
Conclusion
The Hepa-FedBoost framework markedly improved the accuracy and efficiency of multi-center liver cancer classification on CT images. By enabling robust collaborative model training while maintaining patient privacy and minimizing IT resource requirements, it represents a practicable AI solution for real-world hospital networks.
Open Access
A retrospective magnetic resonance imaging-based radiomics study for predicting lymph node regression status following neoadjuvant therapy in patients with locally advanced rectal cancerTianxu Ma, Di Hao, Ruiqing Liu, Wentao Xie, Mingyu Yang, Junhao Zhang, Jingnong Liu, Zhenying Xu, Rujie Jiang, Xuejun Liu, et al.
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.12.005
Abstract
Background
Lymph node metastasis (LNM) status serves as a key prognostic marker in patients with locally advanced rectal cancer (LARC). Neoadjuvant chemoradiotherapy (nCRT) can influence lymph node status by inducing regression, with treatment responses exhibiting substantial variability among patients. We aimed to develop and validate a random forest-based radiomics model for non-invasively predicting lymph node regression status following neoadjuvant therapy in patients with LARC using pretreatment magnetic resonance imaging (MRI).
Methods
This retrospective study included 285 patients with LARC who were treated with nCRT followed by elective resection at Qingdao University Affiliated Hospital between October 2019 and October 2023. Patients were randomly allocated to training and testing sets in a 7:3 ratio. Baseline characteristics including gender, age, body mass index (BMI), and other 13 clinical variables were compared between the groups, with no significant differences observed (P >0.05). High-resolution T2-weighted MRI scans of the rectal region were collected preoperatively. Regions of interest corresponding to primary tumor lesions were manually delineated on the MRI scans, followed by radiomic feature extraction. Feature selection was performed using the least absolute shrinkage and selection operator (LASSO) regression model. A machine learning random forest predictive model was subsequently constructed, incorporating selected radiomics features and clinical variables.
Results
Model performance for predicting lymph node status was assessed using receiver operating characteristic curve analysis, area under the curve (AUC), and calibration curves. Four features were selected from 1,051 radiomic features to construct a radiomics model, achieving an AUC of 0.743 in the test set. Four features were also extracted from 19 clinical parameters to develop a clinical data model, with an AUC of 0.727. Integrating radiomic features and clinical data yielded a combined model with superior performance in the test set (AUC 0.794).
Conclusion
The radiomics model derived from pretreatment rectal MRI in patients with LARC demonstrated strong predictive capability for assessing metastatic lymph node responses to nCRT.
Open Access
Microscopic image processing platform for multi-class cell segmentation using deep learningYuzhou Wang, Xiaojie Li, Frank Kulwa, Xiaoyan Li, Shuochen Tai, Shuaiyi Tian, Kunyang Teng, Marcin Grzegorzek, Xinyu Huang, Tao Jiang, et al.
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.08.002
Abstract
Background
From lung cancer and heart disease to rare disorders, research on almost every disease is speeding up. Microscopic cell image analysis is an important area of medical research. As the first step in analysis, segmentation is a significant clinical concern. However, many complexities, such as variations in cell size or shape, overlapping regions, potential poor contrast, and background noise, make automated segmentation of microscopic images a complicated problem. Moreover, there is a notable deficiency in image processing systems that are both user-friendly and capable of delivering credible results.
Methods
This paper proposed a microscopic cell processing platform to enable efficient and accurate microscopic image analysis. First, the 2018 Data Science Bowl (DSB2018) dataset from the cell segmentation competition of Kaggle in 2018 is grouped into training, validation, and test sets. Then, U-Net+ + incorporated with the watershed algorithm was used for microscopic cell image segmentation tasks. Third, a graphics user interface based on QT was designed to display the segmentation process, fine-tune the model according to clinical needs, and automatically generate diagnostic conclusions.
Results
The average of intersection over union, precision, recall, and F1-score achieved 0.846, 0.908, 0.925, and 0.917, respectively, which were highly satisfactory results, indicating the efficacy of this platform. Moreover, the mean error of the watershed algorithm achieved 0.113, a margin acceptable in clinical diagnostics. Compared with traditional methods, the proposed method significantly improved the performance.
Conclusion
With the efficient segmentation based on deep learning and accurate quantitative analysis for the results, this microscopic image processing platform may outperform many existing medical analysis systems and have applications in the field of auxiliary medical diagnosis.
Open Access
One-stop automated diagnostic system for active sacroiliitis in three-dimensional magnetic resonance images using artificial intelligent models: a retrospective studyXiaojian Ji, Zhuofeng Li, Lulu Zeng, An’an Wang, Jing Dong, Lei Sun, Shiwei Yang, Jian Zhu, Feng Huang, Tao Li, et al.
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.07.005
Abstract
Background
Axial spondyloarthritis (axSpA) can potentially progress to ankylosing spondylitis, and diagnostic delay may lead to irreversible structural damage. Although sacroiliac joint magnetic resonance imaging (MRI) can detect bone marrow edema (BME) for early diagnosis, the complex etiologies and reliance on expert interpretation impede efficient identification in chronic low back pain populations. Therefore, we developed an artificial intelligence (AI) system that integrates MRI analysis and automates report generation to facilitate the detection of active sacroiliitis and BME quantification.
Methods
This retrospective study analyzed 691 patients (540 with axSpA and 151 with non-spondyloarthritis) from the Chinese People’s Liberation Army General Hospital (2011–2023). Data were split into training and testing cohorts (4:1 ratio), with five-fold cross-validation applied to the training set. The system comprises four modules: (1) image preprocessing (intensity normalization), (2) coarse-to-fine 3D U-Net-based quadrant segmentation, (3) ResNet18-driven edema recognition (depth/intensity classification), and (4) diagnostic report generation using the QWEN large language model.
Results
Among 691 patients (335 active sacroiliitis-positive and 356 negative), the AI system achieved Spondyloarthritis Research Consortium of Canada (SPARCC) scores of (15.12±10.73) vs. (0.88±1.13) in positive versus negative groups, respectively. The quadrant segmentation module achieved DICE similarity coefficients above 0.7 on the training, validation, and testing datasets. The edema inflammation, depth, and intensity classifiers exhibited good performance on the testing dataset, with balanced accuracies of 77.54%, 76.26%, and 80.91%, and area under the curve values of 0.85, 0.87, and 0.92, respectively. The intraclass correlation coefficient between SPARCC scores by our system and those by rheumatologists was 0.84 on the testing dataset. At the patient-level, our system achieved 90.81% sensitivity for diagnosis of active sacroiliitis on the testing dataset.
Conclusion
The developed automated system enhances axSpA diagnostic efficiency by automating the identification of active sacroiliitis and enabling quantitative assessment of BME.
Review
Open Access
Artificial intelligence-driven multi-cancer screening: Achievements, challenges, and future prospectsYubei Huang, Shuxue Lu, Hongji Dai, Wei Wang, Ben Liu, Maiheliyan Abudoumijiti, Qianyun Jin, Jie Wu, Kexin Chen, Yaogang Wang
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2026.02.001
Abstract
Malignant tumors pose a significant global health challenge. While established screening methods exist, they are largely site-specific, limiting comprehensive prevention. Multi-cancer screening, which detects multiple cancers simultaneously, offers substantial advantages including optimized sample utilization, reduced participant costs, enhanced efficiency, and improved resource allocation. Artificial intelligence (AI) has emerged as a transformative technology, revolutionizing healthcare by analyzing complex biomedical data and enhancing diagnostic accuracy. Integrating AI into multi-cancer screening holds immense potential for advancing early cancer detection. This review provides a comprehensive overview of AI-driven multi-cancer screening. We examine foundational technologies and current applications, including biomarker data and medical imaging data analysis, as well as core AI techniques like machine learning, deep learning, natural language processing, explainable AI and essential preprocessing steps. We assess key technical bottlenecks such as data sparsity, model generalizability, false positives and false negatives, and affordability, alongside solutions like transfer learning, federated learning, and Bayesian optimization. Additionally, we highlight clinical validation, regulatory approval, and ethical considerations for multi-cancer screening. Furthermore, we explore future prospects, envisioning enhanced accuracy and expanded coverage, deeper multi-modal data fusion, personalized and dynamic screening, and intelligent decision-support systems with improved accessibility. We also outline targeted recommendations for developing countries conducting AI-driven multi-cancer screening, building on global best practices while adapting to local realities including those in China. This review offers a forward-looking perspective on how AI will evolve multi-cancer screening into a more personalized, dynamic and accessible cornerstone of cancer prevention.
Open Access
The performance of Covidence: An artificial intelligence-based tool for title and abstract screening in a breast cancer evidence-based clinical practice guidelineXiaomei Yao, Ashirbani Saha, Ashley Low, Shakil Ahmed, Samya Ali, Aditya Misra, Mariam Abdelmalek, Sharan Saravanan, Jonathan Sussman
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.12.008
Abstract
Background
Clinical practice guidelines (CPGs) support evidence-based care but are time-consuming to develop. We aimed to compare artificial intelligence (AI)-assisted versus manual title and abstract screening (Stage I) in Covidence using data from a published breast-cancer CPG.
Methods
This systematic review (SR) included 8,774 articles identified through a medical literature search, after duplicate removal. Three article subsets (n = 500, 1,000, and 2,000) were randomly selected from 8,774 articles to perform 30, 30, and 10 trials, respectively, independent Stage I AI-assisted trials. The primary outcome of each trial is workload savings achieved through AI-assisted identification of 95% and 100% relevant articles (e.g., sensitivity), and 100% of finally-included articles. The secondary outcome is missed finally-included articles when the sensitivity of 95% was reached for each subset.
Results
At 95% sensitivity, 100% relevant articles and 100% finally-included articles were identified, median (minimum, maximum) workload savings were 40.7% (4.4%, 59.4%), 25.0% (0.4%, 55.2%), and 57.6% (6.2%, 76.4%) for n = 500; 38.3% (6.2%, 54.0%), 17.3% (0.0%, 39.1%), and 63.9% (0.4%, 77.5%) for n = 1,000; 16.6% (10.8%, 41.8%), 4.4% (0.3%, 20.9%), and 17.9% (0.8%, 64.6%) for n = 2,000, respectively. Covidence’s performance does not improve as the size of the subsets increases for a CPG with multiple complicated research questions. A potential positive correlation between the proportion of relevant articles in initial training of Covidence and workload savings at Stage I across all 70 trials. At 95% sensitivity, five trials missed one article (n = 500); two trials missed two articles and one trial missed one article (n = 1,000); and one trial missed three articles and five trials missed one article (n= 2,000) .
Conclusion
AI-assistance in Covidence for Stage I screening showed promise and pitfalls in the SR for a breast cancer CPG on a complex topic. Further prospective research is needed to better understand the performance of AI-assistance in Covidence and intricacies of CPG topics.
Open Access
Impact of telemedicine on chronic disease patients: An overview of systematic reviewsKun Zhang, Yanhua Chen, Peicheng Wang, Yanrong He, Yanrong Du, Hailun Liang, Weiguo Zhu, Xiao Long, Leiyu Shi, Jiming Zhu
Intelligent MedicineVol.06,No.022026
DOI: 10.1016/j.imed.2025.01.002
Abstract
Telemedicine has gained prominence in managing chronic diseases, particularly post-COVID-19. However, interventions often lack consistent categorization based on patient engagement levels, and evidence of efficacy remains heterogeneous. This review synthesized findings from systematic reviews to evaluate the benefits and challenges of telemedicine for chronic disease management. A systematic search of the Medline, EMBASE, and Web of Science databases was conducted to identify systematic reviews with meta-analyses published between January 2022 and December 2024. The Triple Aim framework was used to assess the impact of telemedicine on clinical outcomes, patient experience, and healthcare costs. A framework for digital patient engagement levels was developed to explore the relationship between engagement levels and intervention effectiveness. A total of 81 systematic reviews met the inclusion criteria. Among these, 31% focused on cardiovascular diseases, 30% on diabetes, 28% on cancer, 4% on chronic lung diseases, and 7% on general chronic conditions. Intervention levels were categorized as consultation (9%), involvement (42%), partnership/shared leadership (13%), or unspecified (36%). Clinical outcomes were reported in 61 reviews, patient-reported outcomes in 23, and economic impacts in only 4 reviews, indicating substantial research gaps. Telemedicine enhances patient care experiences and population health, demonstrating substantial benefits in managing diabetes, cancer, cardiovascular, and chronic lung diseases. It holds considerable potential for reducing healthcare costs. However, further research is needed to evaluate the feasibility and effectiveness of telemedicine across different engagement levels. Furthermore, standardized methodologies and improved tools are essential to advance telemedicine research and implementation.
CURRENT ISSUE

Download Cover(PDF) Download TOC(PDF)
