Intelligent Medicine
Volume 06 · Issue 01 · 2026
Intell Med
- Sections
- Guideline & Standard
- Editorial
- Research Article
- Review
- Ethical, Legal, and Social Implications
Dry eye, a common eye disease globally, poses significant challenges to clinical diagnosis and management due to its complex pathogenesis and high incidence rate. The development of artificial intelligence (AI) technology has provided new opportunities for the analysis and auxiliary diagnosis of dry eye imaging. This expert consensus focuses on the classification and annotation methods of dry eye imaging, in line with the application needs of AI technology. It summarizes the scope and tasks of research on the classification and annotation of dry eye imaging and provides detailed standards for the principles and methods of classification and annotation of major imaging modalities, including lipid layer of the tear film, tear meniscus height, tear film breakup time, corneal fluorescein staining, and meibomian gland images. It also clarifies the tools and processes for classification and annotation. The consensus proposes systematic quality control requirements, including annotation consistency assessment, multi-round review, and data cleaning methods. Finally, the consensus summarizes the current challenges and proposes targeted solutions. The launch of this consensus aims to provide high-quality data support for the development of AI in dry eye, enhance the application effects of AI in dry eye diagnosis, disease monitoring, and personalized treatment, and offer scientific references and technical support for clinical and research applications of AI in the field of dry eye.
This editorial examines dynamics-driven approaches for analyzing temporal patterns in medical big data to predict critical disease transitions and enable personalized medicine. We present dynamic network biomarker theory, which identifies disease tipping points by monitoring increased fluctuations and correlations within biomolecular networks, achieving over 80% accuracy in predicting tumor progression and enabling early detection of influenza infections. Advanced extensions including individual-specific edge-network analysis enable single-sample assessments with area under curve (AUC) values exceeding 0.9, eliminating the need for control groups and making real-time clinical application feasible. We discuss the integration of hybrid mechanistic data-driven models that combine domain knowledge with deep learning capabilities, including physiology-informed long short-term memory (LSTM) networks achieving mean absolute error of 35.0 mg/dL versus 79.7 mg/dL for traditional simulators in type 1 diabetes management, and temporal graph neural networks that improve diagnosis prediction accuracy on MIMIC-III datasets. Key challenges include data heterogeneity leading to false positives in critical transition detection, fundamental difficulty in distinguishing correlation from causation without experimental validation, interpretability limitations of black-box models that erode clinical trust, and ethical concerns regarding privacy in federated learning and algorithmic bias in underrepresented populations. Future directions emphasize multimodal integration combining omics profiles, imaging sequences, and wearable sensor streams using advanced transformers and graph neural networks, alongside causal inference methods employing instrumental variables and counterfactual simulations. These dynamics-driven approaches aim to augment rather than replace clinical expertise, providing early warning signals for proactive intervention while preserving irreplaceable human judgment in complex medical decision-making, ultimately transforming healthcare from reactive treatment to preventive care.
Optical coherence tomography angiography (OCTA) is a novel, non-invasive imaging technique that enables capillary-level visualization of the retinal vasculature, offering critical insights into various ophthalmic diseases. Accurate segmentation and quantitative analysis of microstructures—specifically the retinal vascular network (RVN) and foveal avascular zone (FAZ)—are essential for diagnosis and treatment planning. We aimed to develop an artificial intelligence-based system to automate segmentation and analysis of OCTA microstructures using a deep learning framework.
OCTA images were retrospectively collected from January 2020 to December 2022, comprising 2 public datasets (Retinal OCTA vessel Segmentation-1 (ROSE-1) and OCTA_3M) and a newly constructed clinical dataset, named Fully Annotated Retinal OCTA Segmentation dataset (FAROS), acquired at Zhongshan Ophthalmic Center. The FAROS dataset included 40 en face OCTA images from 40 eyes (20 healthy and 20 with retinal diseases such as diabetic retinopathy (DR), age-related macular degeneration, and retinal vein occlusion). Additionally, a separate clinical dataset containing 20 eyes with DR was enrolled to verify the clinical consistency of the observed parameter trends. To accurately segment the RVN and FAZ, we first introduced an innovative OCTA microstructure segmentation network (RS_Unet3+) by combining an encoder-decoder-based architecture with full-scale skip connections and the split-attention-based residual network ResNeSt, paying special attention to OCTA microstructural features while facilitating better model convergence and feature representations. We then performed multi-class segmentation on the FAROS dataset using the proposed RS_Unet3+ and automatically calculated multiple RVN and FAZ parameters from the segmented image for quantitative analysis. Primary quantitative parameters included FAZ area (A), perimeter (P), circularity index (CI), vessel perfusion density (VPD), vessel length density (VLD), fractal dimension (FD), and tortuosity (T). Statistical comparisons between healthy and diseased groups were conducted using the Mann-Whitney U test with a significance threshold of P < 0.05.
The proposed RS_Unet3+ was verified through systematic experiments to achieve excellent single-task/multi-class performances for RVN or/and FAZ segmentation on 2 publicly available OCTA datasets, respectively. Moreover, the multi-class segmentation task on the FAROS dataset using RS_Unet3+ achieved excellent segmentation performance that could be used as baseline performance as well. In the independent clinical DR dataset, P of FAZ were significantly higher in the DR group compared to the healthy controls (P < 0.01), while the CI was significantly lower (P < 0.001). Vessel parameters including VLD and FD were significantly reduced (P < 0.05, P < 0.01, respectively), and T increased (P < 0.05) in DR eyes, aligning with the previously reported clinical patterns.
The RS_Unet3+-based deep learning framework enables accurate segmentation and automated quantitative evaluation of retinal microstructures in OCTA images, which is expected to be a potentially reliable and convenient auxiliary tool for clinical disease diagnosis and treatment.
Recurrent laryngeal nerve (RLN) injury is a notable complication in endoscopic thyroidectomy. Although convolutional neural networks (CNNs) have been extensively developed for medical applications, their use in surgery has predominantly focused on identifying static objects such as surgical tools and anatomical landmarks. This study introduces a CNN-based model for real-time RLN detection, aimed at improving surgical safety.
Video data from 2 endoscopic thyroidectomy techniques, the chest-breast approach (ETCB) and trans-axillary approach (ETTA), were retrospectively collected at Peking Union Medical College Hospital between February 2020 and August 2021. Eligible cases were selected using predefined inclusion and exclusion criteria. RLN-specific video segments were annotated by expert thyroid surgeons, and the dataset was randomly divided into training (83.3%) and test (16.7%) sets. A modified YOLOv3 algorithm with oriented bounding boxes was used for real-time RLN detection. Model performance was evaluated using recall and precision at 2 intersection over union (IoU) thresholds (0.1 and 0.5), and statistical analysis was performed using Python-based tools.
A total of 92 videos (113,690 frames) were included, with the training set comprising 75,700 frames from 42 ETCB videos and 23,111 frames from 37 ETTA videos. The test set consisted of 9,912 ETCB and 4,967 ETTA frames. The model achieved recall rates of 80.6% for ETCB and 91.4% for ETTA, and precision rates of 78.4% and 82.6%, respectively, at an IoU threshold of 0.1. These metrics underscore the model’s robustness in detecting RLN across diverse operative settings.
The developed CNN model reliably detected RLN in real-time during endoscopic thyroidectomy procedures, demonstrating the potential of artificial intelligence to dynamically recognize complex anatomical structures in varied surgical contexts.
Hemoglobin (HGB), hematocrit (HCT), and red blood cell (RBC) concentration are critical parameters for diagnosing anemia, which is traditionally measured using automated hematology analyzers. Recent advancements in deep learning in blood cell analysis have primarily focused on the classification of white blood cells, yet there is limited research on the quantitative prediction of HGB and RBC parameters from peripheral blood smear images. We aimed to develop and validate deep learning models to predict HGB, HCT, and RBC levels directly from blood smear images, offering a potential low-cost alternative to traditional methods.
A multicenter dataset of over 13,000 stained blood smear images paired with corresponding hematology analysis results was collected from 3 medical centers between March 2021 and March 2022. Deep learning models, based on the ResNet18 architecture, were trained to quantitatively predict HGB, HCT, and RBC levels. The models were evaluated using Pearson’s correlation coefficient (PCC), R-squared (R2) values, and mean absolute error (MAE). Tenfold cross-validation was employed to ensure robustness, and external validation was performed on 2 independent datasets.
The deep learning models demonstrated strong performance in predicting HGB (PCC = 0.977, R2 = 0.954, MAE = 4.269 g/L), HCT (PCC = 0.982, R2 = 0.965, MAE = 0.013 L/L), and RBC concentration (PCC = 0.982, R2 = 0.964, MAE = 0.155 × 1012/L). The models could also predict RBC concentration from single images with varying cell densities, leveraging the global cell distribution patterns. External validation showed consistent performance, with PCCs exceeding 0.9 for HGB, HCT, and RBC predictions.
This study demonstrated the feasibility of using deep learning models to predict HGB, HCT, and RBC levels directly from blood smear images, offering a novel, cost-effective, and efficient method for hematology analysis.
Overutilization of medical imaging is a significant problem in healthcare, contributing to wasted resources and potentially causing harm to patients. Despite educational efforts and tools, appropriate imaging adoption remains challenging. To address this, we aimed to train an AI model, termed the Appropriate Medical Imaging Recommendations Generative Pre-trained Transformer (AMIR-GPT), to provide precise recommendations for medical imaging, thereby advancing value-based healthcare.
This prospective study used a dataset comprising 1036 paired questions and answers, collected from 26 guidelines in the American College of Radiology Appropriateness Criteria (ACR AC). The dataset, covering common clinical scenarios, was divided into a training set (932 entries) and a test set (104 entries). The OpenAI text-davinci model based on GPT-3 was fine-tuned in four iterations using the training set. The performance of AMIR-GPT was compared to GPT-4 and GPT-3.5 on the test set. Response similarity to standard answers was scored from 1 to 5 using a weighted Cohen’s kappa to measure inter-rater reliability between the model-generated responses and expert reviewers. Statistical significance was assessed using a chi-square test to compare categorical performance metrics across the models.
AMIR-GPT achieved the highest perfect score rate (33.33%), outperforming GPT-4, Gemini, and GPT-3.5. In the high match category, GPT-3.5 led with 25%, while Gemini excelled in the medium match category at 37.5%. ANOVA confirmed significant differences among models (f = 6.49, P = 0.0004). Notable pairwise results included significant differences between AMIR-GPT and GPT-3.5 (P = 0.018) and between GPT-3.5 and Gemini (P = 0.000), indicating varied model performance.
Fine-tuning GPT models for specific medical domains enhances their ability to provide accurate imaging recommendations. However, further validation is needed to confirm the broader applicability of these findings in various clinical settings.
Emergency medical services (EMSs) management requires maintaining a delicate balance between time, resources, and quality of care. Rapid and effective decision-making is crucial for patient outcomes. Our goal is to integrate advanced large language models (LLMs) into EMS systems to assist in triage decisions and test their practicality and benefits.
This method is designed for emergency triage scenarios. By designing specific prompts to introduce heuristic emergency strategies, it makes full use of the multi-turn dialogue capability and contextual understanding characteristics of LLMs to achieve a comprehensive assessment of the dynamic changes in the condition of the injured and emergency resources. Thus, it forms dynamic triage decisions for a large number of injured people, and can also provide detailed explanations of the decision reasons. This method was evaluated and verified using 4 different LLMs (GPT-4, GLM-4, Qwen-max-0428, and Baichuan2-7b-chat-v1) in various scenarios, including different numbers of injured individuals and various types of large-scale casualty events on our self-built emergency medical dispatch simulation platform, and was compared with the nearest transport method. Additionally, the differences between doctors and LLMs in terms of triage decisions were compared, and emergency experts were invited to evaluate the triage decision results and processes.
We conducted experiments on EMSs under 6 different resource environment conditions. With comprehensive patient information and hospital treatment capacity information, GLM-4, GPT-4, and Qwen-max-0428 demonstrated decision-making capabilities far surpassing traditional evacuation methods. GLM-4 and Qwen-max-0428 improved survival rates by an average of 15% after prompt optimization, whereas GPT-4 performed even better, with an average improvement in survival rates reaching 23% after prompt optimization. The consistency level of manual controlled trials (as high as 0.67) reveals that LLMs have guiding and training significance for inexperienced triage personnel in making triage decisions. However, in clinicians’ evaluations, it was revealed that LLMs possess good decision-making abilities, but there is still scope for improvement compared to the level of emergency experts.
This study highlights the potential of LLMs in EMS diversion decision-making and suggests that more comprehensive emergency information can further enhance their decision-making abilities.
Antifreeze proteins (AFPs) are key in combating cold in living organisms and preventing ice morphogenesis. These proteins have applications in cryopreservation, food preservation, and biotechnology. Factors such as accurate prediction of AFP are considered essential for advancing these fields.
In this study, a novel method, StackAFP, was developed using the stacking method and latent semantic analysis as the feature extraction technique for predicting antifreeze proteins. Four machine learning algorithms, random forest, XGboost, CatBoost, and LightGBM (LGBM), were used as the baseline models, and LGBM was employed as the meta-classifier to develop StackAFP. StackAFP was compared with different conventional machine learning methods to ensure the robustness of the proposed method.
StackAFP showed potentiality with an accuracy of 0.9997, a Matthews correlation coefficient, and a Kappa value of 0.9944. StackAFP outperformed the entire applied conventional machine learning model. Furthermore, StackAFP also outperformed the existing methods for identifying AFPs.
The performance of StackAFP demonstrated its effectiveness, highlighted its potential in bioinformatics, and advanced our knowledge of AFPs.
The successful utilization of artificial intelligence (AI) systems in medical and healthcare systems has substantially advanced research and publication in scientific journals. This field encompasses a wide range of studies, including image processing, natural language processing, medical physics, patient data analysis, and clinical assistance tools. The current progress in AI methods can be attributed to substantial improvements in computational capacity and data processing capabilities. Notably, computer vision and image processing have emerged as highly successful AI applications. The U-Net convolutional neural network has emerged as a powerful and efficient tool for medical image segmentation and processing. This model features an encoder-decoder configuration interconnected by a bridging element, with skip connections between layers that enhance the value of the original training data. Its impressive efficiency in image processing stems from its rapid processing capability, ability to extract relationships from data, and high training velocity.
Medical imagery often comprises multiple cross-sectional slices, providing a volumetric perspective of the observed region. Analyzing such imaging data requires substantial computational power and storage capacity, especially for 3-dimensional (3D) analysis. In this regard, 3D U-Net networks excel by concurrently processing numerous slices in voxel space. This attribute considerably reduces computational expenses while simultaneously improving precision. Recent advancements have brought AI technology developers to a level of stability and positive predictive values that make routine use of these systems in medical devices feasible. Currently, the implementation of AI systems appears more realistic than in previous decades. This review article focuses on state-of-the-art technologies in medical imaging for monitoring and diagnostic purposes, specifically using U-Net. We have reviewed the quality measurement of AI imaging systems using gold standards and explored novel technologies that have not been discussed in previous U-Net review papers. Additionally, we discuss the promising future development of AI systems for medical imaging purposes.
The integration of artificial intelligence (AI) with healthcare is beneficial for expanding the accessibility of medical services, benefiting doctors and patients alike. Currently, research on the intelligentization of traditional Chinese medicine (TCM) is also underway. Patient information is collected using AI in TCM and converted into data storage. Digital technology and intelligent algorithms can construct non-private information fragments into a form of privacy problem. We explored the privacy dilemma that the combination of TCM and AI may face, based on case studies, and focused on analyzing the following issues: First, from the perspective of Chinese culture, the cultural sensitivity and privacy conflict in the diagnosis and treatment of diseases using AI in TCM are examined, emphasizing the need to comprehensively consider the cultural background, personal privacy, and diagnostic accuracy in the application of AI in TCM to find a suitable balance point. Second, the philosophical significance behind the combination of TCM and AI includes the diversity of body cognition and local knowledge. Third, the combination of TCM and AI may introduce social governance issues, along with several governance suggestions that can be referred.
CURRENT ISSUE

