Intelligent Medicine
Volume 04 · Issue 02 · 2024
Intell Med
- Sections
- Research Article
Tuberculosis (TB) is among the most frequent causes of infectious-disease-related mortality. Despite being treatable by antibiotics, tuberculosis often goes misdiagnosed and untreated, especially in rural and low-resource areas. Chest X-rays are frequently used to aid diagnosis; however, this presents additional challenges because of the possibility of abnormal radiological appearance and a lack of radiologists in areas where the infection is most prevalent. Implementing deep-learning-based imaging techniques for computer-aided diagnosis has the potential to enable accurate diagnoses and lessen the burden on medical specialists. In the present work, we aimed to develop deep-learning-based segmentation and classification models for accurate and precise detection of tuberculosis in chest X-ray images, with visualization of infection using gradient-weighted class activation mapping (Grad-CAM) heatmaps.
First, we trained the UNet segmentation model using 704 chest X-ray radiographs taken from the Montgomery County and Shenzhen Hospital datasets. Next, we implemented the trained UNet model on 1,400 tuberculosis and control chest X-ray scans to segment the lung region. The images were taken from the National Institute of Allergy and Infectious Diseases (NIAID) TB portal program dataset. Then, we applied the deep learning Xception model to classify the segmented lung region into tuberculosis and normal classes. We further investigated the visualization capabilities of the model using Grad-CAM to view tuberculosis abnormalities in chest X-rays and discuss them from radiological perspectives.
For segmentation by the UNet model, we achieved accuracy, Jaccard index, Dice coefficient, and area under the curve (AUC) values of 96.35%, 90.38%, 94.88%, and 0.99, respectively. For classification by the Xception model, we achieved classification accuracy, precision, recall, F1-score, and AUC values of 99.29%, 99.30%, 99.29%, 99.29%, and 0.999, respectively. The Grad-CAM heatmap images from the tuberculosis class showed similar heatmap patterns, where lesions were primarily present in the upper part of the lungs.
The findings may verify our system’s efficacy and superiority to clinician precision in tuberculosis diagnosis using chest X-rays and raise the possibility of a valuable setup, particularly in environments with a scarcity of radiological expertise.
With the gradual increase of infertility in the world, among which male sperm problems are the main factor for infertility, more and more couples are using computer-assisted sperm analysis (CASA) to assist in the analysis and treatment of infertility. Meanwhile, the rapid development of deep learning (DL) has led to strong results in image classification tasks. However, the classification of sperm images has not been well studied in current deep learning methods, and the sperm images are often affected by noise in practical CASA applications. The purpose of this article is to investigate the anti-noise robustness of deep learning classification methods applied on sperm images.
The SVIA dataset is a publicly available large-scale sperm dataset containing three subsets. In this work, we used subset-C, which provides more than 125,000 independent images of sperms and impurities, including 121,401 sperm images and 4,479 impurity images. To investigate the anti-noise robustness of deep learning classification methods applied on sperm images, we conducted a comprehensive comparative study of sperm images using many convolutional neural network (CNN) and visual transformer (VT) deep learning methods to find the deep learning model with the most stable anti-noise robustness.
This study proved that VT had strong robustness for the classification of tiny object (sperm and impurity) image datasets under some types of conventional noise and some adversarial attacks. In particular, under the influence of Poisson noise, accuracy changed from 91.45% to 91.08%, impurity precison changed from 92.7% to 91.3%, impurity recall changed from 88.8% to 89.5%, and impurity F1-score changed 90.7% to 90.4%. Meanwhile, sperm precision changed from 90.9% to 90.5%, sperm recall changed from 92.5% to 93.8%, and sperm F1-score changed from 92.1% to 90.4%.
Sperm image classification may be strongly affected by noise in current deep learning methods; the robustness with regard to noise of VT methods based on global information is greater than that of CNN methods based on local information, indicating that the robustness with regard to noise is reflected mainly in global information.
Parastomal hernia is one of the potential complications after enterostomy. There is currently no early risk assessment tool for parastomal hernia.
The current investigation was conducted using retrospective studies. A total of 302 cases were used develop and internally to validate a nomogram prediction model, and 67 cases were used for external validation. Independent risk factors for parastomal hernia after permanent sigmoid colostomy were assessed via univariate analysis and binary logistic regression analysis. The nomogram prediction model was established based on independent risk factors.
Body mass index, serum albumin, age, sex, and stoma diameter were independent risk factors for parastomal hernia. The areas under the receiver operating characteristic curves were 0.909 in the development group and 0.801 in the validation group. The Hosmer-Lemeshow test (P > 0.05) and calibration curves indicated good consistency between actual observations and predicted probabilities.
A nomogram prediction model was constructed and validated based on risk factors for parastomal hernia. The nomogram could be generalized to patients undergoing surgery for stoma by specialized surgeons to provide relevant references for stoma patients.
Accurate infant brain parcellation is crucial for understanding early brain development; however, it is challenging due to the inherent low tissue contrast, high noise, and severe partial volume effects in infant magnetic resonance images (MRIs). The aim of this study was to develop an end-to-end pipeline that enabled accurate parcellation of infant brain MRIs.
We proposed an end-to-end pipeline that employs a two-stage global-to-local approach for accurate parcellation of infant brain MRIs. Specifically, in the global regions of interest (ROIs) localization stage, a combination of transformer and convolution operations was employed to capture both global spatial features and fine texture features, enabling an approximate localization of the ROIs across the whole brain. In the local ROIs refinement stage, leveraging the position priors from the first stage along with the raw MRIs, the boundaries of the ROIs are refined for a more accurate parcellation.
We utilized the Dice ratio to evaluate the accuracy of parcellation results. Results on 263 subjects from National Database for Autism Research (NDAR), Baby Connectome Project (BCP) and Cross-site datasets demonstrated the better accuracy and robustness of our method than other competing methods.
Our end-to-end pipeline may be capable of accurately parcellating 6-month-old infant brain MRIs.
Brain stroke is a serious health issue that requires timely and accurate prediction for effective treatment and prevention. This study described a hybrid system that used the best feature selection method and classifier to predict brain stroke.
The Stroke Prediction Dataset from Kaggle was used for this study. Synthetic minority over-sampling technique (SMOTE) analysis was used to accomplish class balancing. Accuracy, sensitivity, specificity, precision, and the F-Measure were the main performance parameters considered for investigation. To determine the best combination for predicting brain stroke, the performance of five classifiers, Naïve Bayes (NB), support vector machine (SVM), random forest (RF), adaptive boosting (Adaboost), and extreme gradient boosting (XGBoost), was compared along with three feature selection techniques, mutual information (MI), Pearson correlation (PC), and feature importance (FI). The performance parameters were assessed using k-fold cross-validation.
The hybrid system proposed in this study identified a reduced set of features that were able to effectively predict brain stroke. FI provided a feature reduction ratio of 36.3%. The most successful hybrid system for predicting brain stroke used FI as the feature selection technique and RF as the classifier, achieving an accuracy of 97.17%.
The proposed system predicted brain stroke with high accuracy. These findings could be used to inform the early detection and prevention of brain stroke, allowing healthcare professionals to provide timely and targeted care to at-risk patients.
The incidence of cardiovascular diseases (CVD) is rising rapidly worldwide. Some forms of CVD, such as stroke and heart attack, are more common among patients with certain conditions. Atherosclerosis development is a major factor underlying cardiovascular events, such as heart attack and stroke, and its early detection may prevent such events. Ultrasound imaging of carotid arteries is a useful method for diagnosis of atherosclerotic plaques; however, an automated method to classify atherosclerotic plaques for evaluation of early-stage CVD is needed. Here, we propose an automated method for classification of high-risk atherosclerotic plaque ultrasound images.
Five deep learning (DL) models (VGG16, ResNet-50, GoogLeNet, XceptionNet, and SqueezeNet) were used for automated classification and the results compared with those of a machine learning (ML)-based technique, involving extraction of 23 texture features from ultrasound images and classification using a Support Vector Machine classifier. To enhance model interpretability, output gradient-weighted convolutional activation maps (GradCAMs) were generated and overlayed on original images.
A series of indices, including accuracy, sensitivity, specificity, F1-score, Cohen-kappa index, and area under the curve values, were calculated to evaluate model performance. GradCAM output images allowed visualization of the most significant ultrasound image regions. The GoogLeNet model yielded the highest accuracy (98.20%).
ML models may be also suitable for applications requiring low computational resource. Further, DL models could be more completely automated than ML models.
Speech recognition technology is widely used as a mature technical approach in many fields. In the study of depression recognition, speech signals are commonly used due to their convenience and ease of acquisition. Though speech recognition is popular in the research field of depression recognition, it has been little studied in somatisation disorder recognition. The reason for this is the lack of a publicly accessible database of relevant speech and benchmark studies. To this end, we introduced our somatisation disorder speech database and gave benchmark results.
By collecting speech samples of somatisation disorder patients, in cooperation with the Shenzhen University General Hospital, we introduced our somatisation disorder speech database, the Shenzhen Somatisation Speech Corpus (SSSC). Moreover, a benchmark for SSSC using classic acoustic features and a machine learning model was proposed in our work.
To obtain a more scientific benchmark, we compared and analysed the performance of different acoustic features, i. e., the full ComPare feature set, or only Mel frequency cepstral coefficients (MFCCs), fundamental frequency (F0), and frequency and bandwidth of the formants (F1-F3). By comparison, the best result of our benchmark was the 76.0% unweighted average recall achieved by a support vector machine with formants F1-F3.
The proposal of SSSC may bridge a research gap in somatisation disorder, providing researchers with a publicly accessible speech database. In addition, the results of the benchmark could show the scientific validity and feasibility of computer audition for speech recognition in somatization disorders.
CURRENT ISSUE

