中华儿科杂志
2021年 · 第59卷第12期
中华儿科杂志
- 全部
- 述评
- 标准·方案·指南
- 指南解读
- 免疫疾病研究
- 神经系统疾病研究
- 优先出版
- 临床研究与实践
- 病例报告
- 综述
- 会议纪要
- 临床研究方法学园地
Predictive model research is one of the important directions of clinical research in recent years, and there are many application scenarios in clinical practice. One is the prediction model for the population, which is mainly used to predict the epidemic trend of diseases and judge the epidemic situation of diseases in the population at a specific time point. The other is the prediction model for individuals, which is mainly used to classify or predict individual situations. The predictive model for guiding clinical decision-making is usually the latter, mainly including diagnostic model, prognostic model, etc. In the Clinical Predictive Model Reporting Specification TRIPOD Statement, the validation process of the predictive model is divided into four major categories and six subcategories according to the differences in the data sets used in the model establishment and validation process. There are three main situations of data sets involved: only one data set is available, and all data need to be used for modeling (corresponding to category 1 in the classification, including 1a and 1b); Only one dataset is available, one part of which is used for modeling and the other for validation (corresponding to 2 categories in the classification, including 2a, 2b); Multiple datasets are available (corresponding to classes 3, 4 in the classification). (1) Type 1a: Data is limited and only one dataset is available. A predictive model is built based on the entire data, and then the predictive power of the model is directly evaluated using the exact same data. Because the modeling and validation use a unified data set, the predictive performance of the model is usually overestimated. (2) Type 1b: Data is limited and only one dataset is available. The prediction model is built based on the entire dataset, and then the performance of the prediction model is evaluated using repeated sampling techniques such as Bootstrapping or cross-validation. Repeated sampling technique is usually regarded as "internal validation", which is the basic condition and method for carrying out predictive model research. It is commonly used when the data is limited. (3) Type 2a: The data is relatively large, which can be randomly divided into two groups, one group is used to establish the prediction model, and the other group is used to evaluate the prediction effect of the model. Although Type 2a studies are widely used, TRIPOD believes that modification studies are not superior to Type 1b because Type 2a has low sample utilization rate, which may lead to insufficient efficacy in modeling and validation processes. (4) Type 2b: The data is relatively large and can be divided into two groups non-randomly, one group is used to build a prediction model, and the other group is used to evaluate the prediction effect of the model. TRIPOD believes that the Type 2b study is superior to Type 2a because it allows non-random changes between the 2 datasets, and the validation results of the extrapolation ability of the model are more robust at this time. (5) Type 3: There are more data, and more than 2 datasets are available. One dataset is used to build a predictive model, and another completely different dataset is used to evaluate the predictive effect of the model. For example, two independent studies are carried out before and after, one for modeling and the other for verification. (6) Type 4: Only for existing (published) prediction models, its prediction effect is evaluated on an independent dataset.
本期目次

