中华儿科杂志
2021年 · 第59卷第06期
中华儿科杂志
- 全部
- 述评
- 标准·方案·指南
- 内分泌遗传代谢疾病研究
- 论著
- 临床研究与实践
- 临床研究方法学园地
- 病例报告
- 综述
- 继续教育园地
Many clinical researchers will add prediction models after multivariate analysis to observe whether the selected associated factors can accurately predict the occurrence of an event. Whether the prediction model is useful or not depends on the results of validation, so the selection of validation set is very important. At present, the selection methods of commonly used validation sets are as follows: (1) Randomly select the validation set, that is, randomly divide the original data into two groups, one group is used as the training set and the other group is used as the validation set. The grouping ratio can be 5:5, 6:4, 7:3, 8:2, etc., depending on the actual sample size. Generally, the sample size of the training set is more than that of the validation set, because the two data sets are randomly divided, the data sets are also very "similar", and the validation results are usually good, so it is most popular among clinical researchers. (2) Select the validation set by center, that is, in a multicenter study, apply the data of several centers as the training set, and use the data of one or several other centers as the validation set. More extensive selection of validation sets by geographical location, such as the model established by Northern Hospital is validated with data from Southern Hospital, and the model established by Chinese data is validated with foreign data, etc. This is a more recommended method now. (3) Select the validation set according to time, that is, the cases collected before a certain time point are used as the training set, and the data after a certain time point are used as the validation set. This method of determining the validation set is best in line with the original intention of forecasting, that is, to build a model with existing data and verify it in future data.
本期目次

