Academic Journal of Computing & Information Science, 2026, 9(6); doi: 10.25236/AJCIS.2026.090609.
Rui Ge1, Guanting Chen2, Yilin Yang1
1School of Statistics and Mathematics, Henan Finance University, Zhengzhou, China, 450046
2School of Computer Science and Artificial Intelligence, Henan Finance University, Zhengzhou, China, 450046
This paper focuses on solving the problems of timing selection and evaluation for fetal chromosomal abnormality detection in Non-Invasive Prenatal Testing (NIPT). Based on actual clinical data, it comprehensively applies correlation analysis, clustering algorithms, risk control models and machine learning methods. The study successfully constructs a fetal Y-chromosome concentration prediction model, a BMI grouping optimization model, a multi-factor clustering model and a female fetal abnormality judgment model. Furthermore, it obtains the optimal strategies for various problems in NIPT. For the correlation between fetal Y-chromosome concentration and indicators such as gestational age, BMI and age of pregnant women, this paper uses the Pearson correlation coefficient equation to analyze the linear correlation between Y-chromosome concentration and gestational age, BMI and age. Then, it constructs a multiple linear regression model of Y-chromosome concentration with gestational age, BMI and age. Finally, F-test and t-test are used to verify the overall significance of the model and the statistical significance of each variable. To reasonably group the BMI of pregnant women carrying male fetuses, this paper uses the silhouette coefficient to compare and analyze various clustering algorithms, and selects the optimal K-Means++ clustering algorithm to divide them into two groups. The ranges of statistical quantities of the two groups are calculated, and then a dynamic risk control model is constructed using the quantile estimation algorithm to determine the optimal detection time point. The results show that when BMI is in the range of [26.6, 32.4], the optimal detection time is 24.11 gestational weeks; when BMI is in the range of [32.5, 41.1], the optimal detection time is 21.33 gestational weeks. By calculating the actual delayed risk ratio, the risk is controlled below 5% and passes the test. With the help of the linear fitting relationship model for auxiliary analysis, the relationship between BMI and the time to meet the standard within the group is further clarified. Finally, the Bootstrap resampling algorithm is used to verify the stability of the model in predicting the optimal detection time point.
K-Means++ Clustering; Quantile Optimization; Multivariate Clustering
Rui Ge, Guanting Chen, Yilin Yang. Research on NIPT Timing Selection and Fetal Abnormality Judgment Based on Clustering Optimization and Dynamic Risk Control. Academic Journal of Computing & Information Science (2026), Vol. 9, Issue 6: 74-83. https://doi.org/10.25236/AJCIS.2026.090609.
[1] Li J, Liu Y, Zhang N, et al. Optimal planning of energy storage in active distribution network for congestion management and voltage support [J]. IEEE Transactions on Power Systems, 2020, 35 (5): 4120-4133.
[2] Li Z, Wang C, Li Y, et al. A dynamic reconfiguration method for active distribution networks considering distributed generation uncertainty [J]. International Journal of Electrical Power & Energy Systems, 2021, 124: 106337.
[3] Wang Y, Chen J, Chen X, et al. Short-term load forecasting for industrial customers based on TCN-LightGBM [J]. IEEE Transactions on Power Systems, 2021, 36 (3): 1984-1997.
[4] Li F, Han Y. A data-driven approach for customer baseline load estimation based on density peak clustering [J]. IEEE Transactions on Smart Grid, 2022, 13 (1): 653-664.
[5] LifeCycle Project-Maternal Obesity and Childhood Outcomes Study Group. Association of gestational weight gain with adverse maternal and infant outcomes [J]. JAMA, 2019, 321 (17): 1702-1715.
[6] Norton M E, Jacobsson B, Swamy G K, et al. Cell-free DNA analysis for noninvasive examination of trisomy [J]. New England Journal of Medicine, 2015, 372 (17): 1589-1597.
[7] Anifowose F, Labadin J, Abdulraheem A. Improving the prediction of petroleum reservoir characterization with a stacked generalization ensemble model of support vector machines [J]. Applied Soft Computing, 2017, 54: 376-388.
[8] Jain A K, Murty M N, Flynn P J. Data Clustering: A Review [J]. ACM Computing Surveys (CSUR), 2020, 53 (3): 1-38.
[9] Hastie T, Tibshirani R, Friedman J. The Elements of Statistical Learning: Data Mining, Inference, and Prediction [M]. 4th ed. Springer Science & Business Media, 2021: 124-156.
[10] James G, Witten D, Hastie T, et al. An Introduction to Statistical Learning: with Applications in R [M]. 2nd ed. Springer, 2021: 89-112.
[11] Efron B, Tibshirani R J. An Introduction to the Bootstrap [M]. Revised ed. Chapman & Hall/CRC, 2020: 67-92.