← Back to selected work

Health data science · NHANES · Temporal validation

Cardiometabolic Screening Intelligence

Can everyday measurements help identify elevated HbA1c? A survey-aware machine-learning workflow using age, recorded sex, BMI and waist circumference—developed on one NHANES release and evaluated on a separate, later release.

Official CDC/NCHS dataPython · scikit-learnSurvey-weighted evaluationSeparate temporal test

The question

A precise screening task, with a measurable reference outcome.

The outcome is elevated HbA1c (≥5.7%) at examination among U.S. adults aged 20+ without reported diabetes or diabetes medication use. Laboratory HbA1c defines the outcome but never enters the predictor matrix.

This is screening for current elevated HbA1c. It is not prediction of future diabetes, a confirmed diagnosis, or a measurement of every cardiometabolic condition.

Independent temporal test

Performance across a later survey release.

6,644Temporal-test participants
0.744Weighted ROC AUC
80.9%Weighted sensitivity
55.3%Weighted specificity

Validation architecture

Every modelling decision precedes the temporal test.

01Join official NHANES files and validate eligibility
02Split development survey PSUs for fitting and calibration
03Compare four models using group cross-validation
04Freeze calibration and screening threshold
05Evaluate on 2017–March 2020 and audit limitations

Development uses 4,294 adults from 2015–2016: 3,097 for fitting and 1,197 for calibration. The later release supplies 6,644 untouched temporal-test participants. Imputation and scaling are learned within training folds; outcome, glycemic biomarkers, diagnosis and treatment fields are excluded from predictors.

Results and interpretation

Complexity had to earn its place.

01 / Discrimination

The logistic baseline outperformed more complex candidates.

Group cross-validation compared logistic regression, spline logistic regression, random forest and histogram gradient boosting. Logistic regression was selected using development weighted AUC.

Weighted ROC curve on later NHANES release

Temporal AUC 0.744; conditional 95% stratified rescaled PSU bootstrap interval 0.721–0.763. Weighted Brier score 0.171 versus 0.197 for the development-prevalence constant baseline.

02 / Calibration

Good discrimination does not guarantee reliable probabilities.

High predicted probabilities overestimated observed elevated-HbA1c proportions in the later release. This limitation remains visible.

Weighted temporal calibration plot

Sigmoid calibration did not improve temporal Brier score over the uncalibrated selected model. The procedure was not changed retrospectively using test outcomes.

03 / Model reliance

Four accessible inputs, with transparent limitations.

Permutation importance measures predictive reliance, not causal effects. Correlated BMI and waist measures can share importance.

Non-laboratory feature reliance measured by shuffled-feature AUC drop

Age, recorded sex, BMI and waist circumference are the model inputs. Race/ethnicity is used for descriptive audit only; small outcome cells are suppressed and observed differences do not establish fairness.

The operational trade-off

More detection also means more confirmatory testing.

At the development-selected 22.3% threshold: sensitivity 80.9%, specificity 55.3%, positive predictive value 40.2%, and screen-positive fraction 54.5%. Most screen-positive participants therefore did not meet the laboratory outcome.

The threshold targeted at least 80% sensitivity on development calibration data. Clinical utility, cost-effectiveness and the optimal local threshold have not been established.

Scope and limitations

A reproducible research product with clearly bounded claims.

The workflow includes source checksums, validated joins, missingness reports, survey-weighted metrics, conditional uncertainty intervals, subgroup audits, complete-predictor and HbA1c-below-6.5% sensitivity analyses, a saved model and an interactive research application.

These are U.S. survey data, with no Nigerian validation. MEC weights are used without additional laboratory nonresponse reweighting. Bootstrap intervals exclude model-training uncertainty; subgroup estimates have no design-based intervals. HbA1c is a single examination marker and can be affected by haemoglobin and red-cell factors. No clinical deployment or benefit is claimed.

Inspect and reproduce

From source data to a tested research application.

Open the Python pipeline, saved outputs, figures, model artifact, methodology and evidence/application tests. Raw official data are retrieved by the download script with a recorded checksum manifest.