What this research found
How much does it matter whether a complex survey design is honoured? Using the 2017-2018 National Health and Nutrition Examination Survey, the prevalence of diagnosed diabetes among US adults aged 20 and over was estimated twice: once with the sampling weights, masked strata and masked primary sampling units, and once as if the sample were a simple random draw. The design-based figure was 11.85% while the naive one was 16.24%, an upward bias of roughly 4.4 percentage points. A survey-weighted logistic regression then measured the adjusted associations with age, measured body-mass index, physical-activity guideline attainment, and family income.
- Survey-weighted prevalence of diagnosed diabetes among US adults aged 20 and over was 11.85% (95% confidence interval 10.70 to 13.01), with a design-based standard error of 0.54% over a domain of 5,393 respondents.
- Treating the same records as a simple random sample returned 16.24% (95% confidence interval 15.26 to 17.23), overstating national prevalence by about 4.4 percentage points, a relative overestimate of roughly 37%. The intervals do not overlap, so sampling variability cannot explain the gap: the survey deliberately oversamples older and higher-risk adults.
- Age dominated the adjusted model. Against adults aged 20 to 39, those aged 40 to 59 had 5.53 times the adjusted odds of diagnosed diabetes (95% confidence interval 3.93 to 7.78) and those 60 and over had 13.57 times the odds (8.44 to 21.80). Unadjusted weighted prevalence followed the same path, from 2.2% to 11.6% to 24.5%.
- Measured adiposity showed a clear dose-response. Relative to normal weight, overweight carried an adjusted odds ratio of 1.99 (95% confidence interval 1.36 to 2.91) and obesity 4.43 (3.12 to 6.28), while weighted prevalence climbed from 4.5% at normal weight to 10.3% among overweight adults and 17.3% among those with obesity.
- Meeting the guideline of at least 150 minutes per week of moderate-to-vigorous activity was independently associated with roughly 28% lower adjusted odds (odds ratio 0.72, 95% confidence interval 0.59 to 0.88), against weighted prevalence of 9.0% versus 16.9%. Low family income carried higher adjusted odds than high income (1.43, 1.04 to 1.96); the middle-income estimate was not significant.
How it was done
Four public-use files from the 2017-2018 survey cycle — demographics, body measures, the diabetes questionnaire, and the physical-activity module — were merged on respondent sequence number and restricted to adults aged 20 and over, leaving 5,569 of 9,254 examined participants. Weekly moderate-to-vigorous activity was derived from four questionnaire domains covering moderate and vigorous work and recreation, with vigorous minutes double-weighted so that 75 vigorous minutes count as 150 moderate ones, then thresholded at 150 minutes per week. With no complex-survey package available in the computing environment, the Taylor-series linearization sandwich variance estimator was implemented directly and used for both the weighted prevalence and the logistic regression, keeping out-of-domain records at zero contribution so the stratum and cluster structure stayed intact. Every interval uses a t-distribution with 15 design degrees of freedom, from 15 masked pseudo-strata each holding two masked primary sampling units.
Data sources
- NHANES 2017-2018 public-use files from the US National Center for Health Statistics — demographics, body measures, diabetes questionnaire, and physical-activity modules (9,254 examined participants, 5,569 adults aged 20 and over)
- Physical Activity Guidelines for Americans, JAMA 320:2020 (2018) and WHO 2020 guidelines on physical activity and sedentary behaviour — source of the 150 minutes per week threshold
- WHO body-mass index cut-points, Technical Report Series 854 (1995)
- Lumley, Journal of Statistical Software 9(8) (2004) — Taylor-linearization variance estimation for complex survey samples
Limitations
Because the design is cross-sectional, none of the associations establish causality, and reverse causation cannot be excluded for physical activity in particular. The outcome is self-reported physician-diagnosed diabetes and therefore misses undiagnosed cases, activity is likewise self-reported and omits transport-related movement, and the regression was fitted on 4,328 complete cases after 7.1% of body-mass index and 14.2% of income values were found missing.
How this research was produced
K-Dense Web planned and ran this public health investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.


