edithatogo/nhanes-2021-2023-diabetes
NHANES August 2021–August 2023 — diabetes and demographics Unofficial mirror of two real CDC/NCHS public-use files. This is not the complete NHANES cycle, a monogenic-diabetes dataset or a custodian deployment. The train labels are hosting conventions, not study-design splits. Attribution and permitted use Centers for Disease Control and Prevention (CDC). National Center for Health Statistics (NCHS). National Health and Nutrition Examination Survey Data, August… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/nhanes-2021-2023-diabetes.
NHANES August 2021–August 2023 — diabetes and demographics
Unofficial mirror of two real CDC/NCHS public-use files. This is not the complete NHANES cycle, a monogenic-diabetes dataset or a custodian deployment. The train labels are hosting conventions, not study-design splits.
Attribution and permitted use
Centers for Disease Control and Prevention (CDC). National Center for Health Statistics (NCHS). National Health and Nutrition Examination Survey Data, August 2021–August 2023. Hyattsville, MD: U.S. Department of Health and Human Services, CDC. DIQL and DEMOL, first published September 2024.
NHANES citation and reproduction guidance permits reproduction of federal public-domain materials. The NCHS Data User Agreement still applies. Source terms checked 2026-09-06; no new unrestricted licence is asserted. No CDC/NCHS endorsement is implied.
Use these data only for statistical reporting and analysis. Do not attempt to identify people or establishments, link to individually identifiable information, or research re-identification or assess the source's disclosure-protection methods. This mirror does not waive these restrictions. RareBurden must not use NHANES for privacy attack testing; invented fixtures are used for adversarial testing instead.
Contents and processing
raw/DIQ_L.xptandraw/DEMO_L.xpt: unchanged official SAS transport files.data/*.parquet: separate table conversions with original variable names, numerical codes and missing values; no imputation, recoding, filtering or join.manifest.json: exact source URLs, hashes, parser versions and round-trip checks.prepare.py: offline conversion script shared with the UCI mirror; run from the parent directory containing both dataset folders.
The Parquet tables are conveniences, not substitutes for SAS labels/formats and the DIQ_L codebook and DEMO_L codebook. Preserve the XPT originals for source fidelity. Missing values are not negative diagnoses. SEQN is the source respondent key, not a cross-dataset identity.
Statistical limitations
Diabetes history is questionnaire-based, not molecular confirmation. Use the appropriate survey weights, strata and PSU design variables for population inference. Questionnaire-only analyses use interview weights; adding examination or laboratory data changes the applicable weighting rules. Unweighted sample counts are not population prevalence. This mirror supplies no prevalence estimate, clinical validation, family-level inference or genetic classification.
