CoolFace
Datasetpublic

electricsheepasia/asia-who-historical-data-for-turkmenistan

Turkmenistan - Historical Health Indicators Publisher: World Health Organization · Source: HDX · License: hdx-other · Updated: 2025-02-07 Abstract This dataset contains historical data from WHO's data portal. Each row in this dataset represents first-level administrative unit observations. Data was last updated on HDX on 2025-02-07. Geographic scope: TKM. Curated into ML-ready Parquet format by Electric Sheep Africa. Dataset Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-historical-data-for-turkmenistan.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes12downloads
Dataset Card

Turkmenistan - Historical Health Indicators

Publisher: World Health Organization · Source: HDX · License: hdx-other · Updated: 2025-02-07


Abstract

This dataset contains historical data from WHO's data portal.

Each row in this dataset represents first-level administrative unit observations. Data was last updated on HDX on 2025-02-07. Geographic scope: TKM.

Curated into ML-ready Parquet format by [Electric Sheep Africa](https://huggingface.co/electricsheepafrica).


Dataset Characteristics

DomainHumanitarian and development data
Unit of observationFirst-level administrative unit observations
Rows (total)9,902
Columns18 (6 numeric, 12 categorical, 0 datetime)
Train split7,921 rows
Test split1,980 rows
Geographic scopeTKM
PublisherWorld Health Organization
HDX last updated2025-02-07

Variables

Geographic — gho_display (Mean BMI (kg/m²) (crude estimate), Adolescent mortality rate (per 1 000 age specific cohort), Alcohol, recorded per capita (15+) consumption (in litres of pure alcohol)), year_display (range 1970.0–2025.0), startyear (range 1970.0–2025.0), endyear (range 1970.0–2025.0), region_code (EUR, #region+code) and 4 others.

Outcome / Measurement — value (No data, No, Yes).

Identifier / Metadata — gho_code (NCDBMIMEANC, CHILDMORT10TO19, SA0000001400ARCHIVED), dimension_code (SEXFMLE, SEXMLE, SEXBTSX), `dimensionname (Female, Male, Both sexes), esasource`, `esaprocessed`.

Other — numeric (range 0.0–266862800.0), low (range 0.0–37051.0), high (range 0.0–173952.0).


Quick Start

python
from datasets import load_dataset

ds    = load_dataset("electricsheepafrica/asia-who-historical-data-for-turkmenistan")
train = ds["train"].to_pandas()
test  = ds["test"].to_pandas()

print(train.shape)
train.head()

Schema

ColumnTypeNull %Range / Sample Values
gho_codeobject0.0%NCDBMIMEANC, CHILDMORT10TO19, SA0000001400ARCHIVED
gho_displayobject0.0%Mean BMI (kg/m²) (crude estimate), Adolescent mortality rate (per 1 000 age specific cohort), Alcohol, recorded per capita (15+) consumption (in litres of pure alcohol)
year_displayfloat640.0%1970.0 – 2025.0 (mean 2009.9776)
startyearfloat640.0%1970.0 – 2025.0 (mean 2009.9657)
endyearfloat640.0%1970.0 – 2025.0 (mean 2009.9776)
region_codeobject0.0%EUR, #region+code
region_displayobject0.0%Europe, #region+name
country_codeobject0.0%TKM, #country+code
country_displayobject0.0%Turkmenistan, #country+name
dimension_typeobject19.2%SEX, WEALTHDECILE, WEALTHQUINTILE
dimension_codeobject19.2%SEXFMLE, SEXMLE, SEX_BTSX
dimension_nameobject20.4%Female, Male, Both sexes
numericfloat6434.8%0.0 – 266862800.0 (mean 41711.3983)
valueobject1.3%No data, No, Yes
lowfloat6449.8%0.0 – 37051.0 (mean 114.7989)
highfloat6449.9%0.0 – 173952.0 (mean 289.3685)
esa_sourceobject0.0%
esa_processedobject0.0%

Numeric Summary

ColumnMinMaxMeanMedian
year_display1970.02025.02009.97762013.0
startyear1970.02025.02009.96572012.0
endyear1970.02025.02009.97762013.0
numeric0.0266862800.041711.398334.5564
low0.037051.0114.798925.7971
high0.0173952.0289.368544.5

Curation

Raw data was downloaded from HDX via the CKAN API and converted to Parquet. Column names were lowercased and standardised to snakecase. Common missing-value markers (`N/A`, `null`, `none`, `-`, `unknown`, `no data`, `#N/A`) were unified to `NaN`. 1 column(s) with >80% missing values were removed: `ghourl`. 67 exact duplicate rows were removed. 6 column(s) were cast from string to numeric or datetime based on parse-success rate (>85% threshold). The dataset was split 80/20 into train and test partitions using a fixed random seed (42) and saved as Snappy-compressed Parquet.


Limitations

  • —Data originates from World Health Organization and has not been independently validated by ESA.
  • —Automated cleaning cannot correct for misreported values, definitional inconsistencies, or sampling bias in the original collection.
  • —The following columns have >20% missing values and should be treated with caution in modelling: dimension_name, numeric, low, high.
  • —Refer to the original HDX dataset page for the publisher's own methodology notes and caveats.

Citation

bibtex
@dataset{hdx_asia_who_historical_data_for_turkmenistan,
  title     = {Turkmenistan - Historical Health Indicators},
  author    = {World Health Organization},
  year      = {2025},
  url       = {https://data.humdata.org/dataset/who-historical-data-for-tkm},
  note      = {Repackaged for machine learning by Electric Sheep Africa (https://huggingface.co/electricsheepafrica)}
}

[Electric Sheep Africa](https://huggingface.co/electricsheepafrica) — Africa's ML dataset infrastructure. Lagos, Nigeria.