CoolFace
Datasetpublic

electricsheepasia/asia-gender-philippines-other-0-0-0-0-0-0-0-0-0-0-0

Philippines - Employed Persons by major Industry group and by sex Publisher: OCHA Philippines · Source: HDX · License: hdx-other · Updated: 2025-07-22 Abstract This dataset shows the employed Persons by major Industry group and by sex Each row in this dataset represents tabular records. Data was last updated on HDX on 2025-07-22. Geographic scope: PHL. Curated into ML-ready Parquet format by Electric Sheep Africa. Dataset Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-gender-philippines-other-0-0-0-0-0-0-0-0-0-0-0.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes11downloads
Dataset Card

Philippines - Employed Persons by major Industry group and by sex

Publisher: OCHA Philippines · Source: HDX · License: hdx-other · Updated: 2025-07-22


Abstract

This dataset shows the employed Persons by major Industry group and by sex

Each row in this dataset represents tabular records. Data was last updated on HDX on 2025-07-22. Geographic scope: PHL.

Curated into ML-ready Parquet format by [Electric Sheep Africa](https://huggingface.co/electricsheepafrica).


Dataset Characteristics

DomainHumanitarian and development data
Unit of observationTabular records
Rows (total)55
Columns22 (19 numeric, 3 categorical, 0 datetime)
Train split44 rows
Test split11 rows
Geographic scopePHL
PublisherOCHA Philippines
HDX last updated2025-07-22

Variables

Identifier / Metadata — unnamed_1 (range 2004.0–2011.0), unnamed_2 (range 7.0–61882.0), unnamed_3 (range 11.3–7979.0), unnamed_4 (range 4.2–1121.0), unnamed_5 (range 8.1–3467.0) and 15 others.

Other — table_11_1 (HOUSEHOLD POPULATION 15 YEARS OLD AND OVER BY EMPLOYMENT STATUS, AND BY REGION, 2004 to 2011), table_11_1_continued (range 3.9–1875.0).


Quick Start

python
from datasets import load_dataset

ds    = load_dataset("electricsheepafrica/asia-gender-philippines-other-0-0-0-0-0-0-0-0-0-0-0")
train = ds["train"].to_pandas()
test  = ds["test"].to_pandas()

print(train.shape)
train.head()

Schema

ColumnTypeNull %Range / Sample Values
table_11_1object69.1%HOUSEHOLD POPULATION 15 YEARS OLD AND OVER BY EMPLOYMENT STATUS, AND BY REGION, 2004 to 2011
unnamed_1float6432.7%2004.0 – 2011.0 (mean 2007.7027)
unnamed_2float6427.3%7.0 – 61882.0 (mean 11512.1367)
unnamed_3float6427.3%11.3 – 7979.0 (mean 1570.6767)
unnamed_4float6427.3%4.2 – 1121.0 (mean 245.4002)
unnamed_5float6427.3%8.1 – 3467.0 (mean 672.4007)
unnamed_6float6427.3%2.822 – 2225.0 (mean 459.1946)
unnamed_7float6427.3%7.4 – 6828.0 (mean 1284.0097)
unnamed_8float6427.3%9.2 – 7988.0 (mean 1491.0814)
table_11_1_continuedfloat6427.3%3.9 – 1875.0 (mean 373.8818)
unnamed_10float6427.3%4.9131 – 3463.0 (mean 687.3598)
unnamed_11float6427.3%6.0953 – 5073.0 (mean 974.85)
unnamed_12float6427.3%6.5 – 4679.0 (mean 899.7042)
unnamed_13float6427.3%4.2005 – 2777.0 (mean 557.4235)
unnamed_14float6427.3%3.2 – 2245.0 (mean 452.1456)
unnamed_15float6427.3%4.5 – 2874.0 (mean 570.9145)
unnamed_16float6427.3%5.4 – 2889.0 (mean 576.5064)
unnamed_17float6427.3%3.9 – 2640.0 (mean 525.4821)
unnamed_18float6427.3%5.0098 – 1627.0 (mean 345.0814)
unnamed_19float6427.3%2.3451 – 2131.0 (mean 420.1198)
esa_sourceobject0.0%HDX
esa_processedobject0.0%2026-05-06

Numeric Summary

ColumnMinMaxMeanMedian
unnamed_12004.02011.02007.70272008.0
unnamed_27.061882.011512.136764.15
unnamed_311.37979.01570.676762.2506
unnamed_44.21121.0245.400267.35
unnamed_58.13467.0672.400761.1752
unnamed_62.8222225.0459.194667.7
unnamed_77.46828.01284.009760.9
unnamed_89.27988.01491.081463.55
table_11_1_continued3.91875.0373.881869.6
unnamed_104.91313463.0687.359865.25
unnamed_116.09535073.0974.8564.85
unnamed_126.54679.0899.704264.2648
unnamed_134.20052777.0557.423566.05
unnamed_143.22245.0452.145665.7
unnamed_154.52874.0570.914570.65

Curation

Raw data was downloaded from HDX via the CKAN API and converted to Parquet. Column names were lowercased and standardised to snake_case. Common missing-value markers (N/A, null, none, -, unknown, no data, #N/A) were unified to NaN. 9 exact duplicate rows were removed. 19 column(s) were cast from string to numeric or datetime based on parse-success rate (>85% threshold). The dataset was split 80/20 into train and test partitions using a fixed random seed (42) and saved as Snappy-compressed Parquet.


Limitations

  • —Data originates from OCHA Philippines and has not been independently validated by ESA.
  • —Automated cleaning cannot correct for misreported values, definitional inconsistencies, or sampling bias in the original collection.
  • —The following columns have >20% missing values and should be treated with caution in modelling: table_11_1, unnamed_1, unnamed_2, unnamed_3, unnamed_4, unnamed_5, unnamed_6, unnamed_7....
  • —Refer to the original HDX dataset page for the publisher's own methodology notes and caveats.

Citation

bibtex
@dataset{hdx_asia_gender_philippines_other_0_0_0_0_0_0_0_0_0_0_0,
  title     = {Philippines - Employed Persons by major Industry group and by sex},
  author    = {OCHA Philippines},
  year      = {2025},
  url       = {https://data.humdata.org/dataset/philippines-other-0-0-0-0-0-0-0-0-0-0-0-0-0-0-0-0-0-0-0-0-0-0},
  note      = {Repackaged for machine learning by Electric Sheep Africa (https://huggingface.co/electricsheepafrica)}
}

[Electric Sheep Africa](https://huggingface.co/electricsheepafrica) — Africa's ML dataset infrastructure. Lagos, Nigeria.