datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
typhoon-intensity-classification
Typhoon - Image Classification Dataset
This dataset comes from PTIT AI Challenge and is organized for a multi-class image classification task focusing on tropical cyclone (typhoon) intensity estimation.
Dataset Structure
The directory structure is organized as follows:
train/
├── images/
│ ├── image1.jpg
│ └── ...
└── annotations.csv (only present in the train folder)
The public_test and private_test sets are used to evaluate and score the… See the full description on the dataset page: https://huggingface.co/datasets/star092304/typhoon-intensity-classification.philippines-typhoon-tracks
Philippines Typhoon Tracks
IBTrACS v04r01 Western Pacific basin tracks joined with intensity classification.
Fields
storm_id — IBTrACS Storm ID
time — ISO timestamp (3-hourly)
lat, lon — position (decimal degrees)
wind_kt — max sustained wind (knots)
WMO_PRES — min central pressure (mb, may be NaN)
intensity — Saffir-Simpson class (0=TD/TS, 1=Cat 1-2, 2=Cat 3-4, 3=Cat 5)
Stats
26,786 rows
895 storms
1980–2025
Source
IBTrACS v04r01… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/philippines-typhoon-tracks.isan-phonetic-dictionary
Isan Phonetic Dictionary Dataset
Summary
This dataset is a phonetic dictionary focused on Isan (Northeastern Thai) pronunciations. It is structured to handle linguistic complexities such as:
Phonetic Variations (เสียงแปร): Words that have multiple valid pronunciations without changing the meaning.
Homographs (คำพ้องรูป): Words that are spelled the same but have different pronunciations and meanings depending on the context.
The data is provided in TSV (Tab-Separated… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/isan-phonetic-dictionary.Typhoon
