datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-elderly-asr
Final gathered Persian elderly speech
Final corpus: 1,980 train / 294 validation / 329 test chunks. Another 956 uncertain chunks are quarantined under portable/review/ and excluded from these splits. This revision replaces the earlier gathered corpus; earlier data remains available through repository commit history.
80 paired recordings from four speaker folders. Reference transcripts were aligned with a historical Persian Wav2Vec2-base checkpoint, then cut at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/AliAvd/persian-elderly-asr.thai-elderly-speech
Thai Elderly Speech Dataset (Combined Evaluation Set)
This dataset contains evaluation recordings for Thai elderly speech, combined from Healthcare and Smarthome domains.
Dataset Structure
After extracting Combined_Dataset.zip, the directory structure will look like this:
Combined_Dataset/
├── Train/ # 80% of the dataset (15,360 files)
│ ├── Accuracy_100/ # Files with 100% baseline accuracy
│ ├── Accuracy_50_99/ #… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/thai-elderly-speech.
