datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-elderly-asr
Final gathered Persian elderly speech
Final corpus: 1,980 train / 294 validation / 329 test chunks. Another 956 uncertain chunks are quarantined under portable/review/ and excluded from these splits. This revision replaces the earlier gathered corpus; earlier data remains available through repository commit history.
80 paired recordings from four speaker folders. Reference transcripts were aligned with a historical Persian Wav2Vec2-base checkpoint, then cut at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/AliAvd/persian-elderly-asr.Masri-Elders
Masri Elders Speech-to-Text Dataset
Dataset Description
This dataset contains Egyptian Arabic (Masri) speech recordings from elderly speakers, paired with their transcriptions. It is designed to help improve Automatic Speech Recognition (ASR) models for this specific dialect and demographic, which is often underrepresented in standard datasets.
Language: Egyptian Arabic (Masri)
Demographic: Elders
Sampling Rate: 16kHz (Mono)
Format: WAV audio + Text transcripts… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/Masri-Elders.farahan-yoruba-elder-corpus
Apakose Ezekiel Imoleayo — Farahàn
Appear. Be found.
Lagos, Nigeria · UNILAG Yoruba Studies ·
Graduating 2030
What I Build
Yoruba Oral Knowledge Corpus
I document primary-source Yoruba knowledge
directly from elder speakers in Lagos and
southwest Nigeria. Structured interviews
covering proverbs (Owe), oral history (Itan),
praise poetry (Oriki), and cultural knowledge
systems that no web scrape produces.
What makes this corpus different:… See the full description on the dataset page: https://huggingface.co/datasets/Apakose-Ezekiel/farahan-yoruba-elder-corpus.thai-elderly-speech
Thai Elderly Speech Dataset (Combined Evaluation Set)
This dataset contains evaluation recordings for Thai elderly speech, combined from Healthcare and Smarthome domains.
Dataset Structure
After extracting Combined_Dataset.zip, the directory structure will look like this:
Combined_Dataset/
├── Train/ # 80% of the dataset (15,360 files)
│ ├── Accuracy_100/ # Files with 100% baseline accuracy
│ ├── Accuracy_50_99/ #… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/thai-elderly-speech.
