datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-elderly-asr
Final gathered Persian elderly speech
Final corpus: 1,980 train / 294 validation / 329 test chunks. Another 956 uncertain chunks are quarantined under portable/review/ and excluded from these splits. This revision replaces the earlier gathered corpus; earlier data remains available through repository commit history.
80 paired recordings from four speaker folders. Reference transcripts were aligned with a historical Persian Wav2Vec2-base checkpoint, then cut at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/AliAvd/persian-elderly-asr.persian-asr-dataset
Ganjoor Persian Speech Dataset
مجموعهای فارسی برای بازشناسی گفتار که از محتوای صوتی و متنهای متناظر وبسایت گنجور گردآوری و به قالب Hugging Face تبدیل شده است. این مجموعه شامل گفتار سالمندان نیست و از دیتاست اختصاصی سالمندان پروژه مستقل است.
This is a Persian automatic speech recognition dataset derived from aligned audio and text crawled from Ganjoor. It is separate from the project's elderly-speech dataset.
Dataset structure
Split: train
Samples: 5,036
Audio:… See the full description on the dataset page: https://huggingface.co/datasets/AliAvd/persian-asr-dataset.
