datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.IEMO_Audio_Text_Mergedquran-audio-text
QuranLab — Verse-Aligned Quran Text + Recitation References
This dataset joins QuranLab's canonical Hafs Arabic text to its
per-ayah recitation references. Every row is one exact
(recitation_id, verse_key) pair: the Uthmani transcript, a search-friendly
Simple-Clean transcript, and the corresponding audio_url.
QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio-text.MSPI_Audio_Text_Mergedtext-2-audio-human-preference-benchmark
Text to Audio Human Benchmark
In this dataset, ~32k human responses collected in less than 1h using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
The annotators were asked Which voice is more friendly? and Which voice sounds more natural? respectively.
Check out the Benchmark!
Testing_audio_text_pairs_Sinhala_v1ACD-Audios-meta-to-text-v1random500_text_audio_arrowaudioliburuak_with_text_tagslow-resource-audio-textACD-Audios-text-tags-v1acg-audio-tags-and-text-generatedmy_audio_dataset_tags_textACD-Audios-text-tags-v2-Event-ClassificattionACD-Audios-text-tags-v2-Event-Classificattion-modifiedsanskrit_audio_dataset_under_30_taged_meta_to_text_from_edgehinglish_audio_with_textMSPP_Audio_Text_Mergedaudiobench_text_data
