datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quran-tajweed-phonetics
The complete phonetic layer of the Quran in the riwaya of Hafs 'an
'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every
phone carrying its tajweed attribution: madd class with its transmitted
length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt,
the seventeen sifat, and the rule that produced it.
Built and maintained by Quran Lab, a waqf building open technology in
the service of the Quran.
How it was built and verified
Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.phonetico-speech
Phonetico Speech
v2605 · Tigrinya · 14.7 hours · CC-BY-4.0
Phonetico Speech is a speech corpus for automatic speech recognition (ASR) in Ethiopian languages. Each language is available as a separate config. Load only what you need. v2605 contains 14.7 hours of transcribed Tigrinya audio.
This dataset is part of a long-term effort to build foundational speech technology for Ethiopian languages.
Dataset Summary
Language
Tigrinya (tir, ISO 639-3)
Script… See the full description on the dataset page: https://huggingface.co/datasets/phoneticoai/phonetico-speech.Phonetic
[!NOTE]
Dataset origin: https://github.com/Ofis-publik-ar-brezhoneg/phonetic-breton-corpus
Description
Base de données phonétique du breton - Diaz roadennoù fonetek ar brezhonegEnviron 4h30 d'audio.
