CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /ewe-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 48775 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes894 downloads3mo agoHugging Face02ghanaopenai /dagbani-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 53410 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes633 downloads3mo agoHugging Face03quranlab /quran-audio-text QuranLab — Verse-Aligned Quran Text + Recitation References This dataset joins QuranLab's canonical Hafs Arabic text to its per-ayah recitation references. Every row is one exact (recitation_id, verse_key) pair: the Uthmani transcript, a search-friendly Simple-Clean transcript, and the corresponding audio_url. QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio-text.tabularautomatic-speech-recognition100K<n<1M1 likes299 downloads2mo agoHugging Face04risaleinur /risalei-nur-text-audio Risale-i Nur Text–Audio Kaynak · Source: RNK Neşriyat — yazılı izinle · used with written permission. Her satırda gerçek insan okuması ile o sesin kanonik metni birlikte bulunur. Sesler dış bağlantı değildir: WAV baytları Parquet dosyalarının içindedir. Kaynak sitesi veya başka bir ses sunucusu gerekmez. Each row pairs a human reading with its canonical transcript. Audio is stored as WAV bytes inside the Parquet files; no source website or external audio server is required.… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risalei-nur-text-audio.audioautomatic-speech-recognition10K<n<100K1 likes292 downloads2mo agoHugging Face05vnahata /waxal-audio-text-retrieval WAXAL speech–text retrieval (MTEB) Multilingual speech↔text retrieval over 16 Sub-Saharan African languages, derived from WAXAL (Google and partners). Most of these languages have no presence in mteb's existing multilingual audio tasks, which skew European and South/East Asian. Prepared as WaxalA2TRetrieval and WaxalT2ARetrieval. Contents One config per language, each with id, audio, text, speaker_id, gender, language. 1,722 utterances total. code language… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/waxal-audio-text-retrieval.audioautomatic-speech-recognition1K<n<10K0 likes133 downloads26d agoHugging Face06Sin2pi /JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample. -Unnecessary and inaccurate punctuation have been removed. -Text has been normalized. Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja. audioautomatic-speech-recognition100K<n<1M9 likes103 downloads9mo agoHugging Face07fiifinketia /ewe-bible-audio-text-tts Twi 16-Word Speech Segments 48775 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for word-level timestamps Words grouped into 16-word segments Leading/trailing silence trimmed with VAD (-40 dBFS threshold) Filtered: min 1.0s, max 15.0s Original sample rate preserved (24kHz) Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/ewe-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes90 downloads6mo agoHugging Face08fiifinketia /dagbani-bible-audio-text-tts Twi 16-Word Speech Segments 53410 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for word-level timestamps Words grouped into 16-word segments Leading/trailing silence trimmed with VAD (-40 dBFS threshold) Filtered: min 1.0s, max 15.0s Original sample rate preserved (24kHz) Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/dagbani-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes21 downloads6mo agoHugging Face09snorbyte /indic-text-audio-samplegated Dataset Card for Indic Text Audio Sample Dataset Dataset Details Dataset Description The IndicTextAudioSample Dataset is a multilingual, text-speech pair sample dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi. Curated by: snorbyte Funded by: snorbyte Shared by: snorbyte Language(s) (NLP): hi, ta, te, pa, ml, kn, bn, gu, mr License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-text-audio-sample.audioaudio-classification10K<n<100K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.