CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /kasem-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Kasem Speech-Text Parallel Dataset Dataset Description This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes2.9k downloads3mo agoHugging Face02ghanaopenai /ga-speech-text-parallel-90k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Ga Speech-Text Parallel Dataset Dataset Description This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.audioautomatic-speech-recognition10K<n<100K0 likes1.4k downloads3mo agoHugging Face03ghanaopenai /twi-trigrams-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes1.3k downloads3mo agoHugging Face04ghanaopenai /ewe-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 48775 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes892 downloads3mo agoHugging Face05ghanaopenai /dagbani-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 53410 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes631 downloads3mo agoHugging Face06ghanaopenai /fante-speech-text-multispeaker_lds Fante Speech-Text Multispeaker Dataset (LDS) Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations. Dataset Statistics Split Clips Hours Talks Train 29,992 58.32 405 Eval 2,028 4.09 28 Total 32,020 62.41 433 Features audio: 16 kHz mono FLAC sentence-level clips text: Fante transcript (sentence-aligned) talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-speech-text-multispeaker_lds.audioautomatic-speech-recognition10K<n<100K1 likes625 downloads2mo agoHugging Face07Appenlimited /700h-tr-turkish-text-to-speechaudioautomatic-speech-recognition1K<n<10K17 likes505 downloads1y agoHugging Face08michsethowusu /yoruba-speech-text-parallel Yoruba Speech-Text Parallel Dataset Dataset Description This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Yoruba - yo Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.audioautomatic-speech-recognition1M<n<10M3 likes447 downloads1y agoHugging Face09ghanaopenai /vagla-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Vagla Speech-Text Parallel Dataset Dataset Description This dataset contains 48605 parallel speech-text pairs for Vagla, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/vagla-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes376 downloads3mo agoHugging Face10ghanaopenai /ga-multispeaker-speech-text-20k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Ga Multispeaker Audio Transcribed Dataset Overview The Ga… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-multispeaker-speech-text-20k.audioautomatic-speech-recognition10K<n<100K1 likes314 downloads3mo agoHugging Face11ghanaopenai /fante-multispeaker_speech-text-20k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Fante Multispeaker Audio Transcribed Dataset Overview The Fante… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-multispeaker_speech-text-20k.audioautomatic-speech-recognition10K<n<100K1 likes306 downloads3mo agoHugging Face12quranlab /quran-audio-text QuranLab — Verse-Aligned Quran Text + Recitation References This dataset joins QuranLab's canonical Hafs Arabic text to its per-ayah recitation references. Every row is one exact (recitation_id, verse_key) pair: the Uthmani transcript, a search-friendly Simple-Clean transcript, and the corresponding audio_url. QuranLab is a volunteer effort. Our aim is to present these works carefully and at high quality, and to help them travel faithfully — in the spirit in which they were… See the full description on the dataset page: https://huggingface.co/datasets/quranlab/quran-audio-text.tabularautomatic-speech-recognition100K<n<1M1 likes298 downloads2mo agoHugging Face13risaleinur /risalei-nur-text-audio Risale-i Nur Text–Audio Kaynak · Source: RNK Neşriyat — yazılı izinle · used with written permission. Her satırda gerçek insan okuması ile o sesin kanonik metni birlikte bulunur. Sesler dış bağlantı değildir: WAV baytları Parquet dosyalarının içindedir. Kaynak sitesi veya başka bir ses sunucusu gerekmez. Each row pairs a human reading with its canonical transcript. Audio is stored as WAV bytes inside the Parquet files; no source website or external audio server is required.… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risalei-nur-text-audio.audioautomatic-speech-recognition10K<n<100K1 likes285 downloads2mo agoHugging Face14michsethowusu /makhuwa-trigrams-speech-text-parallel Makhuwa Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 154253 parallel speech-text pairs for Makhuwa, a language spoken primarily in Mozambique. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Makhuwa - vmw Task: Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/makhuwa-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes280 downloads1y agoHugging Face15michsethowusu /twi-words-speech-text-parallel-400k Twi Words Speech-Text Parallel Dataset Dataset Description This dataset contains 413463 parallel speech-text pairs for Twi (Akan), a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Twi (Akan) - tw Task: Speech Recognition, Text-to-Speech Size: 413463 audio files >… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-words-speech-text-parallel-400k.audioautomatic-speech-recognition100K<n<1M1 likes249 downloads1y agoHugging Face16michsethowusu /vai-speech-text-parallel Vai Speech-Text Parallel Dataset Dataset Description This dataset contains 23286 parallel speech-text pairs for Vai, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Vai - vai Task: Speech Recognition, Text-to-Speech Size: 23286 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/vai-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes240 downloads1y agoHugging Face17adiren7 /darija_speech_to_textaudioautomatic-speech-recognition10K<n<100K13 likes213 downloads2y agoHugging Face18michsethowusu /deg-speech-text-parallel Deg Speech-Text Parallel Dataset Dataset Description This dataset contains 125958 parallel speech-text pairs for Deg, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Deg - mzw Task: Speech Recognition, Text-to-Speech Size: 125958 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/deg-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes208 downloads1y agoHugging Face19michsethowusu /swahili-words-speech-text-parallel Swahili Words Speech-Text Parallel Dataset Dataset Description This dataset contains 411048 parallel speech-text pairs for Swahili, a widely spoken language in East Africa. The dataset consists of audio recordings paired with corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Swahili - sw Task: Speech Recognition, Text-to-Speech Size: 411048 audio files > 1KB… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-words-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M1 likes208 downloads1y agoHugging Face20ghananlpcommunity /fante-speech-text-multispeaker_lds Fante Speech-Text Multispeaker Dataset (LDS) Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations. Dataset Statistics Split Clips Hours Talks Train 29,992 58.32 405 Eval 2,028 4.09 28 Total 32,020 62.41 433 Features audio: 16 kHz mono FLAC sentence-level clips text: Fante transcript (sentence-aligned) talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/fante-speech-text-multispeaker_lds.audioautomatic-speech-recognition10K<n<100K0 likes139 downloads2mo agoHugging Face21vnahata /waxal-audio-text-retrieval WAXAL speech–text retrieval (MTEB) Multilingual speech↔text retrieval over 16 Sub-Saharan African languages, derived from WAXAL (Google and partners). Most of these languages have no presence in mteb's existing multilingual audio tasks, which skew European and South/East Asian. Prepared as WaxalA2TRetrieval and WaxalT2ARetrieval. Contents One config per language, each with id, audio, text, speaker_id, gender, language. 1,722 utterances total. code language… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/waxal-audio-text-retrieval.audioautomatic-speech-recognition1K<n<10K0 likes132 downloads25d agoHugging Face22pujanpaudel /nepali_speech_to_text Nepali Speech-to-Text Dataset This repository contains a dataset for Automatic Speech Recognition (ASR) in the Nepali language. The dataset is designed for supervised learning tasks and includes audio files along with their corresponding transcriptions. The audio samples have been collected from various open-source platforms and other publicly available sources on the internet. Each audio file has an average length of 15 seconds and has been converted into a consistent WAV format… See the full description on the dataset page: https://huggingface.co/datasets/pujanpaudel/nepali_speech_to_text.audioautomatic-speech-recognition1K<n<10K1 likes122 downloads2y agoHugging Face23fiifinketia /twi-trigrams-speech-text-parallel Twi Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Twi - twi Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/twi-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes119 downloads6mo agoHugging Face24Sin2pi /JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample. -Unnecessary and inaccurate punctuation have been removed. -Text has been normalized. Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja. audioautomatic-speech-recognition100K<n<1M9 likes103 downloads9mo agoHugging Face25BrunoHays /darija-speech-to-text Speech To Text Darija dataset Reupload of adiren7/darija_speech_to_text audioautomatic-speech-recognition1K<n<10K6 likes102 downloads2y agoHugging Face26SDAIANCAI /Ar-En-Code-Switching-Textual-Dataset ArE-CSTD: Arabic-English Code-Switching Textual Dataset The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”. This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4. TXT Files There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.texttext-generation100K<n<1M2 likes93 downloads2y agoHugging Face27fiifinketia /ewe-bible-audio-text-tts Twi 16-Word Speech Segments 48775 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for word-level timestamps Words grouped into 16-word segments Leading/trailing silence trimmed with VAD (-40 dBFS threshold) Filtered: min 1.0s, max 15.0s Original sample rate preserved (24kHz) Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/ewe-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes90 downloads6mo agoHugging Face28michsethowusu /chichewa-trigrams-speech-text-parallel Chichewa Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 132549 parallel speech-text pairs for Chichewa, a language spoken primarily in Malawi. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Chichewa - ny Task: Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M1 likes79 downloads1y agoHugging Face29EMINES /Tamazight-Speech-to-Arabic-Text Tamazight-Arabic Speech Recognition Dataset Overview This is the EMINES organization-hosted version of the Tamazight-Arabic Speech Recognition Dataset, synchronized with the original dataset. It contains ~15.5 hours of Tamazight speech (Tachelhit dialect) paired with Arabic transcriptions, designed for developing ASR and translation systems. Quick Start from datasets import load_dataset # Load the dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/EMINES/Tamazight-Speech-to-Arabic-Text.audioautomatic-speech-recognition10K<n<100K6 likes72 downloads2y agoHugging Face30Reza2kn /persian-asr-text-2.69M-deduped 🗂️ persian-asr-text-2.69M-deduped English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Deduplicated Persian ASR text dataset used by the training stack. پیکرهٔ متنی فارسیِ حذف‌تکرارشده برای ساخت واژگان، مدل‌سازی زبانی و پشتیبانی از آموزش ASR. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 4 files; approximately 109.64 MB 4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.tabularautomatic-speech-recognition1M<n<10M0 likes68 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.