CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /ghana-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Language Statistics Language Subset Segments Duration Akuapem_Twi Akuapem_Twi_twi 52,650 63.25h Anyin Anyin_any 5,568 13.24h Asante_Twi Asante_Twi_twi 143,383 200.02h Avatime Avatime_avn 9,956 21.62h Bassar_Ntcham… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-speech.audio1M<n<10M0 likes6.6k downloads29d agoHugging Face02ghanaopenai /ghana-speech-ipa Ghana Speech — Audio with IPA Transcripts Speech with both transcript forms: the original orthography and the IPA phoneme sequence read off the audio by ASR. Each language is a subset, with real train/validation splits. from datasets import load_dataset ds = load_dataset("ghanaopendata/ghana-speech-ipa", "Akuapem_Twi_twi", split="train") ds[0]["audio"] # decoded waveform, 16 kHz ds[0]["text"] # original orthography ds[0]["ipa"] # IPA phonemes 369,347 clips · ~747 h… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-speech-ipa.audiotext-to-speech100K<n<1M0 likes5.8k downloads1mo agoHugging Face03ghananlpcommunity /ghana-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Language Statistics Language Subset Segments Duration Akuapem_Twi Akuapem_Twi_twi 52,650 63.25h Anyin Anyin_any 5,568 13.24h Asante_Twi Asante_Twi_twi 143,383 200.02h Avatime Avatime_avn 9,956 21.62h Bassar_Ntcham… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-speech.audio1M<n<10M3 likes4.1k downloads2mo agoHugging Face04ghanaopenai /kasem-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Kasem Speech-Text Parallel Dataset Dataset Description This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes2.9k downloads3mo agoHugging Face05ghanaopenai /ghana-english-asr-2700hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. 🇬🇭 Ghana English ASR Dataset A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Automatic Speech Recognition (ASR) models on West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-asr-2700hrs.audioautomatic-speech-recognition100K<n<1M7 likes2.4k downloads3mo agoHugging Face06ghanaopenai /twi-health-asr-gemini-500hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Health Speech Dataset Gemini (500 hours) A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages, sourced from publicly available video content on health and wellness. Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.audioautomatic-speech-recognition10K<n<100K0 likes2k downloads2mo agoHugging Face07ghanaopenai /ghana-english-speech-600hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. 🇬🇭 Ghana English ASR Dataset A speech dataset of Ghanaian English extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Automatic Speech Recognition (ASR) models on West African English accents.… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-600hrs.audioautomatic-speech-recognition100K<n<1M1 likes1.7k downloads3mo agoHugging Face08ghanaopenai /new-twi-tts-aligned-ipa new-twi-tts-aligned + IPA phonemes ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme transcription for every clip, produced with ghananlpcommunity/ghana-speech-phoneme-asr. Audio included — this is self-contained, no join with the source dataset needed. Contents split clips hours phoneme units mean units/clip test 16,140 17.24 663,140 41.1 train 145,258 155.21 5,945,389 40.9 Columns column type meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.audioautomatic-speech-recognition100K<n<1M0 likes1.5k downloads2mo agoHugging Face09ghanaopenai /ga-speech-text-parallel-90k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Ga Speech-Text Parallel Dataset Dataset Description This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.audioautomatic-speech-recognition10K<n<100K0 likes1.4k downloads3mo agoHugging Face10ghanaopenai /asante-twi-bible-speech-text This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. audio10K<n<100K1 likes1.3k downloads3mo agoHugging Face11ghanaopenai /twi-trigrams-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes1.3k downloads3mo agoHugging Face12ghananlpcommunity /ghana-speech-eval ghana-speech-eval ⚠️ Evaluation only. This is a held-out benchmark for measuring ASR quality on Ghanaian languages. Please do not train or fine-tune on it — doing so contaminates the benchmark and makes reported scores meaningless. Every subset ships a single eval split. Multi-source ASR evaluation benchmark for Ghanaian languages. All subsets are 16 kHz mono with standardised schema (audio, text, language, country, length, iso, subset). Source groups… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-speech-eval.audio10K<n<100K0 likes1.2k downloads1mo agoHugging Face13ghanaopenai /ghana-english-tts-clean2 Ghana English TTS Filtered Clean v2 Filtered subset of ghananlpcommunity/ghana-english-tts-filtered using PANNs CNN14. Filtering Second-pass filtering with PANNs CNN14 (soundclassifier with music, applause, and speech tags): Keep if: music_prob ≤ 0.2 AND applause_prob ≤ 0.2 AND speech_prob ≥ 0.5 Batch size 32 on NVIDIA H200, float16 inference 282,096 / 303,204 kept (93.0%) Fields corrected_text: utterance text bytes: raw WAV bytes (16-bit PCM… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-tts-clean2.audiotext-to-speech100K<n<1M0 likes1.2k downloads2mo agoHugging Face14ghananlpcommunity /ghana-speech-ipa Ghana Speech — Audio with IPA Transcripts Speech with both transcript forms: the original orthography and the IPA phoneme sequence read off the audio by ASR. Each language is a subset, with real train/validation splits. from datasets import load_dataset ds = load_dataset("ghanaopendata/ghana-speech-ipa", "Akuapem_Twi_twi", split="train") ds[0]["audio"] # decoded waveform, 16 kHz ds[0]["text"] # original orthography ds[0]["ipa"] # IPA phonemes 369,347 clips · ~747 h… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-speech-ipa.audiotext-to-speech100K<n<1M0 likes1.1k downloads1mo agoHugging Face15ghanaopenai /navigation-corpus-dagbani-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Dag Speech Segments (sentence splitting) 52799 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-dagbani-speech.audioautomatic-speech-recognition10K<n<100K0 likes1.1k downloads2mo agoHugging Face16ghanaopenai /navigation-corpus-twi-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Speech Segments (sentence splitting) 52562 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-twi Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-twi-speech.audioautomatic-speech-recognition10K<n<100K0 likes1k downloads2mo agoHugging Face17ghanaopenai /twi-health-asr This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Health Speech Dataset A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages, sourced from publicly available video content on health and wellness. Created by Mich-Seth Owusu and… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr.audioautomatic-speech-recognition10K<n<100K0 likes930 downloads3mo agoHugging Face18ghanaopenai /ewe-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 48775 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes894 downloads3mo agoHugging Face19ghanaopenai /ghana-english-speech-ipa Ghanaian English Speech — Audio with IPA Transcripts Speech with both transcript forms: the original orthography and the IPA phoneme sequence read off the audio by ASR. Each language is a subset, with real train/validation splits. from datasets import load_dataset ds = load_dataset("ghanaopendata/ghana-english-speech-ipa", "English_eng", split="train") ds[0]["audio"] # decoded waveform, 16 kHz ds[0]["text"] # original orthography ds[0]["ipa"] # IPA phonemes 52,855… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-speech-ipa.audiotext-to-speech10K<n<100K0 likes789 downloads1mo agoHugging Face20ghananlpcommunity /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes759 downloads23d agoHugging Face21ghananlpcommunity /nsanku-tts-benchmark-audioaudio10K<n<100K2 likes738 downloads10d agoHugging Face22ghanaopenai /ghana-named-entities-tts-twi This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana Named Entities TTS — Twi A Twi-language speech dataset built from descriptions of Ghana named entities (people, places, organisations, and concepts). Each audio clip is a synthesised reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.audiotext-to-speech1K<n<10K0 likes711 downloads3mo agoHugging Face23ghanaopenai /new-twi-tts-aligned This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi TTS Dataset A speech dataset of Twi (Akan) extracted from Ghanaian news media broadcasts, designed for training and fine-tuning Text-To-Speech (TTS) models. 📂 Dataset Structure Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned.audio100K<n<1M0 likes703 downloads3mo agoHugging Face24ghananlpcommunity /twi-health-asr-gemini-500hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Health Speech Dataset Gemini (500 hours) A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages, sourced from publicly available video content on health and wellness. Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs.audioautomatic-speech-recognition10K<n<100K1 likes646 downloads3mo agoHugging Face25ghanaopenai /dagbani-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 53410 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes633 downloads3mo agoHugging Face26ghanaopenai /fante-speech-text-multispeaker_lds Fante Speech-Text Multispeaker Dataset (LDS) Sentence-level aligned Fante (fat) speech dataset sourced from the Church of Jesus Christ of Latter-day Saints General Conference translations. Dataset Statistics Split Clips Hours Talks Train 29,992 58.32 405 Eval 2,028 4.09 28 Total 32,020 62.41 433 Features audio: 16 kHz mono FLAC sentence-level clips text: Fante transcript (sentence-aligned) talk_id: Source conference talk… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/fante-speech-text-multispeaker_lds.audioautomatic-speech-recognition10K<n<100K1 likes629 downloads2mo agoHugging Face27ghanaopenai /navigation-corpus-ewe-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ewe Speech Segments (sentence splitting) 49348 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-ewe Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-ewe-speech.audioautomatic-speech-recognition10K<n<100K0 likes621 downloads3mo agoHugging Face28ghanaopenai /kumawood-speech-transcriptions Kumawood Speech Transcriptions Speech segments from Ghanaian films, each paired with the film's human-authored English subtitle and a machine Twi transcript. Total number of hours 249.8 hours Fields field meaning audio 16 kHz mono FLAC segment text English subtitle displayed during the segment (human-authored, recovered by OCR) twi_text Twi transcript from Google STT (ak) — machine output twi_words_per_sec transcript words per second of audio… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kumawood-speech-transcriptions.audioautomatic-speech-recognition100K<n<1M0 likes601 downloads23d agoHugging Face29ghanaopenai /ghana-speech-eval ghana-speech-eval ⚠️ Evaluation only. This is a held-out benchmark for measuring ASR quality on Ghanaian languages. Please do not train or fine-tune on it — doing so contaminates the benchmark and makes reported scores meaningless. Every subset ships a single eval split. Multi-source ASR evaluation benchmark for Ghanaian languages. All subsets are 16 kHz mono with standardised schema (audio, text, language, country, length, iso, subset). Source groups… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-speech-eval.audio10K<n<100K0 likes581 downloads29d agoHugging Face30ghanaopenai /twi-agriculture-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Agriculture Speech Dataset A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages, sourced from publicly available video content on agriculture.Created by Mich-Seth Owusu and published… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-agriculture-speech.audioautomatic-speech-recognition10K<n<100K0 likes580 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.