CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /navigation-corpus-dagbani-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Dag Speech Segments (sentence splitting) 52799 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-dagbani-speech.audioautomatic-speech-recognition10K<n<100K0 likes978 downloads2mo agoHugging Face02ghanaopenai /navigation-corpus-twi-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Speech Segments (sentence splitting) 52562 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-twi Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-twi-speech.audioautomatic-speech-recognition10K<n<100K0 likes932 downloads2mo agoHugging Face03ghanaopenai /navigation-corpus-ewe-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ewe Speech Segments (sentence splitting) 49348 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-ewe Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/navigation-corpus-ewe-speech.audioautomatic-speech-recognition10K<n<100K0 likes558 downloads3mo agoHugging Face04ghanaopenai /navigation-corpus-speech-full-dagbani This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana TTS Navigation Corpus — Dagbani Synthetic speech dataset for navigation. Structure audio/ – all .wav audio files text/ – matching .txt files with transcriptions metadata.csv – full metadata table audiotext-to-speech1K<n<10K0 likes307 downloads3mo agoHugging Face05ghananlpcommunity /navigation-corpus-dagbani-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Dag Speech Segments (sentence splitting) 52799 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani Full-file CTC forced… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/navigation-corpus-dagbani-speech.audioautomatic-speech-recognition10K<n<100K0 likes290 downloads3mo agoHugging Face06ghananlpcommunity /navigation-corpus-twi-speech This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Speech Segments (sentence splitting) 52562 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-twi Full-file CTC forced alignment… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/navigation-corpus-twi-speech.audioautomatic-speech-recognition10K<n<100K0 likes279 downloads3mo agoHugging Face07ghanaopenai /navigation-corpus-speech-full-twi This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana TTS Navigation Corpus — Twi Synthetic speech dataset for navigation. Structure audio/ – all .wav audio files text/ – matching .txt files with transcriptions metadata.csv – full metadata table audiotext-to-speech1K<n<10K0 likes194 downloads3mo agoHugging Face08ghanaopenai /navigation-corpus-speech-full-ewe This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana TTS Navigation Corpus — Ewe Synthetic speech dataset for navigation. Structure audio/ – all .wav audio files text/ – matching .txt files with transcriptions metadata.csv – full metadata table audiotext-to-speech1K<n<10K0 likes127 downloads3mo agoHugging Face09fiifinketia /navigation-corpus-ewe-speech Ewe Speech Segments (sentence splitting) 49348 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-ewe Full-file CTC forced alignment (MMS-300M) for word-level timestamps Sentence-boundary splits (. ? !) — long sentences re-chunked to 16 words Leading/trailing silence trimmed with VAD (-40 dBFS threshold) Filtered: min 1.0s, max 15.0s Original sample rate preserved Usage from… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/navigation-corpus-ewe-speech.audioautomatic-speech-recognition10K<n<100K1 likes46 downloads6mo agoHugging Face10fiifinketia /navigation-corpus-dagbani-speech Dag Speech Segments (sentence splitting) 52799 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/navigation-corpus-speech-full-dagbani Full-file CTC forced alignment (MMS-300M) for word-level timestamps Sentence-boundary splits (. ? !) — long sentences re-chunked to 16 words Leading/trailing silence trimmed with VAD (-40 dBFS threshold) Filtered: min 1.0s, max 15.0s Original sample rate preserved Usage from… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/navigation-corpus-dagbani-speech.audioautomatic-speech-recognition10K<n<100K0 likes37 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.