CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /kasem-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Kasem Speech-Text Parallel Dataset Dataset Description This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes2.9k downloads3mo agoHugging Face02ghanaopenai /ga-speech-text-parallel-90k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Ga Speech-Text Parallel Dataset Dataset Description This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.audioautomatic-speech-recognition10K<n<100K0 likes1.4k downloads3mo agoHugging Face03ghanaopenai /twi-trigrams-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes1.3k downloads3mo agoHugging Face04michsethowusu /yoruba-speech-text-parallel Yoruba Speech-Text Parallel Dataset Dataset Description This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Yoruba - yo Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.audioautomatic-speech-recognition1M<n<10M3 likes447 downloads1y agoHugging Face05ghanaopenai /vagla-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Vagla Speech-Text Parallel Dataset Dataset Description This dataset contains 48605 parallel speech-text pairs for Vagla, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/vagla-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes376 downloads3mo agoHugging Face06michsethowusu /makhuwa-trigrams-speech-text-parallel Makhuwa Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 154253 parallel speech-text pairs for Makhuwa, a language spoken primarily in Mozambique. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Makhuwa - vmw Task: Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/makhuwa-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes280 downloads1y agoHugging Face07michsethowusu /twi-words-speech-text-parallel-400k Twi Words Speech-Text Parallel Dataset Dataset Description This dataset contains 413463 parallel speech-text pairs for Twi (Akan), a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Twi (Akan) - tw Task: Speech Recognition, Text-to-Speech Size: 413463 audio files >… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-words-speech-text-parallel-400k.audioautomatic-speech-recognition100K<n<1M1 likes249 downloads1y agoHugging Face08michsethowusu /vai-speech-text-parallel Vai Speech-Text Parallel Dataset Dataset Description This dataset contains 23286 parallel speech-text pairs for Vai, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Vai - vai Task: Speech Recognition, Text-to-Speech Size: 23286 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/vai-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes240 downloads1y agoHugging Face09michsethowusu /deg-speech-text-parallel Deg Speech-Text Parallel Dataset Dataset Description This dataset contains 125958 parallel speech-text pairs for Deg, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Deg - mzw Task: Speech Recognition, Text-to-Speech Size: 125958 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/deg-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes208 downloads1y agoHugging Face10michsethowusu /swahili-words-speech-text-parallel Swahili Words Speech-Text Parallel Dataset Dataset Description This dataset contains 411048 parallel speech-text pairs for Swahili, a widely spoken language in East Africa. The dataset consists of audio recordings paired with corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Swahili - sw Task: Speech Recognition, Text-to-Speech Size: 411048 audio files > 1KB… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-words-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M1 likes208 downloads1y agoHugging Face11prince4332 /twi-words-speech-text-parallel-splittimeseries100K<n<1M0 likes207 downloads1y agoHugging Face12fiifinketia /twi-trigrams-speech-text-parallel Twi Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Twi - twi Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/twi-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes119 downloads6mo agoHugging Face13michsethowusu /chichewa-trigrams-speech-text-parallel Chichewa Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 132549 parallel speech-text pairs for Chichewa, a language spoken primarily in Malawi. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Chichewa - ny Task: Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M1 likes79 downloads1y agoHugging Face14tiny-aya-translate /tr-hi-parallel-text TR↔HI Parallel Text 65,662 aligned text triples — English pivot plus Turkish and Hindi (en_text / tr_text / hi_text), each tagged with its source. This is the text layer the speech corpora were synthesised from: these sentences were sent to TTS to produce tr-hi-parallel-speech-v2, which was then Mimi-encoded into tr-hi-mimi-encoded. Text-only, ~10 MB, no audio. Sources include FLORES, OPUS-100, and machine-translated conversational data — check source per row, since the licence… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-text.texttranslation10K<n<100K0 likes25 downloads2mo agoHugging Face15Thermostatic /texts_parallel_corpus_europarl_english_spanish Dataset Card for Dataset Name A massive parallel corpus of English-Spanish pairs. It hasn't a specified license, but there doesn't seem to be any copyrighted material in the corpus. I have personally merged rows using a pseudo-random algorithm making the dataset useful in training LLMs, reducing the risk of overfitting. Dataset Details Dataset Description Curated by: Philipp Koehn Funded by [optional]: In part funded by the European Commission (7th… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/texts_parallel_corpus_europarl_english_spanish.texttranslation10K<n<100K2 likes22 downloads2y agoHugging Face16michsethowusu /twi-trigrams-speech-text-parallel-cleanaudio10K<n<100K0 likes22 downloads10mo agoHugging Face17alakxender /dhivehi-legal-text-parallelgated Dhivehi-English Legal Parallel Corpus Dataset Description A high-quality parallel corpus of 56,556 Dhivehi-English sentence pairs extracted from 200 Maldivian legal documents. This dataset is deduplicated and cleaned for machine translation and bilingual model training. Dataset Summary Languages: Dhivehi (dv) ↔ English (en) Total Pairs: 56,556 Source Laws: 200 Duplicates Removed: 31,235 Average Dhivehi Length: 173.6 characters Average English… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-legal-text-parallel.tabulartranslation10K<n<100K0 likes12 downloads9mo agoHugging Face18tonative /swahili-parallel-text-extension Swahili–English Extension of Kidaw’ida, Kalenjin and Dholuo Parallel Corpus Dataset Description This dataset is an extension of the parallel corpora introduced in Building Low-Resource African Language Corpora: A Case Study of Kidaw’ida, Kalenjin and Dholuo by: Ambrose N. Ambogho Quin Awuor Andrew Kipkebut Lilian Wanzare Vivian Oloo The original dataset provides Swahili–[Kidaw’ida/Kalenjin/Dholuo] parallel sentence pairs. English was not included in the original corpus.… See the full description on the dataset page: https://huggingface.co/datasets/tonative/swahili-parallel-text-extension.text10K<n<100K0 likes12 downloads7mo agoHugging Face19ghananlpcommunity /twi-english-parallel-text This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi–English Parallel Text Dataset A parallel corpus of Asante Twi and English sentence pairs compiled by Ghana NLP Community. Sources Source Description youversion_bible Verse-level parallel pairs scraped… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-english-parallel-text.texttranslation100K<n<1M0 likes10 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.