CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /kasem-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Kasem Speech-Text Parallel Dataset Dataset Description This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes2.9k downloads3mo agoHugging Face02ghanaopenai /ga-speech-text-parallel-90k This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github Ga Speech-Text Parallel Dataset Dataset Description This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.audioautomatic-speech-recognition10K<n<100K0 likes1.4k downloads3mo agoHugging Face03ghanaopenai /twi-trigrams-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes1.3k downloads3mo agoHugging Face04michsethowusu /yoruba-speech-text-parallel Yoruba Speech-Text Parallel Dataset Dataset Description This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Yoruba - yo Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.audioautomatic-speech-recognition1M<n<10M3 likes447 downloads1y agoHugging Face05ghanaopenai /vagla-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Vagla Speech-Text Parallel Dataset Dataset Description This dataset contains 48605 parallel speech-text pairs for Vagla, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/vagla-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes376 downloads3mo agoHugging Face06michsethowusu /makhuwa-trigrams-speech-text-parallel Makhuwa Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 154253 parallel speech-text pairs for Makhuwa, a language spoken primarily in Mozambique. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Makhuwa - vmw Task: Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/makhuwa-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes280 downloads1y agoHugging Face07michsethowusu /twi-words-speech-text-parallel-400k Twi Words Speech-Text Parallel Dataset Dataset Description This dataset contains 413463 parallel speech-text pairs for Twi (Akan), a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Twi (Akan) - tw Task: Speech Recognition, Text-to-Speech Size: 413463 audio files >… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-words-speech-text-parallel-400k.audioautomatic-speech-recognition100K<n<1M1 likes249 downloads1y agoHugging Face08michsethowusu /vai-speech-text-parallel Vai Speech-Text Parallel Dataset Dataset Description This dataset contains 23286 parallel speech-text pairs for Vai, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Vai - vai Task: Speech Recognition, Text-to-Speech Size: 23286 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/vai-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes240 downloads1y agoHugging Face09ImruQays /Rasaif-Classical-Arabic-English-Parallel-texts Introduction This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period. Content Details Contained within this dataset are English translations of the following texts, sourced from the Rasaif website: A Muslim Manual of War Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.texttranslation10K<n<100K8 likes209 downloads3y agoHugging Face10michsethowusu /deg-speech-text-parallel Deg Speech-Text Parallel Dataset Dataset Description This dataset contains 125958 parallel speech-text pairs for Deg, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Deg - mzw Task: Speech Recognition, Text-to-Speech Size: 125958 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/deg-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes208 downloads1y agoHugging Face11michsethowusu /swahili-words-speech-text-parallel Swahili Words Speech-Text Parallel Dataset Dataset Description This dataset contains 411048 parallel speech-text pairs for Swahili, a widely spoken language in East Africa. The dataset consists of audio recordings paired with corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Swahili - sw Task: Speech Recognition, Text-to-Speech Size: 411048 audio files > 1KB… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-words-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M1 likes208 downloads1y agoHugging Face12prince4332 /twi-words-speech-text-parallel-splittimeseries100K<n<1M0 likes207 downloads1y agoHugging Face13Ghana-NLP /ENGLISH_TWI_PARALLEL_TEXT GhanaNLP Twi and English Parallel Data Twi_to_English • 1 MB • XLS English_to_Twi • 1 MB • XLS The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.text1K<n<10K3 likes158 downloads10mo agoHugging Face14fiifinketia /twi-trigrams-speech-text-parallel Twi Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Twi - twi Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/twi-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M0 likes119 downloads6mo agoHugging Face15michsethowusu /chichewa-trigrams-speech-text-parallel Chichewa Trigrams Speech-Text Parallel Dataset Dataset Description This dataset contains 132549 parallel speech-text pairs for Chichewa, a language spoken primarily in Malawi. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Chichewa - ny Task: Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-trigrams-speech-text-parallel.audioautomatic-speech-recognition100K<n<1M1 likes79 downloads1y agoHugging Face16Ghana-NLP /FANTE_ENGLISH_PARALLEL_TEXT GhanaNLP Fante and English Data Fante_to_English • 1 MB • XLS The GhanaNLP Fante dataset contains sentence pairs in Fante and English, designed to support translation models between these two languages. Fante is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for language processing tools. The… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/FANTE_ENGLISH_PARALLEL_TEXT.text1K<n<10K0 likes70 downloads10mo agoHugging Face17Ghana-NLP /TWI_ENGLISH_PARALLEL_TEXT GhanaNLP Twi and English Parallel Data Twi_to_English • 1 MB • XLS English_to_Twi • 1 MB • XLS The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/TWI_ENGLISH_PARALLEL_TEXT.text1K<n<10K1 likes62 downloads10mo agoHugging Face18Ghana-NLP /EWE_ENGLISH_PARALLEL_TEXT GhanaNLP Ewe and English Data Ewe_to_English • 1 MB • XLS The GhanaNLP Ewe dataset contains sentence pairs in Ewe and English, designed to support translation models between these two languages. Ewe is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for language processing tools. The sentence… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/EWE_ENGLISH_PARALLEL_TEXT.text1K<n<10K0 likes57 downloads10mo agoHugging Face19Ghana-NLP /GA_ENGLISH_PARALLEL_TEXT GhanaNLP Ga and English Data Ga_to_English • 1 MB • XLS The GhanaNLP Ga dataset contains sentence pairs in Ga and English, designed to support translation models between these two languages. Ga is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for language processing tools. The sentence… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/GA_ENGLISH_PARALLEL_TEXT.text1K<n<10K0 likes38 downloads10mo agoHugging Face20Ghana-NLP /KUSAAL_ENGLISH_PARALLEL_TEXT GhanaNLP Kusaal and English Data Kusaal_to_English • 1 MB • XLS The GhanaNLP Kusaal dataset contains sentence pairs in Kusaal and English, designed to support translation models between these two languages. Kusaal is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for language processing tools. The… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/KUSAAL_ENGLISH_PARALLEL_TEXT.text1K<n<10K0 likes36 downloads10mo agoHugging Face21ImruQays /Quran-Classical-Arabic-English-Parallel-texts Introduction This dataset presents a collection of parallel texts of the Holy Quran in Arabic (Imla'ei & Uthmanic scripts) alongside 17 different English translations. Contents The dataset includes the Holy Quran in Classical Arabic with the following English translations: "al-Qur’ân: A Contemporary Translation" by Ahmed Ali "Kanz-ul-Iman" by Ahmed Raza Khan "The Koran Interpreted" by Arthur John Arberry "The Message of The Qur'an" by Muhammad Asad "Qur'an English… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Quran-Classical-Arabic-English-Parallel-texts.texttranslation1K<n<10K6 likes28 downloads3y agoHugging Face22tiny-aya-translate /tr-hi-parallel-text TR↔HI Parallel Text 65,662 aligned text triples — English pivot plus Turkish and Hindi (en_text / tr_text / hi_text), each tagged with its source. This is the text layer the speech corpora were synthesised from: these sentences were sent to TTS to produce tr-hi-parallel-speech-v2, which was then Mimi-encoded into tr-hi-mimi-encoded. Text-only, ~10 MB, no audio. Sources include FLORES, OPUS-100, and machine-translated conversational data — check source per row, since the licence… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-text.texttranslation10K<n<100K0 likes25 downloads2mo agoHugging Face23ImruQays /Thaqalayn-Classical-Arabic-English-Parallel-texts Introduction This dataset represents a comprehensive collection of parallel Arabic-English texts from the Thaqalayn Hadith Library, a premier source for exploring the classical hadith tradition of the Imāmī Shia Muslim school of thought. The library focuses on making primary historical sources accessible, serving as a bridge between past wisdom and contemporary study. The dataset features translations of significant classical Imāmī hadith texts, allowing for a deep dive into the… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Thaqalayn-Classical-Arabic-English-Parallel-texts.texttranslation10K<n<100K8 likes24 downloads3y agoHugging Face24Thermostatic /texts_parallel_corpus_europarl_english_spanish Dataset Card for Dataset Name A massive parallel corpus of English-Spanish pairs. It hasn't a specified license, but there doesn't seem to be any copyrighted material in the corpus. I have personally merged rows using a pseudo-random algorithm making the dataset useful in training LLMs, reducing the risk of overfitting. Dataset Details Dataset Description Curated by: Philipp Koehn Funded by [optional]: In part funded by the European Commission (7th… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/texts_parallel_corpus_europarl_english_spanish.texttranslation10K<n<100K2 likes22 downloads2y agoHugging Face25DNivalis /parallel-complexity-med-texttabular10K<n<100K0 likes22 downloads1y agoHugging Face26michsethowusu /twi-trigrams-speech-text-parallel-cleanaudio10K<n<100K0 likes22 downloads10mo agoHugging Face27narinzar /parallel-image-text-dataset-builder parallel-image-text-dataset-builder (sample) A small representative sample from the parallel-image-text-dataset-builder pipeline: it ingests image-text pairs, removes near-duplicates with perceptual-hash (dhash) LSH-style bucketing, filters weak pairs by CLIP image-text similarity, and writes fixed-size WebDataset-style tar shards. Contents shard-00002.tar - one WebDataset-style shard (536 samples). Each sample is two members sharing a key: {key}.jpg (image) and… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/parallel-image-text-dataset-builder.tabularimage-to-textn<1K0 likes13 downloads3mo agoHugging Face28alakxender /dhivehi-legal-text-parallelgated Dhivehi-English Legal Parallel Corpus Dataset Description A high-quality parallel corpus of 56,556 Dhivehi-English sentence pairs extracted from 200 Maldivian legal documents. This dataset is deduplicated and cleaned for machine translation and bilingual model training. Dataset Summary Languages: Dhivehi (dv) ↔ English (en) Total Pairs: 56,556 Source Laws: 200 Duplicates Removed: 31,235 Average Dhivehi Length: 173.6 characters Average English… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-legal-text-parallel.tabulartranslation10K<n<100K0 likes12 downloads9mo agoHugging Face29tonative /swahili-parallel-text-extension Swahili–English Extension of Kidaw’ida, Kalenjin and Dholuo Parallel Corpus Dataset Description This dataset is an extension of the parallel corpora introduced in Building Low-Resource African Language Corpora: A Case Study of Kidaw’ida, Kalenjin and Dholuo by: Ambrose N. Ambogho Quin Awuor Andrew Kipkebut Lilian Wanzare Vivian Oloo The original dataset provides Swahili–[Kidaw’ida/Kalenjin/Dholuo] parallel sentence pairs. English was not included in the original corpus.… See the full description on the dataset page: https://huggingface.co/datasets/tonative/swahili-parallel-text-extension.text10K<n<100K0 likes12 downloads7mo agoHugging Face30shangzx /Chinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat Chinese-English Parallel Translation Corpus (Chinese Source Text & English Translation) A Chinese–English parallel corpus resource for translation and cross-lingual alignment applications, providing one-to-one bilingual text pairs: Chinese source texts aligned with their corresponding English translations. The data covers common writing styles and domains, making it suitable for parallel alignment, translation modeling, and cross-lingual representation learning. It supports… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Chinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat.texttext-classificationn<1K0 likes12 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.