datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kasem-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Kasem Speech-Text Parallel Dataset
Dataset Description
This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.ga-speech-text-parallel-90k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ga Speech-Text Parallel Dataset
Dataset Description
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ga-speech-text-parallel-90k.twi-trigrams-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.yoruba-speech-text-parallel
Yoruba Speech-Text Parallel Dataset
Dataset Description
This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Yoruba - yo
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/yoruba-speech-text-parallel.vagla-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Vagla Speech-Text Parallel Dataset
Dataset Description
This dataset contains 48605 parallel speech-text pairs for Vagla, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/vagla-speech-text-parallel.makhuwa-trigrams-speech-text-parallel
Makhuwa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 154253 parallel speech-text pairs for Makhuwa, a language spoken primarily in Mozambique. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Makhuwa - vmw
Task: Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/makhuwa-trigrams-speech-text-parallel.twi-words-speech-text-parallel-400k
Twi Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 413463 parallel speech-text pairs for Twi (Akan), a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi (Akan) - tw
Task: Speech Recognition, Text-to-Speech
Size: 413463 audio files >… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/twi-words-speech-text-parallel-400k.vai-speech-text-parallel
Vai Speech-Text Parallel Dataset
Dataset Description
This dataset contains 23286 parallel speech-text pairs for Vai, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Vai - vai
Task: Speech Recognition, Text-to-Speech
Size: 23286 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/vai-speech-text-parallel.Rasaif-Classical-Arabic-English-Parallel-texts
Introduction
This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period.
Content Details
Contained within this dataset are English translations of the following texts, sourced from the Rasaif website:
A Muslim Manual of War
Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.deg-speech-text-parallel
Deg Speech-Text Parallel Dataset
Dataset Description
This dataset contains 125958 parallel speech-text pairs for Deg, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Deg - mzw
Task: Speech Recognition, Text-to-Speech
Size: 125958 audio files > 1KB (small/corrupted… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/deg-speech-text-parallel.swahili-words-speech-text-parallel
Swahili Words Speech-Text Parallel Dataset
Dataset Description
This dataset contains 411048 parallel speech-text pairs for Swahili, a widely spoken language in East Africa. The dataset consists of audio recordings paired with corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Swahili - sw
Task: Speech Recognition, Text-to-Speech
Size: 411048 audio files > 1KB… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/swahili-words-speech-text-parallel.twi-words-speech-text-parallel-splitENGLISH_TWI_PARALLEL_TEXT
GhanaNLP Twi and English Parallel Data
Twi_to_English
• 1 MB • XLS
English_to_Twi
• 1 MB • XLS
The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.twi-trigrams-speech-text-parallel
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi - twi
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/twi-trigrams-speech-text-parallel.chichewa-trigrams-speech-text-parallel
Chichewa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 132549 parallel speech-text pairs for Chichewa, a language spoken primarily in Malawi. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Chichewa - ny
Task: Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-trigrams-speech-text-parallel.FANTE_ENGLISH_PARALLEL_TEXT
GhanaNLP Fante and English
Data
Fante_to_English
• 1 MB • XLS
The GhanaNLP Fante dataset contains sentence pairs in Fante and English,
designed to support translation models between these two languages.
Fante is a Ghanaian local language that lacks extensive digital resources,
making this dataset useful for language processing tools. The… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/FANTE_ENGLISH_PARALLEL_TEXT.TWI_ENGLISH_PARALLEL_TEXT
GhanaNLP Twi and English Parallel Data
Twi_to_English
• 1 MB • XLS
English_to_Twi
• 1 MB • XLS
The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/TWI_ENGLISH_PARALLEL_TEXT.EWE_ENGLISH_PARALLEL_TEXT
GhanaNLP Ewe and English
Data
Ewe_to_English
• 1 MB • XLS
The GhanaNLP Ewe dataset contains sentence pairs in Ewe and English,
designed to support translation models between these two languages.
Ewe is a Ghanaian local language that lacks extensive digital resources,
making this dataset useful for language processing tools. The sentence… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/EWE_ENGLISH_PARALLEL_TEXT.GA_ENGLISH_PARALLEL_TEXT
GhanaNLP Ga and English
Data
Ga_to_English
• 1 MB • XLS
The GhanaNLP Ga dataset contains sentence pairs in Ga and English,
designed to support translation models between these two languages.
Ga is a Ghanaian local language that lacks extensive digital resources,
making this dataset useful for language processing tools. The sentence… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/GA_ENGLISH_PARALLEL_TEXT.KUSAAL_ENGLISH_PARALLEL_TEXT
GhanaNLP Kusaal and English
Data
Kusaal_to_English
• 1 MB • XLS
The GhanaNLP Kusaal dataset contains sentence pairs in Kusaal and English,
designed to support translation models between these two languages.
Kusaal is a Ghanaian local language that lacks extensive digital resources,
making this dataset useful for language processing tools. The… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/KUSAAL_ENGLISH_PARALLEL_TEXT.Quran-Classical-Arabic-English-Parallel-texts
Introduction
This dataset presents a collection of parallel texts of the Holy Quran in Arabic (Imla'ei & Uthmanic scripts) alongside 17 different English translations.
Contents
The dataset includes the Holy Quran in Classical Arabic with the following English translations:
"al-Qur’ân: A Contemporary Translation" by Ahmed Ali
"Kanz-ul-Iman" by Ahmed Raza Khan
"The Koran Interpreted" by Arthur John Arberry
"The Message of The Qur'an" by Muhammad Asad
"Qur'an English… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Quran-Classical-Arabic-English-Parallel-texts.tr-hi-parallel-text
TR↔HI Parallel Text
65,662 aligned text triples — English pivot plus Turkish and Hindi
(en_text / tr_text / hi_text), each tagged with its source.
This is the text layer the speech corpora were synthesised from: these
sentences were sent to TTS to produce
tr-hi-parallel-speech-v2,
which was then Mimi-encoded into
tr-hi-mimi-encoded.
Text-only, ~10 MB, no audio. Sources include FLORES, OPUS-100, and
machine-translated conversational data — check source per row, since the
licence… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-text.Thaqalayn-Classical-Arabic-English-Parallel-texts
Introduction
This dataset represents a comprehensive collection of parallel Arabic-English texts from the Thaqalayn Hadith Library, a premier source for exploring the classical hadith tradition of the Imāmī Shia Muslim school of thought. The library focuses on making primary historical sources accessible, serving as a bridge between past wisdom and contemporary study. The dataset features translations of significant classical Imāmī hadith texts, allowing for a deep dive into the… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Thaqalayn-Classical-Arabic-English-Parallel-texts.texts_parallel_corpus_europarl_english_spanish
Dataset Card for Dataset Name
A massive parallel corpus of English-Spanish pairs. It hasn't a specified license, but there doesn't seem to be any copyrighted material in the corpus. I have personally merged rows using a pseudo-random algorithm making the dataset useful in training LLMs, reducing the risk of overfitting.
Dataset Details
Dataset Description
Curated by: Philipp Koehn
Funded by [optional]: In part funded by the European Commission (7th… See the full description on the dataset page: https://huggingface.co/datasets/Thermostatic/texts_parallel_corpus_europarl_english_spanish.parallel-complexity-med-texttwi-trigrams-speech-text-parallel-cleanparallel-image-text-dataset-builder
parallel-image-text-dataset-builder (sample)
A small representative sample from the
parallel-image-text-dataset-builder
pipeline: it ingests image-text pairs, removes near-duplicates with
perceptual-hash (dhash) LSH-style bucketing, filters weak pairs by CLIP
image-text similarity, and writes fixed-size WebDataset-style tar shards.
Contents
shard-00002.tar - one WebDataset-style shard (536 samples). Each sample is
two members sharing a key: {key}.jpg (image) and… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/parallel-image-text-dataset-builder.dhivehi-legal-text-parallel
Dhivehi-English Legal Parallel Corpus
Dataset Description
A high-quality parallel corpus of 56,556 Dhivehi-English sentence pairs extracted from 200 Maldivian legal documents. This dataset is deduplicated and cleaned for machine translation and bilingual model training.
Dataset Summary
Languages: Dhivehi (dv) ↔ English (en)
Total Pairs: 56,556
Source Laws: 200
Duplicates Removed: 31,235
Average Dhivehi Length: 173.6 characters
Average English… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-legal-text-parallel.swahili-parallel-text-extension
Swahili–English Extension of Kidaw’ida, Kalenjin and Dholuo Parallel Corpus
Dataset Description
This dataset is an extension of the parallel corpora introduced in Building Low-Resource African Language Corpora: A Case Study of Kidaw’ida, Kalenjin and Dholuo by:
Ambrose N. Ambogho
Quin Awuor
Andrew Kipkebut
Lilian Wanzare
Vivian Oloo
The original dataset provides Swahili–[Kidaw’ida/Kalenjin/Dholuo] parallel sentence pairs. English was not included in the original corpus.… See the full description on the dataset page: https://huggingface.co/datasets/tonative/swahili-parallel-text-extension.Chinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat
Chinese-English Parallel Translation Corpus (Chinese Source Text & English Translation)
A Chinese–English parallel corpus resource for translation and cross-lingual alignment applications, providing one-to-one bilingual text pairs: Chinese source texts aligned with their corresponding English translations. The data covers common writing styles and domains, making it suitable for parallel alignment, translation modeling, and cross-lingual representation learning.
It supports… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Chinese-English-Parallel-Translation-Corpus-Chinese-Source-Text-English-Translat.
