datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xone-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..)
A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/xone-repository-parallel-en-id-corpus.human-ai-parallel-corpus
Human-AI Parallel English Corpus (HAP-E) 🙃
Purpose
The HAP-E corpus is designed for comparisions of the writing produced by humans and the writing produced by large language models (LLMs).
The corpus was created by seeding an LLM with an approximately 500-word chunk of human-authored text and then prompting the model to produce an additional 500 words.
Thus, a second 500-word chunk of human-authored text (what actually comes next in the original text) can be compared to… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus.UFAL_Parallel_Corpus_of_North_Levantine_1.0
[!NOTE]
Dataset origin: https://zenodo.org/records/4012218
UFAL Parallel Corpus of North Levantine 1.0
March 10, 2023
Authors
Shadi Saleh <saleh@ufal.mff.cuni.cz>
Hashem Sellat <sellat@ufal.mff.cuni.cz>
Mateusz Krubiński <krubinski@ufal.mff.cuni.cz>
Adam Posppíšil <adam.pospisil@ff.cuni.cz>
Petr Zemánek <petr.zemanek@ff.cuni.cz>
Pavel Pecina <pecina@ufal.mff.cuni.cz>
Overview
This is the first release of the UFAL Parallel Corpus of North Levantine… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/UFAL_Parallel_Corpus_of_North_Levantine_1.0.khakas-russian-parallel-corpus
Khakas-Russian Parallel Corpus
The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and
machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing
high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people.
Dataset Overlap:
The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.english-danish-parallel-corpus
DanishMedicinesAgencyBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
A Bilingual English-Danish parallel corpus from The Danish Medicines Agency.
Task category
t2t
Domains
Medical, Written
Reference
https://sprogteknologi.dk/dataset/bilingual-english-danish-parallel-corpus-from-the-danish-medicines-agency
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/english-danish-parallel-corpus.english_karakalpak_parallel_corpus_v5
English-Karakalpak Parallel Corpus
This dataset contains parallel sentences in English and Karakalpak language.
It is created to support AI development for the Karakalpak language.
Dataset Description
English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences.
Egyptian-Arabic-English-Parallel-Corpus
Egyptian Arabic-English Parallel Corpus
Author: Mohamed Abdalkader · LinkedIn · GitHub
A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation.
Dataset Structure
egyptian-arabic-english-parallel-corpus/
├── SFT/
│ ├── Train/
│ │ ├── topics/ # 1,800 individual topic JSON files
│ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.human-ai-parallel-corpus-biber
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles
@misc{reinhart2024llmswritelikehumans,
title={Do LLMs write like humans? Variation in grammatical and rhetorical styles},
author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg},
year={2024},
eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-biber.Ekegusii-English-Kiswahili-Parallel-Corpus
Ekegusii - English - Kiswahili Parallel Master Corpus
This repository contains the official consolidated, cleaned, and deduplicated Multilingual Parallel Corpus for Ekegusii (guz), Kiswahili (sw), and English (en) machine translation research, with a focus on Public Service Announcements (PSAs) in Kenya.
Dataset Summary
Total Records: 149,849 unique parallel alignment concepts.
Languages: English, Kiswahili (Swahili), Ekegusii (Gusii).
Public Service… See the full description on the dataset page: https://huggingface.co/datasets/aykgeh/Ekegusii-English-Kiswahili-Parallel-Corpus.circassian-parallel-corpus
Circassian-Russian Parallel Corpus v1.0
This is a high-quality dataset containing over 330,000 parallel text pairs for machine translation between Russian and the Circassian language in its two literary dialects: East Circassian (Kabardian, kbd) and West Circassian (Adyghe, ady).
About Circassian
Circassian is an indigenous language of the Northwest Caucasus region. The language is notable for its complex phonological system (featuring 50+ consonants)… See the full description on the dataset page: https://huggingface.co/datasets/adiga-ai/circassian-parallel-corpus.human-ai-parallel-corpus-2
Human-AI Parallel English Corpus-2 (HAP-E-2) 🙃
Purpose
The HAP-E-2 corpus is an extension of the original HAP-E corpus, with the addition of updated models. is designed for comparisions of the writing produced by humans and the writing produced by large language models (LLMs).
The corpus was created by seeding an LLM with an approximately 500-word chunk of human-authored text and then prompting the model to produce an additional 500 words.
Thus, a second 500-word chunk… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-2.basque-parallel-corpus
Sources of the corpus used
Opus
Orai
coca-ai-parallel-corpus-biber
COCA-AI Parallel Corpus (Biber Parsed)
R users can import the data directly using r-polars:
library(polars)
df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet')
df <- df$to_data_frame()
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles@misc{reinhart2024llmswritelikehumans,
title={Do LLMs write like humans? Variation in grammatical… See the full description on the dataset page: https://huggingface.co/datasets/browndw/coca-ai-parallel-corpus-biber.Filtered-Japanese-English-Parallel-Corpusdef prompt(japanese, english):
system_prompt = cleandoc("""<s>[INST]Your role is to evaluate the accuracy of the provided Japanese to English translation.
- Translations with parts missing should be rejected.
- Incomplete translations should be rejected.
- Inaccurate translations should be rejected.
- Poor grammar should be rejected.
- Any kind of mistake should be rejected.
- Bad spelling should be rejected.
- Low quality english should be rejected.
- Low… See the full description on the dataset page: https://huggingface.co/datasets/Moleys/Filtered-Japanese-English-Parallel-Corpus.african-language-parallel-corpus
African Language Parallel Corpus
Human-created, human-validated parallel sentence pairs for three African languages,
released openly by Okwu. Version 1.0.
Dataset summary
A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and
Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own
language-learning curriculum — content authored and reviewed by native-speaker educators —
supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.tunisian-msa-parallel-corpus
Dataset Description
This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models.
The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.ES-AN_Parallel_Corpus
Dataset Card for ES-AN Parallel Corpus
Dataset Summary
The ES-AN Parallel Corpus is a Spanish-Aragonese dataset created to support the use of under-resourced languages from Spain,
such as Aragonese, in NLP tasks, specifically Machine Translation.
Supported Tasks and Leaderboards
The dataset can be used to train Bilingual Machine Translation models between Aragonese and Spanish in any direction, as well as Multilingual Machine Translation models.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ES-AN_Parallel_Corpus.human-ai-parallel-corpus-spacy
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles
@misc{reinhart2024llmswritelikehumans,
title={Do LLMs write like humans? Variation in grammatical and rhetorical styles},
author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg},
year={2024},
eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-spacy.cantonese-chinese-parallel-corpus
Dataset Summary
This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation.
The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation.
Languages
Cantonese (yue)
Simplified Chinese (zh)
Dataset Structure
Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page: https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.CA-ZH_Parallel_Corpus
Dataset Card for CA-ZH Parallel Corpus
Dataset Summary
The CA-ZH Parallel Corpus is a Catalan-Chinese dataset of parallel sentences created to
support Catalan in NLP tasks, specifically Machine Translation.
Supported Tasks and Leaderboards
The dataset can be used to train Bilingual Machine Translation models between Chinese and Catalan in any direction,
as well as Multilingual Machine Translation models.
Languages
The sentences included in the… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-ZH_Parallel_Corpus.human-ai-parallel-corpus-docuscope
COCA-AI Parallel Corpus (Biber Parsed)
Data were tagged with the en_docusco_spacy model.
R users can import the data directly using r-polars:
library(polars)
df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet')
df <- df$to_data_frame()
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles
@misc{reinhart2024llmswritelikehumans,
title={Do… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-docuscope.kashmiri_English_parallel_corpus_49K
license: apache-2.0
task_categories:
translation
language:
ks
Usage Terms for this Dataset
Purpose of UseThis dataset is made available for the purpose of training machine learning models, academic research, and other non-commercial uses and its applications.
Citation RequirementIf you use this dataset for research, training models, or any other purpose, you must provide proper attribution by citing the following:
@misc {haq_nawaz_malik_2024,
author = { {HAQ NAWAZ… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/kashmiri_English_parallel_corpus_49K.CA-DE_Parallel_Corpus
Dataset Card for CA-DE Parallel Corpus
Dataset Summary
The CA-DE Parallel Corpus is a Catalan-German dataset of parallel sentences created to support Catalan in NLP tasks, specifically
Machine Translation.
Supported Tasks and Leaderboards
The dataset can be used to train Bilingual Machine Translation models between German and Catalan in any direction,
as well as Multilingual Machine Translation models.
Languages
The sentences included in the… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-DE_Parallel_Corpus.Kashmiri-English-Parallel-Corpus
Kashmiri-English Parallel Corpus
This dataset is a parallel corpus of Kashmiri and English, consisting of 30,000 sentence pairs(also contains filtered already available corpus). The original dataset contains 270,000 sentence pairs, which are available upon request. This dataset can be utilized for various NLP tasks, including machine translation, alignment studies, and linguistic research.
Corpus Structure
The dataset is organized into several directories to differentiate… See the full description on the dataset page: https://huggingface.co/datasets/SMUQamar/Kashmiri-English-Parallel-Corpus.Catalan-Aranese_Parallel_Corpus
Dataset Card for Catalan-Aranese Parallel Corpus
Dataset Summary
A bilingual parallel corpus for the low-resource language pair Catalan-Aranese. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic Catalan translations generated from… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Catalan-Aranese_Parallel_Corpus.CA-GL_Parallel_Corpus
Dataset Card for CA-GL Parallel Corpus
Dataset Description
Dataset Summary
The CA-GL Parallel Corpus is a Catalan-Galician synthetic dataset parallel sentences created to
support the use of co-official languages from Spain, such as Catalan and Galician,
in NLP tasks, specifically Machine Translation.
Supported Tasks and Leaderboards
The dataset can be used to train Bilingual Machine Translation models between Galician and Catalan in any direction… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-GL_Parallel_Corpus.msmarco-corpus-en-id-parallel-sentences
Dataset Card for "msmarco-corpus-en-id-parallel-sentences"
More Information needed
ES-OC_Parallel_Corpus
Dataset Card for ES-OC Parallel Corpus
Dataset Summary
The ES-OC Parallel Corpus is a Spanish-Aranese dataset created to support the use of under-resourced languages from Spain,
such as Aranese, in NLP tasks, specifically Machine Translation.
Supported Tasks and Leaderboards
The dataset can be used to train Bilingual Machine Translation models between Aranese and Spanish in any direction, as well as Multilingual Machine Translation models.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ES-OC_Parallel_Corpus.azerbaijani-english-parallel-corpus
Azerbaijani English Parallel Corpus
This dataset contains 4,141,966 pairs of high-quality sentences translated from Azerbaijani to English. The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc.
License
CC-BY-4.0
Contact
For more information, questions, or issues, please contact LocalDoc at [v.resad.89@gmail.com].
