datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.paradetox
ParaDetox: Text Detoxification with Parallel Data (English)
This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.Crop-Recommendation-Parameters
🌱 Crop Recommendation Dataset
A machine learning dataset for crop recommendation based on soil properties and environmental conditions. The dataset contains measurements of essential soil nutrients and climatic parameters, along with the crop label that is suitable for those conditions.
This dataset can be used for machine learning classification, agricultural analytics, decision-support systems, and smart farming applications.
📌 Dataset Overview
Property… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/Crop-Recommendation-Parameters.CDial-BiasOfficial release of CDial-Bias dataset.
Notation:
Before downloading the dataset, please be aware that: The CDial-Bias Dataset is released for research purpose only and other usages require further permission. Please ensure the usage contributes to improving the safety and fairness of AI technologies. No malicious usage is allowed.
Paper:
https://aclanthology.org/2022.findings-emnlp.262/
Github Repo:
https://github.com/para-zhou/CDial-Bias
Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/para-zhou/CDial-Bias.paranoia
Paranoia
Naturalistic fMRI dataset: 22 subjects listened to a three-part ambiguous
social narrative (~22 minutes total) designed to elicit varying levels of
paranoid interpretation. TR = 1.0 s. 3-second fixation before each run.
This repo mirrors the fmriprep-preprocessed dataset originally distributed via
DataLad at https://gin.g-node.org/ljchang/Paranoia. fmriprep version
1.2.6-1.
Layout
derivatives/fmriprep/sub-tbXXXX/
anat/ func/ figures/
participants.tsv… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/paranoia.ARPA-Armenian-Paraphrase-Corpus
Dataset Description
We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language.
Dataset Summary
The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.chatgpt-paraphrasesThis is a dataset of paraphrases created by ChatGPT.
Model based on this dataset is avaible: model
We used this prompt to generate paraphrases
Generate 5 similar paraphrases for this question, show it like a numbered list without commentaries: {text}
This dataset is based on the Quora paraphrase question, texts from the SQUAD 2.0 and the CNN news dataset.
We generated 5 paraphrases for each sample, totally this dataset has about 420k data rows. You can make 30 rows from a row from… See the full description on the dataset page: https://huggingface.co/datasets/humarin/chatgpt-paraphrases.scipar_parallel_docs
SciPar Parallel Documents
Dataset Description
This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts.
In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories.
This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences.
To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.ru_paradetox
ParaDetox: Text Detoxification with Parallel Data (Russian)
This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit
[2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.Rasaif-Classical-Arabic-English-Parallel-texts
Introduction
This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period.
Content Details
Contained within this dataset are English translations of the following texts, sourced from the Rasaif website:
A Muslim Manual of War
Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.ParaSCIReformatted version of the ParaSCI dataset from ParaSCI: A Large Scientific Paraphrase Dataset for Longer Paraphrase Generation. Data retrieved from dqxiu/ParaSCI.
english_karakalpak_parallel_corpus_v5
English-Karakalpak Parallel Corpus
This dataset contains parallel sentences in English and Karakalpak language.
It is created to support AI development for the Karakalpak language.
Dataset Description
English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.NMT_Rwandan-Gazette_parallel_data_en_kin
Dataset Details
Dataset Description
This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix
Curated by: Digital Umuganda
Language(s) (NLP): Kinyarwanda and English
License: cc-by-4.0
Dataset Sources [optional]
The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.paraphrases_pawsENGLISH_TWI_PARALLEL_TEXT
GhanaNLP Twi and English Parallel Data
Twi_to_English
• 1 MB • XLS
English_to_Twi
• 1 MB • XLS
The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.african-language-parallel-corpus
African Language Parallel Corpus
Human-created, human-validated parallel sentence pairs for three African languages,
released openly by Okwu. Version 1.0.
Dataset summary
A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and
Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own
language-learning curriculum — content authored and reviewed by native-speaker educators —
supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.Brihat_Parashara_Hora_ShastraEnglish-Persian-Parallel-Dataset
English-Persian Parallel Dataset
This repository provides access to a high-quality parallel dataset for English-to-Persian translation. The dataset has been curated for research purposes and is suitable for training and evaluating Neural Machine Translation (NMT) models.
Download Link
You can download the dataset using the following link:
Download English-Persian Parallel Dataset
Description
The dataset contains aligned sentence pairs in English and Persian… See the full description on the dataset page: https://huggingface.co/datasets/shenasa/English-Persian-Parallel-Dataset.korean-parallel-corporaParallelXNLIvarplasma-parallel-dbd-air
Non-thermal Plasma Parallel DBD Air Dataset
Overview
This dataset contains experimental time-series measurements from a parallel Dielectric Barrier Discharge (DBD) plasma system in air at NTP.
The dataset was collected using a digital oscilloscope and includes current-voltage waveforms measurements for plasma discharge characterization.
Data Acquisition
The experiments were conducted in the Physics Laboratory, Department of Physics, Kathmandu… See the full description on the dataset page: https://huggingface.co/datasets/Funghang/plasma-parallel-dbd-air.paraphrases_mrpcflores-parallelmachine-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Machine Paraphrase Corpus (MPC) consists of ~200k examples of original, and paraphrases using two online paraphrasing tools.
It uses two paraphrasing tools (SpinnerChief, SpinBot) on three source texts (Wikipedia, arXiv, student theses).
The examples are not aligned, i.e., we sample different paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset.paradehttps://github.com/heyunh2015/PARADE_dataset
@inproceedings{he-etal-2020-parade,
title = "{PARADE}: {A} {N}ew {D}ataset for {P}araphrase {I}dentification {R}equiring {C}omputer {S}cience {D}omain {K}nowledge",
author = "He, Yun and
Wang, Zhuoer and
Zhang, Yin and
Huang, Ruihong and
Caverlee, James",
booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)",
month = nov,
year = "2020",
address… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/parade.Paratope
Nanobody Paratope Prediction Dataset
Dataset Overview
This dataset helps predict which amino acids in nanobody sequences directly bind to antigens (these binding sites are called "paratopes"). Knowing the paratope residues is important for understanding how antibodies interact with their targets and for designing better therapeutic nanobodies.
Data Collection
The data is based on solved 3D structures of nanobody-antigen complexes. These structures come from the… See the full description on the dataset page: https://huggingface.co/datasets/ZYMScott/Paratope.secondary-screen-dose-response-curve-parameters
Citation
DepMap, Broad; Corsello, Steven; Kocak, Mustafa; Golub, Todd (2019). PRISM Repurposing 19Q4 Dataset. figshare. Dataset.
Current dataset: https://doi.org/10.6084/m9.figshare.9393293.v4
General guidance: https://doi.org/10.1101/730119
Dataset specific README: prism_repurposing_secondary
Dataset Fields
broad_id: ID used to identify drug-batch combinations
name: Name of the drug
depmap_id: ID to identify a cell line
ccle_name: ID to identify a cell line… See the full description on the dataset page: https://huggingface.co/datasets/donb-hf/secondary-screen-dose-response-curve-parameters.autoencoder-paraphrase-dataset
Dataset Card for Machine Paraphrase Dataset (MPC)
Dataset Summary
The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models.
It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses).
The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions.
How to use it
You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.PARADISEkashmiri_English_parallel_corpus_49K
license: apache-2.0
task_categories:
translation
language:
ks
Usage Terms for this Dataset
Purpose of UseThis dataset is made available for the purpose of training machine learning models, academic research, and other non-commercial uses and its applications.
Citation RequirementIf you use this dataset for research, training models, or any other purpose, you must provide proper attribution by citing the following:
@misc {haq_nawaz_malik_2024,
author = { {HAQ NAWAZ… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/kashmiri_English_parallel_corpus_49K.
