CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mrlbenchmarks /global-piqa-parallel Global PIQA Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems. In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.imagequestion-answering10K<n<100K10 likes3.4k downloads4mo agoHugging Face02s-nlp /paradetox ParaDetox: Text Detoxification with Parallel Data (English) This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.imagetext-generation10K<n<100K9 likes901 downloads1y agoHugging Face03Samarth-27 /Crop-Recommendation-Parameters 🌱 Crop Recommendation Dataset A machine learning dataset for crop recommendation based on soil properties and environmental conditions. The dataset contains measurements of essential soil nutrients and climatic parameters, along with the crop label that is suitable for those conditions. This dataset can be used for machine learning classification, agricultural analytics, decision-support systems, and smart farming applications. 📌 Dataset Overview Property… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/Crop-Recommendation-Parameters.tabular1K<n<10K0 likes532 downloads2mo agoHugging Face04para-zhou /CDial-BiasOfficial release of CDial-Bias dataset. Notation: Before downloading the dataset, please be aware that: The CDial-Bias Dataset is released for research purpose only and other usages require further permission. Please ensure the usage contributes to improving the safety and fairness of AI technologies. No malicious usage is allowed. Paper: https://aclanthology.org/2022.findings-emnlp.262/ Github Repo: https://github.com/para-zhou/CDial-Bias Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/para-zhou/CDial-Bias.tabulartext-classification10K<n<100K8 likes465 downloads2y agoHugging Face05dartbrains /paranoia Paranoia Naturalistic fMRI dataset: 22 subjects listened to a three-part ambiguous social narrative (~22 minutes total) designed to elicit varying levels of paranoid interpretation. TR = 1.0 s. 3-second fixation before each run. This repo mirrors the fmriprep-preprocessed dataset originally distributed via DataLad at https://gin.g-node.org/ljchang/Paranoia. fmriprep version 1.2.6-1. Layout derivatives/fmriprep/sub-tbXXXX/ anat/ func/ figures/ participants.tsv… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/paranoia.imagefeature-extractionn<1K0 likes283 downloads3mo agoHugging Face06Karavet /ARPA-Armenian-Paraphrase-Corpus Dataset Description We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language. Dataset Summary The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.text1K<n<10K3 likes282 downloads4y agoHugging Face07humarin /chatgpt-paraphrasesThis is a dataset of paraphrases created by ChatGPT. Model based on this dataset is avaible: model We used this prompt to generate paraphrases Generate 5 similar paraphrases for this question, show it like a numbered list without commentaries: {text} This dataset is based on the Quora paraphrase question, texts from the SQUAD 2.0 and the CNN news dataset. We generated 5 paraphrases for each sample, totally this dataset has about 420k data rows. You can make 30 rows from a row from… See the full description on the dataset page: https://huggingface.co/datasets/humarin/chatgpt-paraphrases.text100K<n<1M61 likes270 downloads3y agoHugging Face08ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K2 likes223 downloads3y agoHugging Face09s-nlp /ru_paradetox ParaDetox: Text Detoxification with Parallel Data (Russian) This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit [2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.imagetext-generation10K<n<100K4 likes209 downloads1y agoHugging Face10ImruQays /Rasaif-Classical-Arabic-English-Parallel-texts Introduction This dataset represents a curated collection of parallel Arabic-English texts, featuring the translations of 24 historically and culturally significant books. These texts provide a portal to the intellectual and literary heritage of the Arabic-speaking world during its classical period. Content Details Contained within this dataset are English translations of the following texts, sourced from the Rasaif website: A Muslim Manual of War Al-Hanin Ila'l-Awtan… See the full description on the dataset page: https://huggingface.co/datasets/ImruQays/Rasaif-Classical-Arabic-English-Parallel-texts.texttranslation10K<n<100K8 likes206 downloads3y agoHugging Face11HHousen /ParaSCIReformatted version of the ParaSCI dataset from ParaSCI: A Large Scientific Paraphrase Dataset for Longer Paraphrase Generation. Data retrieved from dqxiu/ParaSCI. text100K<n<1M2 likes173 downloads5y agoHugging Face12bekan /english_karakalpak_parallel_corpus_v5 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa). Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.texttranslation10K<n<100K4 likes166 downloads13d agoHugging Face13DigitalUmuganda /NMT_Rwandan-Gazette_parallel_data_en_kin Dataset Details Dataset Description This is a curated parallel dataset from the Official Gazette of the Republic of Rwanda. It has been curated to extract corresponding English and Kinyarwanda text and in the future we shall add French to the mix Curated by: Digital Umuganda Language(s) (NLP): Kinyarwanda and English License: cc-by-4.0 Dataset Sources [optional] The dataset original content was retrieved from the Rwandan ministry of Justice website… See the full description on the dataset page: https://huggingface.co/datasets/DigitalUmuganda/NMT_Rwandan-Gazette_parallel_data_en_kin.texttranslation100K<n<1M3 likes147 downloads3y agoHugging Face14allegrolab /paraphrases_pawstext1K<n<10K0 likes143 downloads1y agoHugging Face15Ghana-NLP /ENGLISH_TWI_PARALLEL_TEXT GhanaNLP Twi and English Parallel Data Twi_to_English • 1 MB • XLS English_to_Twi • 1 MB • XLS The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.text1K<n<10K3 likes143 downloads10mo agoHugging Face16Okwu /african-language-parallel-corpus African Language Parallel Corpus Human-created, human-validated parallel sentence pairs for three African languages, released openly by Okwu. Version 1.0. Dataset summary A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own language-learning curriculum — content authored and reviewed by native-speaker educators — supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.texttranslation10K<n<100K0 likes143 downloads4d agoHugging Face17JDhruv14 /Brihat_Parashara_Hora_Shastratabular1K<n<10K0 likes137 downloads9mo agoHugging Face18shenasa /English-Persian-Parallel-Dataset English-Persian Parallel Dataset This repository provides access to a high-quality parallel dataset for English-to-Persian translation. The dataset has been curated for research purposes and is suitable for training and evaluating Neural Machine Translation (NMT) models. Download Link You can download the dataset using the following link: Download English-Persian Parallel Dataset Description The dataset contains aligned sentence pairs in English and Persian… See the full description on the dataset page: https://huggingface.co/datasets/shenasa/English-Persian-Parallel-Dataset.text1M<n<10M11 likes135 downloads1y agoHugging Face19Moo /korean-parallel-corporatexttranslation10K<n<100K21 likes131 downloads4y agoHugging Face20jaio98 /ParallelXNLIvartext10K<n<100K0 likes130 downloads9d agoHugging Face21Funghang /plasma-parallel-dbd-air Non-thermal Plasma Parallel DBD Air Dataset Overview This dataset contains experimental time-series measurements from a parallel Dielectric Barrier Discharge (DBD) plasma system in air at NTP. The dataset was collected using a digital oscilloscope and includes current-voltage waveforms measurements for plasma discharge characterization. Data Acquisition The experiments were conducted in the Physics Laboratory, Department of Physics, Kathmandu… See the full description on the dataset page: https://huggingface.co/datasets/Funghang/plasma-parallel-dbd-air.textfeature-extraction10K<n<100K3 likes128 downloads4mo agoHugging Face22allegrolab /paraphrases_mrpctext1K<n<10K0 likes127 downloads1y agoHugging Face23israel /flores-paralleltabular1K<n<10K0 likes113 downloads2y agoHugging Face24jpwahle /machine-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Machine Paraphrase Corpus (MPC) consists of ~200k examples of original, and paraphrases using two online paraphrasing tools. It uses two paraphrasing tools (SpinnerChief, SpinBot) on three source texts (Wikipedia, arXiv, student theses). The examples are not aligned, i.e., we sample different paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset.texttext-classification100K<n<1M7 likes111 downloads1y agoHugging Face25tasksource /paradehttps://github.com/heyunh2015/PARADE_dataset @inproceedings{he-etal-2020-parade, title = "{PARADE}: {A} {N}ew {D}ataset for {P}araphrase {I}dentification {R}equiring {C}omputer {S}cience {D}omain {K}nowledge", author = "He, Yun and Wang, Zhuoer and Zhang, Yin and Huang, Ruihong and Caverlee, James", booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)", month = nov, year = "2020", address… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/parade.tabularsentence-similarity10K<n<100K0 likes102 downloads3y agoHugging Face26ZYMScott /Paratope Nanobody Paratope Prediction Dataset Dataset Overview This dataset helps predict which amino acids in nanobody sequences directly bind to antigens (these binding sites are called "paratopes"). Knowing the paratope residues is important for understanding how antibodies interact with their targets and for designing better therapeutic nanobodies. Data Collection The data is based on solved 3D structures of nanobody-antigen complexes. These structures come from the… See the full description on the dataset page: https://huggingface.co/datasets/ZYMScott/Paratope.text1K<n<10K0 likes95 downloads1y agoHugging Face27donb-hf /secondary-screen-dose-response-curve-parameters Citation DepMap, Broad; Corsello, Steven; Kocak, Mustafa; Golub, Todd (2019). PRISM Repurposing 19Q4 Dataset. figshare. Dataset. Current dataset: https://doi.org/10.6084/m9.figshare.9393293.v4 General guidance: https://doi.org/10.1101/730119 Dataset specific README: prism_repurposing_secondary Dataset Fields broad_id: ID used to identify drug-batch combinations name: Name of the drug depmap_id: ID to identify a cell line ccle_name: ID to identify a cell line… See the full description on the dataset page: https://huggingface.co/datasets/donb-hf/secondary-screen-dose-response-curve-parameters.tabular100K<n<1M2 likes89 downloads2y agoHugging Face28jpwahle /autoencoder-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models. It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses). The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.tabulartext-classification1M<n<10M2 likes87 downloads1y agoHugging Face29GGLab /PARADISEtabular100K<n<1M2 likes83 downloads3y agoHugging Face30Omarrran /kashmiri_English_parallel_corpus_49Kgated license: apache-2.0 task_categories: translation language: ks Usage Terms for this Dataset Purpose of UseThis dataset is made available for the purpose of training machine learning models, academic research, and other non-commercial uses and its applications. Citation RequirementIf you use this dataset for research, training models, or any other purpose, you must provide proper attribution by citing the following: @misc {haq_nawaz_malik_2024, author = { {HAQ NAWAZ… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/kashmiri_English_parallel_corpus_49K.text10K<n<100K2 likes82 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.