CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes21k downloads1y agoHugging Face02Columbia-NLP /PUPAThis dataset contains the data presented in the paper PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles. Code: https://github.com/siyan-sylvia-li/PAPILLON texttext-generationn<1K3 likes6.6k downloads1y agoHugging Face03MaartenGr /arxiv_nlp arXiv Abstracts Abstracts for the cs.CL category of ArXiv between 1991 and 2024. This dataset was created as an instructional tool for the Clustering and Topic Modeling chapter in the upcoming "Hands-On Large Language Models" book. The original dataset was retrieved here. This subset will be updated towards the release of the book to make sure it captures relatively recent articles in the domain. text10K<n<100K12 likes1.1k downloads3y agoHugging Face04DAMO-NLP-SG /MultiJail Multilingual Jailbreak Challenges in Large Language Models This repo contains the data for our paper "Multilingual Jailbreak Challenges in Large Language Models". [Github repo] Annotation Statistics We collected a total of 315 English unsafe prompts and annotated them into nine non-English languages. The languages were categorized based on resource availability, as shown below: High-resource languages: Chinese (zh), Italian (it), Vietnamese (vi) Medium-resource languages:… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/MultiJail.textn<1K12 likes1.1k downloads3y agoHugging Face05s-nlp /paradetox ParaDetox: Text Detoxification with Parallel Data (English) This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.imagetext-generation10K<n<100K9 likes842 downloads1y agoHugging Face06SALT-NLP /ImplicitHate Implicit Hate Speech Latent Hatred: A Benchmark for Understanding Implicit Hate Speech [Read the Paper] | [Take a Survey to Access the Data] | [Download the Data] Why Implicit Hate? It is important to consider the subtle tricks that many extremists use to mask their threats and abuse. These more implicit forms of hate speech may easily go undetected by keyword detection systems, and even the most advanced architectures can fail if they have not been trained on… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/ImplicitHate.text1K<n<10K9 likes467 downloads4y agoHugging Face07AUEB-NLP /lar-echr Dataset Card for LAR-ECHR Dataset Details Dataset Description Curated by: Odysseas S. Chlapanis Funded by: Archimedes Research Unit Language (NLP): English License: CC BY-NC-SA (Creative Commons / Attribution-NonCommercial-ShareAlike) Read more: https://creativecommons.org/licenses/by-nc-sa/4.0/ Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Uses… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/lar-echr.textquestion-answeringn<1K3 likes464 downloads1y agoHugging Face08sixuexing /FAERS-NLP FAERS-NLP Version: 1.0Author: sixuexing GitHub: FAERS-NLP Repository Dataset Summary FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction. Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks. Dataset Structure Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/sixuexing/FAERS-NLP.tabular1M<n<10M1 likes431 downloads1y agoHugging Face09nlpatunt /D_persuade_2 Persuade_2 The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and Understanding Argumentative and Discourse Elements) contains over 25,000 argumentative essays written by 6th–12th grade students in the United States, covering 15 distinct prompts across two writing tasks: independent and source-based writing. The corpus also provides detailed individual and demographic information for each writer. This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.tabular10K<n<100K0 likes370 downloads6mo agoHugging Face10tum-nlp /neural-news-benchmark AI-generated News Detection Benchmark neural-news is a benchmark dataset designed for human/AI news authorship classification in English, Turkish, Hungarian, and Persian. Presented in Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian @ NLP for Positive Impact Workshop @ EMNLP2024. Dataset Details The dataset includes equal parts human-written and AI-generated news articles, raw and pre-processed. Curated… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/neural-news-benchmark.texttext-classification10K<n<100K4 likes334 downloads2y agoHugging Face11nlpatunt /D_ASAP-AES D_ASAP-AES This is the train, test, and validation split of the ASAP Automated Essay Scoring dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation. For the original dataset with labels, see below. Original Dataset 🔗 ASAP-AES on Kaggle Citation If you use this dataset, please cite the original: @misc{asap_aes, title={ASAP Automated Essay Scoring}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-AES.tabular10K<n<100K0 likes307 downloads6mo agoHugging Face12GroNLP /ik-nlp-22_winemagtabular10K<n<100K6 likes272 downloads5y agoHugging Face13Alibaba-NLP /EcomBench EcomBench: Where Intelligent Agents Conquer Commerce Realms 🚀 Benchmark Overview EcomBench is a domain-specific, real-world evaluation framework designed to rigorously assess the capabilities of AI agents in delivering practical support for the complex, ever-evolving demands of e-commerce. We believe that truly capable AI agents will fundamentally transform how we interact with commerce. E-commerce represents one of the world's most significant economic sectors, with… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/EcomBench.textn<1K7 likes265 downloads10mo agoHugging Face14sinhala-nlp /SOLD SOLD - A Benchmark for Sinhala Offensive Language Identification In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SOLD.texttext-classification10K<n<100K2 likes263 downloads3y agoHugging Face15uoe-nlp /extrinsic_mt_evaltabular10K<n<100K0 likes251 downloads3y agoHugging Face16SALT-NLP /CultureBanktext10K<n<100K20 likes225 downloads2y agoHugging Face17SanaeLaRose /FAERS-NLP FAERS-NLP Version: 1.0Author: sixuexing GitHub: FAERS-NLP Repository Dataset Summary FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction. Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks. Dataset Structure Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/SanaeLaRose/FAERS-NLP.tabular1M<n<10M0 likes209 downloads8mo agoHugging Face18OpenLab-NLP /tiny-singleturn-chat-kotextquestion-answering10K<n<100K0 likes207 downloads10mo agoHugging Face19s-nlp /ru_paradetox ParaDetox: Text Detoxification with Parallel Data (Russian) This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts. 📰 Updates [2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit [2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.imagetext-generation10K<n<100K4 likes203 downloads1y agoHugging Face20NLPLabNTUST /Merged-CWA CWA Benchmark: A Seismic Dataset from Taiwan for Seismic Research Dataset Description This dataset includes a larger number of seismic events, especially high-magnitude. A comprehensive set of events collected by the Central Weather Bureau in Taiwan. The CWA benchmark features over 40 attributes and ∼500,000 seismograms, providing valuable data labels for various seismology-related tasks. In the future, we will keep updating the dataset to ensure its relevance and… See the full description on the dataset page: https://huggingface.co/datasets/NLPLabNTUST/Merged-CWA.tabular1K<n<10K0 likes196 downloads2y agoHugging Face21Finnish-NLP /mc4_fi_cleaned Dataset Card for mC4 Finnish Cleaned Dataset Summary mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split. Supported Tasks and Leaderboards mC4 Finnish is mainly intended to pretrain Finnish language models and word representations. Languages Finnish Dataset Structure Data Instances [Needs More Information] Data Fields The data have several fields: url: url of the source as a string text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.texttext-generation10M<n<100M4 likes192 downloads4y agoHugging Face22nlpscu /Beyond-Flesch Beyond-Flesch: ScienceQA Difficulty Classification with Static and Prompt-Based Metrics A preprocessed subset of ScienceQA for K-12 educational text difficulty classification, along with the static and LLM-derived prompt-based features we use to reproduce Rooein et al. (2024) — Beyond Flesch-Kincaid. This dataset accompanies our class research project (Option 1: reproducing a paper whose original code was not released). What's here File Rows Description… See the full description on the dataset page: https://huggingface.co/datasets/nlpscu/Beyond-Flesch.tabular10K<n<100K0 likes192 downloads4mo agoHugging Face23McGill-NLP /statcan-dialogue-dataset-retrieval Statcan Dialogue Dataset (Processed for Retrieval Tasks) This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately. Quickstart from datasets import load_dataset repo = 'McGill-NLP/statcan-dialogue-dataset-retrieval' # load english queries, training split queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.textquestion-answering10K<n<100K1 likes184 downloads2y agoHugging Face24hamedhf /nlp_twitter_analysistexttext-classification1K<n<10K1 likes182 downloads3y agoHugging Face25Tamazight-NLP /AmaWar Amawal Warayni - ⴰⵎⴰⵡⴰⵍ ⴰⵙⵏⵎⴰⵍⴰⵢ ⵏ ⵉⵏⵓⵎⴰⴽ ⵏ ⵡⴰⵙⵙⴰⵖⵏ ⴷ ⵉⵎⵢⴰⴳⵏ ⵏ ⵜⵎⴰⵣⵉⵖⵜ ⵏ ⴰⵢⵜ ⵡⴰⵔⴰⵢⵏ Bitext scraped from the online AmaWar dictionary of the Tamazight dialect of Ait Warain spoken in northeastern Morocco. Contains sentences, stories, and poems in Tamazight written in the Neo-Tifinagh script along with their translations into Modern Standard Arabic. The dataset is split into the following subsets: examples: Example parallel sentences taken from dictionary entries. idioms: Idiomatic… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/AmaWar.texttranslation1K<n<10K2 likes169 downloads4mo agoHugging Face26recogna-nlp /FakeRecogna FakeRecogna FakeRecogna is a dataset comprised of real and fake news. The real news is not directly linked to fake news and vice-versa, which could lead to a biased classification. The news collection was performed by crawlers developed for mining pages of well-known and of great national importance agency news. The web crawlers were developed based on each analyzed webpage, where the extracted information is first separated into categories and then grouped by dates. The plurality… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/FakeRecogna.texttext-classification10K<n<100K3 likes166 downloads3y agoHugging Face27s-nlp /lc_quad2 Dataset Card for LC-QuAD 2.0 with answers textquestion-answering10K<n<100K1 likes161 downloads3y agoHugging Face28Ghana-NLP /ENGLISH_TWI_PARALLEL_TEXT GhanaNLP Twi and English Parallel Data Twi_to_English • 1 MB • XLS English_to_Twi • 1 MB • XLS The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.text1K<n<10K3 likes143 downloads10mo agoHugging Face29recogna-nlp /fakerecogna2-abstrativa FakeRecogna 2.0 - Abstractive FakeRecogna 2.0 presents the extension for the FakeRecogna dataset in the context of fake news detection. FakeRecogna includes real and fake news texts collected from online media and ten fact-checking sources in Brazil. An important aspect is the lack of relation between the real and fake news samples, i.e., they are not mutually related to each other to avoid intrinsic bias in the data. The Dataset The fake news collection was performed on… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/fakerecogna2-abstrativa.tabulartext-classification10K<n<100K2 likes140 downloads1y agoHugging Face30yassiracharki /Yahoo_Answers_10_categories_for_NLP Dataset Card for Dataset Name The Yahoo! Answers topic classification dataset is constructed using 10 largest main categories. Each class contains 140,000 training samples and 6,000 testing samples. Therefore, the total number of training samples is 1,400,000 and testing samples 60,000 in this dataset. From all the answers and other meta-information, we only used the best answer content and the main category information. Dataset Description The file classes.txt contains a… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yahoo_Answers_10_categories_for_NLP.texttext-classification1M<n<10M3 likes137 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.