datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-reward-bench
AgentRewardBench
💾Code
📄Paper
🌐Website
🤗Dataset
💻Demo
🏆Leaderboard
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor
Loading dataset
You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.PUPAThis dataset contains the data presented in the paper PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles.
Code: https://github.com/siyan-sylvia-li/PAPILLON
arxiv_nlp
arXiv Abstracts
Abstracts for the cs.CL category of ArXiv between 1991 and 2024. This dataset was created as an instructional tool for the Clustering and Topic Modeling chapter in the upcoming
"Hands-On Large Language Models" book.
The original dataset was retrieved here.
This subset will be updated towards the release of the book to make sure it captures relatively recent articles in the domain.
MultiJail
Multilingual Jailbreak Challenges in Large Language Models
This repo contains the data for our paper "Multilingual Jailbreak Challenges in Large Language Models".
[Github repo]
Annotation Statistics
We collected a total of 315 English unsafe prompts and annotated them into nine non-English languages. The languages were categorized based on resource availability, as shown below:
High-resource languages: Chinese (zh), Italian (it), Vietnamese (vi)
Medium-resource languages:… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/MultiJail.paradetox
ParaDetox: Text Detoxification with Parallel Data (English)
This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.ImplicitHate
Implicit Hate Speech
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech
[Read the Paper] | [Take a Survey to Access the Data] | [Download the Data]
Why Implicit Hate?
It is important to consider the subtle tricks that many extremists use to mask their threats and abuse. These more implicit forms of hate speech may easily go undetected by keyword detection systems, and even the most advanced architectures can fail if they have not been trained on… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/ImplicitHate.lar-echr
Dataset Card for LAR-ECHR
Dataset Details
Dataset Description
Curated by: Odysseas S. Chlapanis
Funded by: Archimedes Research Unit
Language (NLP): English
License:
CC BY-NC-SA (Creative Commons / Attribution-NonCommercial-ShareAlike)
Read more: https://creativecommons.org/licenses/by-nc-sa/4.0/
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Uses… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/lar-echr.FAERS-NLP
FAERS-NLP
Version: 1.0Author: sixuexing
GitHub: FAERS-NLP Repository
Dataset Summary
FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction.
Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks.
Dataset Structure
Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/sixuexing/FAERS-NLP.D_persuade_2
Persuade_2
The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and
Understanding Argumentative and Discourse Elements) contains over 25,000
argumentative essays written by 6th–12th grade students in the United States,
covering 15 distinct prompts across two writing tasks: independent and
source-based writing. The corpus also provides detailed individual and
demographic information for each writer.
This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.neural-news-benchmark
AI-generated News Detection Benchmark
neural-news is a benchmark dataset designed for human/AI news authorship classification in English, Turkish, Hungarian, and Persian.
Presented in Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian @ NLP for Positive Impact Workshop @ EMNLP2024.
Dataset Details
The dataset includes equal parts human-written and AI-generated news articles, raw and pre-processed.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/neural-news-benchmark.D_ASAP-AES
D_ASAP-AES
This is the train, test, and validation split of the ASAP Automated Essay Scoring dataset,
prepared for use with the S-GRADES benchmark.
Ground truth labels have been removed to prevent leakage during evaluation.
For the original dataset with labels, see below.
Original Dataset
🔗 ASAP-AES on Kaggle
Citation
If you use this dataset, please cite the original:
@misc{asap_aes,
title={ASAP Automated Essay Scoring}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-AES.ik-nlp-22_winemagEcomBench
EcomBench: Where Intelligent Agents Conquer Commerce Realms
🚀 Benchmark Overview
EcomBench is a domain-specific, real-world evaluation framework designed to rigorously assess the capabilities of AI agents in delivering practical support for the complex, ever-evolving demands of e-commerce.
We believe that truly capable AI agents will fundamentally transform how we interact with commerce. E-commerce represents one of the world's most significant economic sectors, with… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/EcomBench.SOLD
SOLD - A Benchmark for Sinhala Offensive Language Identification
In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SOLD.extrinsic_mt_evalCultureBankFAERS-NLP
FAERS-NLP
Version: 1.0Author: sixuexing
GitHub: FAERS-NLP Repository
Dataset Summary
FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction.
Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks.
Dataset Structure
Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/SanaeLaRose/FAERS-NLP.tiny-singleturn-chat-koru_paradetox
ParaDetox: Text Detoxification with Parallel Data (Russian)
This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit
[2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.Merged-CWA
CWA Benchmark: A Seismic Dataset from Taiwan for Seismic Research
Dataset Description
This dataset includes a larger number of seismic events, especially high-magnitude. A comprehensive set of events collected by the
Central Weather Bureau in Taiwan. The CWA benchmark features over 40 attributes and ∼500,000 seismograms, providing
valuable data labels for various seismology-related tasks. In the future, we will keep updating the dataset to ensure its relevance and… See the full description on the dataset page: https://huggingface.co/datasets/NLPLabNTUST/Merged-CWA.mc4_fi_cleaned
Dataset Card for mC4 Finnish Cleaned
Dataset Summary
mC4 Finnish cleaned is cleaned version of the original mC4 Finnish split.
Supported Tasks and Leaderboards
mC4 Finnish is mainly intended to pretrain Finnish language models and word representations.
Languages
Finnish
Dataset Structure
Data Instances
[Needs More Information]
Data Fields
The data have several fields:
url: url of the source as a string
text: text… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/mc4_fi_cleaned.Beyond-Flesch
Beyond-Flesch: ScienceQA Difficulty Classification with Static and Prompt-Based Metrics
A preprocessed subset of ScienceQA for K-12 educational text difficulty classification, along with the static and LLM-derived prompt-based features we use to reproduce Rooein et al. (2024) — Beyond Flesch-Kincaid.
This dataset accompanies our class research project (Option 1: reproducing a paper whose original code was not released).
What's here
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/nlpscu/Beyond-Flesch.statcan-dialogue-dataset-retrieval
Statcan Dialogue Dataset (Processed for Retrieval Tasks)
This is a variant of the Statcan Dialogue Dataset, which we processed specifically for multilingual retrieval (english, french). It contains everything in CSVs, rather than having metadata hosted separately.
Quickstart
from datasets import load_dataset
repo = 'McGill-NLP/statcan-dialogue-dataset-retrieval'
# load english queries, training split
queries_en = load_dataset(repo, 'queries_english', split='train') #… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/statcan-dialogue-dataset-retrieval.nlp_twitter_analysisAmaWar
Amawal Warayni - ⴰⵎⴰⵡⴰⵍ ⴰⵙⵏⵎⴰⵍⴰⵢ ⵏ ⵉⵏⵓⵎⴰⴽ ⵏ ⵡⴰⵙⵙⴰⵖⵏ ⴷ ⵉⵎⵢⴰⴳⵏ ⵏ ⵜⵎⴰⵣⵉⵖⵜ ⵏ ⴰⵢⵜ ⵡⴰⵔⴰⵢⵏ
Bitext scraped from the online AmaWar dictionary of the Tamazight dialect of Ait Warain spoken in northeastern Morocco.
Contains sentences, stories, and poems in Tamazight written in the Neo-Tifinagh script along with their translations into Modern Standard Arabic.
The dataset is split into the following subsets:
examples: Example parallel sentences taken from dictionary entries.
idioms: Idiomatic… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/AmaWar.FakeRecogna
FakeRecogna
FakeRecogna is a dataset comprised of real and fake news. The real news is not directly linked to fake news and vice-versa, which could lead to a biased classification. The news collection was performed by crawlers developed for mining pages of well-known and of great national importance agency news. The web crawlers were developed based on each analyzed webpage, where the extracted information is first separated into categories and then grouped by dates. The plurality… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/FakeRecogna.lc_quad2
Dataset Card for LC-QuAD 2.0 with answers
ENGLISH_TWI_PARALLEL_TEXT
GhanaNLP Twi and English Parallel Data
Twi_to_English
• 1 MB • XLS
English_to_Twi
• 1 MB • XLS
The GhanaNLP Twi dataset contains sentence pairs in Twi and English, designed to support translation models between these two languages. Twi is a Ghanaian local language that lacks extensive digital resources, making this dataset useful for… See the full description on the dataset page: https://huggingface.co/datasets/Ghana-NLP/ENGLISH_TWI_PARALLEL_TEXT.fakerecogna2-abstrativa
FakeRecogna 2.0 - Abstractive
FakeRecogna 2.0 presents the extension for the FakeRecogna dataset in the context of fake news detection. FakeRecogna includes real and fake news texts collected from online media and ten fact-checking sources in Brazil. An important aspect is the lack of relation between the real and fake news samples, i.e., they are not mutually related to each other to avoid intrinsic bias in the data.
The Dataset
The fake news collection was performed on… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/fakerecogna2-abstrativa.Yahoo_Answers_10_categories_for_NLP
Dataset Card for Dataset Name
The Yahoo! Answers topic classification dataset is constructed using 10 largest main categories. Each class contains 140,000 training samples and 6,000 testing samples. Therefore, the total number of training samples is 1,400,000 and testing samples 60,000 in this dataset. From all the answers and other meta-information, we only used the best answer content and the main category information.
Dataset Description
The file classes.txt contains a… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Yahoo_Answers_10_categories_for_NLP.
