datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FAERS-NLP
FAERS-NLP
Version: 1.0Author: sixuexing
GitHub: FAERS-NLP Repository
Dataset Summary
FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction.
Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks.
Dataset Structure
Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/sixuexing/FAERS-NLP.D_persuade_2
Persuade_2
The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and
Understanding Argumentative and Discourse Elements) contains over 25,000
argumentative essays written by 6th–12th grade students in the United States,
covering 15 distinct prompts across two writing tasks: independent and
source-based writing. The corpus also provides detailed individual and
demographic information for each writer.
This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.D_ASAP-AES
D_ASAP-AES
This is the train, test, and validation split of the ASAP Automated Essay Scoring dataset,
prepared for use with the S-GRADES benchmark.
Ground truth labels have been removed to prevent leakage during evaluation.
For the original dataset with labels, see below.
Original Dataset
🔗 ASAP-AES on Kaggle
Citation
If you use this dataset, please cite the original:
@misc{asap_aes,
title={ASAP Automated Essay Scoring}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-AES.ik-nlp-22_winemagFAERS-NLP
FAERS-NLP
Version: 1.0Author: sixuexing
GitHub: FAERS-NLP Repository
Dataset Summary
FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction.
Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks.
Dataset Structure
Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/SanaeLaRose/FAERS-NLP.Merged-CWA
CWA Benchmark: A Seismic Dataset from Taiwan for Seismic Research
Dataset Description
This dataset includes a larger number of seismic events, especially high-magnitude. A comprehensive set of events collected by the
Central Weather Bureau in Taiwan. The CWA benchmark features over 40 attributes and ∼500,000 seismograms, providing
valuable data labels for various seismology-related tasks. In the future, we will keep updating the dataset to ensure its relevance and… See the full description on the dataset page: https://huggingface.co/datasets/NLPLabNTUST/Merged-CWA.Beyond-Flesch
Beyond-Flesch: ScienceQA Difficulty Classification with Static and Prompt-Based Metrics
A preprocessed subset of ScienceQA for K-12 educational text difficulty classification, along with the static and LLM-derived prompt-based features we use to reproduce Rooein et al. (2024) — Beyond Flesch-Kincaid.
This dataset accompanies our class research project (Option 1: reproducing a paper whose original code was not released).
What's here
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/nlpscu/Beyond-Flesch.extrinsic_mt_evalfakerecogna2-abstrativa
FakeRecogna 2.0 - Abstractive
FakeRecogna 2.0 presents the extension for the FakeRecogna dataset in the context of fake news detection. FakeRecogna includes real and fake news texts collected from online media and ten fact-checking sources in Brazil. An important aspect is the lack of relation between the real and fake news samples, i.e., they are not mutually related to each other to avoid intrinsic bias in the data.
The Dataset
The fake news collection was performed on… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/fakerecogna2-abstrativa.DeepScaleR_Difficulty
Difficulty Estimation on DeepScaleR
We annotate the entire DeepScaleR dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
DeepScaleR is a curated dataset of 40,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/DeepScaleR_Difficulty.fakerecogna2-extrativa
FakeRecogna 2.0 Extractive
FakeRecogna 2.0 presents the extension for the FakeRecogna dataset in the context of fake news detection. FakeRecogna includes real and fake news texts collected from online media and ten fact-checking sources in Brazil. An important aspect is the lack of relation between the real and fake news samples, i.e., they are not mutually related to each other to avoid intrinsic bias in the data.
The Dataset
The fake news collection was performed on… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/fakerecogna2-extrativa.aikyatansinha_cybersecurity-cves-for-nlp-dataset
Cybersecurity CVEs for NLP Dataset
Every CVE since 1999, scrubbed and perfectly formatted for NLP tasks
Dataset Info
Source: Kaggle
Original Size: 38.28 MB
Kaggle Downloads: 36
Files: 1
Files
NVD_Cybersecurity_Dataset.csv
Mirrored from Kaggle
PubMed-Cancer-NLP-Textual-Dataset
PubMed-Cancer-NLP-Textual-Dataset
This dataset has been obtained from PubMed for research purposes. README will be updated with time.
Dataset Details
Dataset Description
It has multiple cancer samples with labels with their title and abstract from PubMed Repository.
Curated by: Om Aryan
Dataset Sources
Repository: https://pubmed.ncbi.nlm.nih.gov
Dream_NLP_FineTuneGSM8K_Difficulty
Difficulty Estimation on DeepScaleR
We annotate the entire GSM8K dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/GSM8K_Difficulty.tr_rteRoJBMO
RoJBMO: Junior Balkan Mathematical Olympiad Benchmark
RoJBMO is a benchmark of 508 competition mathematics problems drawn from the Junior Balkan Mathematical Olympiad (JBMO), its official shortlists, and the Romanian Team Selection Tests (TST). It is designed to evaluate the mathematical reasoning capabilities of large language models on problems that are underrepresented in existing benchmarks and resistant to training data contamination.
Sources
Source… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/RoJBMO.D_ASAP-SAS
D_ASAP-SAS
This is the train, test, and validation split of the ASAP Short Answer Scoring dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation.
For the original dataset with labels, see below.
Original Dataset
🔗 ASAP-SAS on Kaggle
Citation
If you use this dataset, please cite the original:
@misc{asapsas2012,
author={Barbara and Hamner, Ben and Morgan, Jaison and lynnvandev and… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-SAS.NyayaAnumana-Transformers-ResultsFLUE-FiQA
Dataset Summary
Homepage: https://sites.google.com/view/salt-nlp-flang
Models: https://huggingface.co/SALT-NLP/FLANG-BERT
Repository: https://github.com/SALT-NLP/FLANG
FLUE
FLUE (Financial Language Understanding Evaluation) is a comprehensive and heterogeneous benchmark that has been built from 5 diverse financial domain specific datasets.
Sentiment Classification: Financial PhraseBankSentiment Analysis, Question Answering: FiQA 2018New Headlines Classification:… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/FLUE-FiQA.ImplicatureX
ImplicatureX
More information can be found at https://github.com/cesare-spinoso/ImplicatureX.
import pandas as pd
# skiprows=1: the first line is a leading comment, not part of the header
df = pd.read_csv("implicatureX.csv", skiprows=1)
Citation
If you use our data, please cite us:
@misc{piano2026evaluatingcommunicativebeliefupdates,
title={Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation}… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/ImplicatureX.IBM-Debater-ArgKPCHASE-QA
CHASE: Challenging AI with Synthetic Evaluations
The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-QA.docling-nlp-datasetsThis repository contains the models used for docling-nlp.
Contents
This model repository packages the pretrained assets used by Docling’s NLP
components:
CRF models for material classification and English part-of-speech tagging
fastText models for language detection, metadata, semantic, topic, and person-name classification
Regular-expression assets for geographic-location extraction and unit handling
A default tokenizer model
Correct workflow to add new files… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/docling-nlp-datasets.almaz-asr-roster
The ALMAZ ASR Roster
Archived at Zenodo: 10.5281/zenodo.22761871
(concept DOI, always resolves to the latest version).
A curated catalog of Azerbaijani speech-to-text artifacts: corpora, models,
services, benchmarks and tools. Companion to the
ALMAZ Resource Roster,
which does the same for text.
Schema matches the text roster so the two join, plus three columns speech needs
and text does not: hours, condition, and verified.
The verified column
The standard way a… See the full description on the dataset page: https://huggingface.co/datasets/almaz-nlp/almaz-asr-roster.D_Ielts_Writing_Dataset
D_Ielts_Writing_Dataset
This dataset contains IELTS Writing scored essays, prepared for use with the S-GRADES benchmark. The test split ground truth labels have been removed to prevent leakage during evaluation.
Original Dataset
🔗 IELTS Writing Scored Essays Dataset on Kaggle
Citation
If you use this dataset, please cite the original source:
@misc{mazlum2023ielts,
title={IELTS Writing Scored Essays Dataset},
author={Mazlum, Ibrahim},
year={2023}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_Ielts_Writing_Dataset.orz_math_difficulty
Difficulty Estimation on Open Reasoner Zero
We annotate the entire Open Reasoner Zero dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction.
Open Reasoner Zero is a curated a dataset of 57,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models.
Difficulty Scoring Method
Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/orz_math_difficulty.MATH_Difficulty
Difficulty Estimation on MATH
We annotate the entire MATH dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation.
The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/MATH_Difficulty.SemiSOLD
SOLD - A Benchmark for Sinhala Offensive Language Identification
In this repository, we introduce the {S}inhala {O}ffensive {L}anguage {D}ataset (SOLD) and present multiple experiments on this dataset. SOLD is a manually annotated dataset containing 10,000 posts from Twitter annotated as offensive and not offensive at both sentence-level and token-level. SOLD is the largest offensive language dataset compiled for Sinhala. We also introduce SemiSOLD, a larger dataset containing more… See the full description on the dataset page: https://huggingface.co/datasets/sinhala-nlp/SemiSOLD.D_ASAP_plus_plus
D_ASAP_plus_plus
This is the train, test, and validation split of the ASAP++ dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation.
Original Dataset
ASAP++ enriches the original ASAP dataset with attribute-specific essay scores (content, organization, style, etc.).
🔗 ASAP++ Official Page
Citation
If you use this dataset, please cite the original:
@inproceedings{mathias2018asap++… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP_plus_plus.
