datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Calc-asdiv_a
Dataset Card for Calc-asdiv_a
Summary
The dataset is a collection of simple math word problems focused on arithmetics. It is derived from the arithmetic subset of ASDiv (original repo).
The main addition in this dataset variant is the chain column. It was created by converting the solution to a simple html-like language that can be easily
parsed (e.g. by BeautifulSoup). The data contains 3 types of tags:
gadget: A tag whose content is intended to be evaluated by calling… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/Calc-asdiv_a.SWE-chat
SWE-chat: Coding Agent Interactions From Real Users in the Wild
📄 Paper: arxiv.org/abs/2604.20779
🌐 Website: swe-chat.com
Dataset Summary
SWE-chat captures real-world AI coding sessions from developers using AI coding assistants (Claude Code, Codex, Gemini CLI, and others via the Entire.io CLI). Each session includes the full conversation transcript, tool calls, thinking traces, code changes, and attribution of human vs. agent-authored code.
Dataset Size… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SWE-chat.proofwriter_processed_OWAQuRatedPajama-260B
QuRatedPajama
Paper: QuRating: Selecting High-Quality Data for Training Language Models
A 260B token subset of cerebras/SlimPajama-627B, annotated by princeton-nlp/QuRater-1.3B with sequence-level quality ratings across 4 criteria:
Educational Value - e.g. the text includes clear explanations, step-by-step reasoning, or questions and answers
Facts & Trivia - how much factual and trivia knowledge the text contains, where specific facts and obscure trivia are preferred over more… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRatedPajama-260B.german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.Mintaka_Graph_Features_T5-xl-ssm
Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm"
More Information needed
assin2
Dataset Card for ASSIN 2
Dataset Summary
The ASSIN 2 corpus is composed of rather simple sentences. Following the procedures of SemEval 2014 Task 1.
The training and validation data are composed, respectively, of 6,500 and 500 sentence pairs in Brazilian Portuguese,
annotated for entailment and semantic similarity. Semantic similarity values range from 1 to 5, and text entailment
classes are either entailment or none. The test data are composed of approximately 3,000… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/assin2.mc4_3.1.0_fi_cleaned
Dataset Card for "mc4_3.1.0_fi_cleaned"
More Information needed
details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct
Dataset Card for Evaluation run of princeton-nlp/Llama-3-8B-ProLong-512k-Instruct
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-8B-ProLong-512k-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct.EgyMMLU
Dataset Card for EgyMMLU
Dataset Description
Dataset Summary
Languages
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Personal and Sensitive Information
Considerations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Citation Information
Dataset Summary
EgyMMLU is a benchmark created to test the… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/EgyMMLU.corr2cause
Dataset card for corr2cause
TODO
LitSearch
LitSearch: A Retrieval Benchmark for Scientific Literature Search
This dataset contains the query set and retrieval corpus for our paper LitSearch: A Retrieval Benchmark for Scientific Literature Search. We introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers. LitSearch is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/LitSearch.FOLIOoscar_2301_fi_cleaned
Dataset Card for "oscar_2301_fi_cleaned"
More Information needed
lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.hle-context-baseline-deepLow-resource-QE-DA-dataset
Low-resource QE-DA Dataset
Direct Assessment (DA) quality estimation data for English→Indic (Gujarati, Hindi, Marathi, Tamil, Telugu) and related Estonian/Nepali/Sinhala pairs, released with the ALOPE work on LLM-based QE.
Paper: Sindhujan, A., Qian, S., Matthew, C.C.C., Orasan, C., and Kanojia, D. (2024). ALOPE: Adaptive Layer Optimization for Translation Quality Estimation using Large Language Models. In Second Conference on Language Modeling. (arXiv)
Task: Sentence-level quality… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/Low-resource-QE-DA-dataset.KGQASubgraphsRanking
📰 News
[12/2023] Publishing of the original paper "Large Language Models Meets Knowledge Graph to Answer Factoid Questions". This paper first introduces the novelty of the extracted subgraphs; which provide valuable information for different methods of ranking. The paper leveraged T5-like models, and achieve SOTA results with Graph2Text ranking.
Dataset Summary
KGQASubgraphsRanking is the total-packaged dataset for both publications mentioned in the News section.… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/KGQASubgraphsRanking.multilingual-medical-reasoning-tracesThis datasets containes the traces generated to answer multiple-choice medical questions in Italian, Englihs, and Spanish.
The dataset is structured in 3 parts, one per language. Each part is composed by 2 splits, one containing the examples generated from medqa, one from medmcqa.
The columns are:
id, representing an unique identifier
full_question, representing the medical question
options, a dictionary of options to answer the question and their identifiers
list_of_options, a list of the… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/multilingual-medical-reasoning-traces.mathdial
Mathdial dataset
https://arxiv.org/abs/2305.14536
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems.
MathDial is grounded in math word problems as well as student confusions which provide a challenging testbed for creating faithful and equitable dialogue tutoring models able to reason over complex information. Current models achieve high accuracy in solving such problems but they fail in the task of teaching.
Data… See the full description on the dataset page: https://huggingface.co/datasets/eth-nlped/mathdial.FAERS-NLP
FAERS-NLP
Version: 1.0Author: sixuexing
GitHub: FAERS-NLP Repository
Dataset Summary
FAERS-NLP is a cleaned and processed version of the FDA Adverse Event Reporting System (FAERS), formatted for natural language retrieval and drug–adverse effect–disease relation extraction.
Each record corresponds to a single adverse event report, including structured and semi-structured fields suitable for NLP tasks.
Dataset Structure
Each CSV row contains the following… See the full description on the dataset page: https://huggingface.co/datasets/sixuexing/FAERS-NLP.FLD.v2
Dataset Card for "FLD.v2"
For the schema of the dataset, see here.
For the whole of the project, see our project page.
More Information needed
lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private
Dataset Card for Evaluation run of hkust-nlp/dart-math-llama3-8b-prop2diff
Dataset automatically created during the evaluation run of model hkust-nlp/dart-math-llama3-8b-prop2diff
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private.assin
Dataset Card for ASSIN
Dataset Summary
The ASSIN (Avaliação de Similaridade Semântica e INferência textual) corpus is a corpus annotated with pairs of sentences written in
Portuguese that is suitable for the exploration of textual entailment and paraphrasing classifiers. The corpus contains pairs of sentences
extracted from news articles written in European Portuguese (EP) and Brazilian Portuguese (BP), obtained from Google News Portugal
and Brazil, respectively. To… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/assin.D_persuade_2
Persuade_2
The PERSUADE 2.0 corpus (Persuasive Essays for Rating, Selecting, and
Understanding Argumentative and Discourse Elements) contains over 25,000
argumentative essays written by 6th–12th grade students in the United States,
covering 15 distinct prompts across two writing tasks: independent and
source-based writing. The corpus also provides detailed individual and
demographic information for each writer.
This is the train, test, and validation split of the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_persuade_2.EnokiQA
EnokiQA
EnokiQA is an annotated dataset for fine-grained hallucination detection in long-form question answering. Each example contains a factual question, a no-context LLM answer, the full Wikipedia article used as verification evidence, sentence-grouped factual triples, and per-triple NLI and hallucination probabilities.
The dataset is dual-granularity: every hallucination label is attached to a claim (an extracted triple) and projected to a character span of the answer. The… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/EnokiQA.D_ASAP-AES
D_ASAP-AES
This is the train, test, and validation split of the ASAP Automated Essay Scoring dataset,
prepared for use with the S-GRADES benchmark.
Ground truth labels have been removed to prevent leakage during evaluation.
For the original dataset with labels, see below.
Original Dataset
🔗 ASAP-AES on Kaggle
Citation
If you use this dataset, please cite the original:
@misc{asap_aes,
title={ASAP Automated Essay Scoring}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-AES.instagram-political-communication-it
Instagram Political Communication (Italy) — NLP-POL
Dataset Summary
This dataset is part of NLP-POL (NLP for Political Communication), a research project focused on the analysis of political communication strategies through Natural Language Processing.
The dataset contains Instagram posts and comments collected from more than 300 Italian political figures, primarily members of the Italian Parliament (with a strong focus on Deputies). It includes both content published by… See the full description on the dataset page: https://huggingface.co/datasets/NLP-POL/instagram-political-communication-it.TopiOCQATopiOCQA is an information-seeking conversational dataset with challenging topic switching phenomena.hle-context-baseline-gpt41
