datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EnokiQA
EnokiQA
EnokiQA is an annotated dataset for fine-grained hallucination detection in long-form question answering. Each example contains a factual question, a no-context LLM answer, the full Wikipedia article used as verification evidence, sentence-grouped factual triples, and per-triple NLI and hallucination probabilities.
The dataset is dual-granularity: every hallucination label is attached to a claim (an extracted triple) and projected to a character span of the answer. The… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/EnokiQA.CONDAQA
Dataset Card for CondaQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation
Dataset Summary
Data from the EMNLP 2022 paper by Ravichander et al.: "CondaQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation".
If you use this dataset, we would appreciate you citing our work:
@inproceedings{ravichander-et-al-2022-condaqa,
title={CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation}… See the full description on the dataset page: https://huggingface.co/datasets/lasha-nlp/CONDAQA.greek-bar-bench
Dataset Card for GreekBarBench 🇬🇷🏛️⚖️
GreekBarBench is a benchmark designed to evaluate LLMs on challenging legal reasoning questions across five different legal areas from the Greek Bar exams, requiring citations to statutory articles and case facts.
This repository hosts two related benchmarks:
Benchmark
Subsets
Task
GreekBarBench (GBB)
greekbarbench, gbb-jme
Free-text legal reasoning with citations, and LLM-judge meta-evaluation
GreekBarRetrieval (GBR)… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/greek-bar-bench.Kashes
Kashes — a Yiddish Evaluation Benchmark
Kashes (קשיא, pl. קשיות — "questions") is an evaluation benchmark for Yiddish language
models, covering translation, morphosyntax, named-entity recognition, and NLU tasks. It packages the exact frozen evaluation sets used in
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
(COLM 2026), where it is used to evaluate
MameLoshnLM against strong open baselines.
📄 Paper: https://arxiv.org/abs/2608.05850
🧮 Evaluation & reproduction… See the full description on the dataset page: https://huggingface.co/datasets/Yiddish-NLP/Kashes.SciDQA
SciDQA: A Deep Reading Comprehension Dataset over Scientific Papers
📄 Paper | 💻 Code
Scientific literature is typically dense, requiring significant background knowledge and deep comprehension for effective engagement. We introduce SciDQA, a new dataset for reading comprehension that challenges LLMs for a deep understanding of scientific articles, consisting of 2,937 QA pairs. Unlike other scientific QA datasets, SciDQA sources questions from peer reviews by domain experts and… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/SciDQA.EduBench
EduBench 📚
EduBench é um benchmark em português brasileiro para avaliação de Large Language Models (LLMs) em tarefas educacionais, composto por 3,149 questões discursivas extraídas de vestibulares de alta competitividade.
GitHub
Paper
Dataset Description
Fontes
USP: Universidade de São Paulo
UNICAMP: Universidade Estadual de Campinas
UNESP: Universidade Estadual Paulista
Período
2015-2025 (11 anos de provas)
Áreas do… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/EduBench.RoJBMO
RoJBMO: Junior Balkan Mathematical Olympiad Benchmark
RoJBMO is a benchmark of 508 competition mathematics problems drawn from the Junior Balkan Mathematical Olympiad (JBMO), its official shortlists, and the Romanian Team Selection Tests (TST). It is designed to evaluate the mathematical reasoning capabilities of large language models on problems that are underrepresented in existing benchmarks and resistant to training data contamination.
Sources
Source… See the full description on the dataset page: https://huggingface.co/datasets/upb-nlp/RoJBMO.bangla-nlp-catalog
Bangla NLP Catalog
A machine-readable catalog of Bangla (Bengali) NLP resources: 813 papers, 63 datasets, 20 models, and 9 tools across 26 tasks, each tagged by task and carrying a source link.
This is the data behind BanglaNLP Hub. It is metadata about resources, not the resources themselves: no corpora or model weights are redistributed here, only structured records pointing at them.
Why this exists
Bangla is spoken by roughly 240 million people and is still… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/bangla-nlp-catalog.CHASE-QA
CHASE: Challenging AI with Synthetic Evaluations
The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-QA.NLP
r/IPMATtards Reddit Dataset
Dataset Description
This dataset contains scraped posts and comments from the r/IPMATtards subreddit, a community dedicated to aspirants of the Integrated Programme in Management Aptitude Test (IPMAT) in India.
The data is structured into two main components:
Posts: Top-level submissions including titles, body text, scores, and metadata.
Comments: Threaded replies associated with the posts, including recursion depth and parent-child… See the full description on the dataset page: https://huggingface.co/datasets/Nischithhh/NLP.RedirectQA
RedirectQA
RedirectQA is an entity-based factual QA dataset for analyzing how large language models access the same fact through different surface forms of an entity.
This v1.0.0 release is aligned with the dataset described in the paper Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms. The public dataset contains:
61,120 question realizations in the test split
30,560 subject-surface instances
14,672 Wikidata factual triples
14,672… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/RedirectQA.MedQA-CS-ExamBenchmarking LLMs Clinical Skills for Patient-Centered Diagnostics and Documentation
Project github: https://github.com/bio-nlp/MedQA-CS
MedQA-CS-Student dataset: https://huggingface.co/datasets/bio-nlp-umass/MedQA-CS-Student
khmer-nlp-technical-corpus
khmer-nlp-technical-corpus — Khmer Strategic NLP Corpus
Dataset Summary
This dataset contains peer-grade long-form technical treatises (3,000+ words each) in the Khmer language (km / ភាសាខ្មែរ). Every article is normalized and features neural BiGRU+CRF word segmentation with Zero-Width Space (\u200B) injection to prevent token fragmentation in sub-word tokenizers.
Dataset Statistics
Total Documents: 3
Train Documents: 3
Total Words: 8,002
Total… See the full description on the dataset page: https://huggingface.co/datasets/guanvireak/khmer-nlp-technical-corpus.ko.SHP
🚢 Korean Stanford Human Preferences Dataset (Ko.SHP)
이 데이터셋은 자체 구축한 번역기를 활용하여 stanfordnlp/SHP 데이터셋을 번역한 것입니다.
아래의 내용은 해당 번역기로 README 파일을 번역한 것입니다. 참고 부탁드립니다.
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP는 요리에서 법률 조언에 이르기까지 18가지 다른 주제 영역의 질문/지침에 대한 응답에 대한 385K 집단 인간 선호도 데이터 세트이다.
기본 설정은 다른 응답에 대 한 한 응답의 유용성을 반영 하기 위한 것이며 RLHF 보상 모델 및 NLG 평가 모델 (예: SteamSHP)을 훈련 하는 데… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/ko.SHP.abc-multiple-choice
abc-multiple-choice Dataset
abc-multiple-choice は、競技クイズの大会「abc」で使用された4択問題を元に作成された、多肢選択式の質問応答データセットです。
データセットの詳細については、下記の発表資料を参照してください。
鈴木正敏. 4択クイズを題材にした多肢選択式日本語質問応答データセットの構築. 言語処理学会第30回年次大会 (NLP2024) 併設ワークショップ 日本語言語資源の構築と利用性の向上 (JLR2024), 2024. [PDF]
下記の GitHub リポジトリで、本データセットを用いた評価実験のスクリプトを管理しています。
https://github.com/cl-tohoku/abc-multiple-choice
ライセンス
本データセットのクイズ問題の著作権は abc/EQIDEN 実行委員会 に帰属します。
本データセットは研究目的での利用許諾を得ているものです。商用目的での利用は不可とします。
Phunny
Phunny: A Humor-Based QA Benchmark for Evaluating LLM Generalization
Welcome to Phunny, a humor-based question answering (QA) benchmark designed to evaluate the reasoning and generalization abilities of large language models (LLMs) through structured puns.
This repository accompanies our ACL 2025 main track paper:"What do you call a dog that is incontrovertibly true? Dogma: Testing LLM Generalization through Humor"
To reproduce our experiments: Code available on GitHub… See the full description on the dataset page: https://huggingface.co/datasets/disi-unibo-nlp/Phunny.dumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.belebele-fi-filtered-sft
Dataset Card for Finnish-NLP/benebele
Creation process
Finnish subset loaded from facebook/belebele
kowiki-cleaned
NLP-07-ODQA/kowiki-cleaned
Dataset Description
이 데이터셋은 한국어 위키피디아 XML 덤프에서 추출하고 정제한 데이터셋입니다. 청킹 전 단계의 정제된 문서를 포함합니다.
데이터 소스
원본: 한국어 위키피디아 XML 덤프
네임스페이스: 0 (일반 문서)
처리된 페이지 수: 565,484개
전체 페이지 수: 2,157,147개
데이터 정제 과정
위키 마크업 제거: [[링크]], {템플릿}, ==제목== 등 제거
불필요한 섹션 제거: '같이 보기', '외부 링크', '참고 문헌', '각주' 등 제거
필터링: 리다이렉트, 빈 페이지, 스텁 페이지 제거
최소 길이: 200자 이상만 포함
통계
평균 텍스트 길이: 1871.8자
데이터 구조
각 데이터 포인트는 다음 필드를 포함합니다:
content: 정제된 텍스트 (전체… See the full description on the dataset page: https://huggingface.co/datasets/NLP-07-ODQA/kowiki-cleaned.tamil-english-corpus
Tamil-English Retrieval Corpus
A high-quality multilingual retrieval corpus constructed from the Mozhi Tamil Corpus and machine-translated into English using IndicTrans2.
Dataset Summary
This dataset contains Tamil documents paired with English translations.
The corpus was created by filtering high-quality documents from the Mozhi Tamil Corpus and translating them using AI4Bharat's IndicTrans2 translation model.
The resulting corpus is intended to support:… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tamil-english-corpus.aihub_mrc_admin행정 문서 대상 기계독해 데이터
popqa-olmo-3-7b-instruct-temp0.9-samples99-logprobs
OLMo-3-7B-Instruct self-consistency generations with logprobs on PopQA
This dataset contains 99 self-consistency generations per question for the
PopQA benchmark, produced with allenai/OLMo-3-7B-Instruct at temperature
0.9, together with token-level log probabilities for each completion.
The file is intended for post-hoc analysis, self-consistency curves, adaptive
stopping, and related aggregation methods.
Source
Base benchmark: PopQA
Model: allenai/OLMo-3-7B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/popqa-olmo-3-7b-instruct-temp0.9-samples99-logprobs.drbodebench_medicamentos
Medication-Focused Clinical Benchmark from DrBodeBench
Dataset Details
To evaluate retrieval capabilities in higher-level reasoning scenarios, we created a second benchmark derived from the Portuguese medical benchmark DrBodeBench. This benchmark aggregates questions from Brazilian medical examinations, including the Revalida and the FUVEST direct-access residency exam. From DrBodeBench, we curated a specific subset of questions that exclusively pertains to… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/drbodebench_medicamentos.multilingual-repo-qa
UPD: Release in progress, stay tuned.
MultiRepoQA
MultiRepoQA is a multilingual benchmark for question answering over complete software repositories. It contains 783 validated canonical questions across 30 open-source repositories, with aligned English, German, and Ukrainian versions (2,349 language-specific examples). Questions cover local implementation details, cross-file behavior, repository-wide flows, and maintenance impact.
The benchmark was introduced in… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/multilingual-repo-qa.QA_hukum_samplesThis dataset sample was constructed by generating QA pairs from Law No. 17 of 2008 on Shipping (Undang-Undang No 17 Tahun 2008 Tentang Pelayaran). It is then manually verified by human validators and reviewers, resulting in QA pairs with fine-grained label categories.
