datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQuad-MedicalQnADataset
Reference:
"A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019.
leetcode-problem-solutions
LeetCode Solution Dataset
This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling.
Column Descriptions
Column Name
Type
Description
question_slug
string
The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.KMMLU-HARD
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.CogStream
CogStream Dataset
Dataset for CogStream: Context-guided Streaming Video Question Answering.
Overview
CogStream is a streaming video QA dataset designed to evaluate context-guided video reasoning. Models must identify and utilize relevant historical context to answer questions about ongoing video streams.
Statistics:
Split
Videos
QA Pairs
Train
852
55,623
Test
236
15,364
Total
1,088
70,987
Sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%)… See the full description on the dataset page: https://huggingface.co/datasets/SII-KYW/CogStream.leetcode-problem-set
LeetCode Scraper Dataset
This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes.
Dataset Contents
The dataset includes the following files:
problem_set.csv
Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more.
Columns:
acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.healthsearchqa
HealthSearchQA
Dataset of consumer health questions released by Google for the Med-PaLM paper (arXiv preprint).
From the paper:
We curated our own additional dataset consisting of 3,173 commonly searched consumer questions,
referred to as HealthSearchQA. The dataset was curated using seed medical conditions and their
associated symptoms. We used the seed data to retrieve publicly-available commonly searched questions
generated by a search engine, which were displayed to all users… See the full description on the dataset page: https://huggingface.co/datasets/katielink/healthsearchqa.peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.klej-dyk
klej-dyk
Description
The Czy wiesz? (eng. Did you know?) the dataset consists of almost 5k question-answer pairs obtained from Czy wiesz... section of Polish Wikipedia. Each question is written by a Wikipedia collaborator and is answered with a link to a relevant Wikipedia article. In huggingface version of this dataset, they chose the negatives which have the largest token overlap with a question.
Tasks (input, output, and metrics)
The task is to predict if… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-dyk.KazMMLU
MukhammedTogmanov/jana
This dataset contains multiple subsets for different subjects and domains, structured for few-shot learning and evaluation.
Dataset Structure
The dataset is organized into subsets, each corresponding to a specific subject and level. Each subset has two splits:
dev: Few-shot examples (3 rows per subset).
test: The remaining rows for evaluation.
Subsets
The dataset contains the following subsets:
Accounting (University)
Biology (High… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/KazMMLU.CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.hle-multichoiceHumanity Last Exam dataset with extra incorrect answers generated with Qwen3-4B
tiny-singleturn-chat-koKOFFVQA_Data
About this data
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
KOFFVQA is a general-purpose VLM benchmark in the Korean language. For more information, refer to our leaderboard page and the official evaluation code.
This contains the data for the benchmark consisting of images, their corresponding questions, and response grading criteria. The benchmark focuses on free-form visual question answering, evaluating the… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/KOFFVQA_Data.k-beauty-ai-citation-dataset
K-Beauty AI Citation Dataset
Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research.
Canonical source: https://kbeautyanswers.com/dataset/
License: CC BY 4.0
Maintainer: K-Beauty Answers (site)
Initial release: 2026-05-23
What's in it
128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.MolPuzzle_data
MolPuzzle: A Multimodal Benchmark for Molecular Structure Elucidation
Dataset Description:
The MolPuzzle dataset is a newly developed resource designed to challenge Large Language Models with multi-modal, multi-step reasoning tasks (molecular structure elucidation). This dataset consists of 217 diverse and intricate structure elucidation challenges that require LLMs to demonstrate advanced reasoning capabilities, integrating multimodal data and deep chemical understanding… See the full description on the dataset page: https://huggingface.co/datasets/kguo2/MolPuzzle_data.Ncert_datasetKinship
Kinship KG–QA (Hinton)
A lightweight knowledge-graph + question answering (KGQA) resource adapted from the UCI Kinship dataset created by Geoff Hinton.
This dataset provides a small family-tree knowledge graph paired with templated multi-hop QA tasks designed for MultiHop KGQA.
Project: THESEUSPaper: Theseus in the GraphOriginal dataset: UCI Kinship
Key Features
Small knowledge graph with two family trees
Human-readable entities and relations
Templated 1--3 hop… See the full description on the dataset page: https://huggingface.co/datasets/HalcyonSolutions/Kinship.ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
stackexchangeThis dataset is based entirely on HuggingFaceH4/stack-exchange-preferences, but it has been restructured.
All HTML tags have been cleaned out, and the answers column has been turned into the answer column, so instead of answers being stored in JSON format there is now a row for each answer.
Furthermore there is a separate file for every forum instead of a single file.
KorfinQA
FinQA 한국어 번역본
Question, Answer 총 6252 Rows
patient-doctor-qa-tr-321179
Patient Doctor Q&A TR 321179 Veri Kümesi
Patient Doctor Q&A TR 321179 veri kümesi, Patient Doctor Q&A TR 19583, Patient Doctor Q&A TR 167732, Patient Doctor Q&A TR 5695 ve Patient Doctor Q&A TR 95588 veri kümelerinin birleştirilmiş ve karıştırılmış halidir.
Ana Özellikler:
İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları.
Yapı: 2 sütun içerir: Soru, Cevap.Dil: Türkçe.
Potansiyel Kullanım Alanları:
Tıbbi araştırmalar
Doğal Dil… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-321179.ARC-Easy-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC-Easy. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
MoroccanHistory-QA-Datasetethical-framework-UNESCO-Ethics-of-AI
Ethical AI Training Dataset
Introduction
UNESCO's Ethics of Artificial Intelligence, adopted by 193 Member States in November 2021, represents the first global framework for ethical AI development and deployment.
While regional initiatives like The Montréal Declaration for a Responsible Development of Artificial Intelligence emphasize community-driven governance, UNESCO's approach establishes comprehensive international standards through coordinated multi-stakeholder… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework-UNESCO-Ethics-of-AI.TALES-QA
TALES-QA: Cultural Knowledge Question Bank
This repository contains the question bank introduced in our paper “TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories.” The project webpage can be found here: https://cultural-misrepresentations.github.io/
This project evaluates cultural misrepresentations in LLM-generated stories for diverse Indian cultural identities. As part of this effort, we provide a curated question bank of standalone… See the full description on the dataset page: https://huggingface.co/datasets/Kirtibg/TALES-QA.Current_Trivia_Knowledge-benchmark
Current Trivia Knowledge RAG Benchmark
Short Summary:
A 140-QA pair (70 train, 70 test) dataset for real-world RAG evaluation. It features current knowledge questions unavailable to LLMs trained before 2024 (e.g., GPT-4o) across diverse domains, and includes human feedback for the training set, enabling robust assessment of contextual information's critical impact on LLM accuracy.
Introduction & Motivation:
This dataset addresses the critical need for a dynamic… See the full description on the dataset page: https://huggingface.co/datasets/TPelc/Current_Trivia_Knowledge-benchmark.BiomixQA
BiomixQA Dataset
Overview
BiomixQA is a curated biomedical question-answering dataset comprising two distinct components:
Multiple Choice Questions (MCQ)
True/False Questions
This dataset has been utilized to validate the Knowledge Graph based Retrieval-Augmented Generation (KG-RAG) framework across different Large Language Models (LLMs). The diverse nature of questions in this dataset, spanning multiple choice and true/false formats, along with its coverage of various… See the full description on the dataset page: https://huggingface.co/datasets/kg-rag/BiomixQA.NCERT_Physics_11thNCERT_Business_Studies_12th
