datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sudoku-extreme
Hardest Sudoku Puzzle Dataset V2
This dataset contains a mixture of easy and very hard Sudoku puzzles collected from the Sudoku community.
Dataset Composition
Sources
tdoku benchmarks
enjoysudoku
Easy Puzzles (1.1M)
puzzles0_kaggle
puzzles1_unbiased
puzzles2_17_clue
Hard Puzzles (3.1M)
puzzles3_magictour_top1465
puzzles4_forum_hardest_1905
puzzles6_forum_hardest_1106
ph_2010/01_file1.txt
Dataset Characteristics
All… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/sudoku-extreme.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.function_calling_extended
Trelis Function Calling Dataset
UPDATE: As of Dec 5th 2023, there is a v3 of this dataset now available from here.
Allows models to be fine-tuned for function-calling.
The dataset is human generated and does not make use of Llama 2 or OpenAI!
Contains 59 training and 17 test rows
Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file, clear_chat
Access this dataset by purchasing a license HERE.
Alternatively… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_extended.BaisBenchThis is the dataset container for the Biological AI Scientist Benchmark (BAISBench). It's a benchmark designed to assess AI scientists' ability to generate biological discoveries through data analysis and reasoning with external knowledge.
This benchmark contains two tasks:
Data Process and cell Type Annotation task (BAIS-DPTA): This task includes 15 single-cell datasets to assess AI scientists' ability to annotate cell types, a fundamental challenge in single-cell analysis. To enable… See the full description on the dataset page: https://huggingface.co/datasets/EperLuo/BaisBench.ExploreToM
Data sample for ExploreToM: Program-guided adversarial data generation for theory of mind reasoning
ExploreToM is the first framework to allow large-scale generation of diverse and challenging theory of mind data for robust training and evaluation.
Our approach leverages an A* search over a custom domain-specific language to produce complex story structures and novel, diverse, yet plausible scenarios to stress test the limits of LLMs.
Our A* search procedure aims to find… See the full description on the dataset page: https://huggingface.co/datasets/facebook/ExploreToM.APEX-v1-extended
APEX-v1-extended
The AI Productivity Index (APEX) is a benchmark from Mercor for assessing whether frontier models are capable of performing economically valuable tasks across four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD).
APEX-v1-extended doubles the heldout evaluation set from n=200 to n=400, with increased complexity and variety. On average, tasks take over two-and-a-half hours for seasoned professionals to… See the full description on the dataset page: https://huggingface.co/datasets/mercor/APEX-v1-extended.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.thai-onet-m6-exam
Thai O-Net Exams Dataset
Overview
The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems.
Dataset Source
Thai National Institute of Educational Testing Service (NIETS)
Maintainer
Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.paul_graham_essays
Dataset Card for Paul Graham Essay Collection Dataset
Dataset Description
This dataset contains a complete collection of essays written by Paul Graham, a renowned programmer, venture capitalist, and essayist. The essays cover a wide range of topics including startups, programming, technology, entrepreneurship, and personal growth. Each essay has been cleaned and processed to extract the title, date of publication, and the full text content.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/sgoel9/paul_graham_essays.Bitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.EnviroExam
Dataset Summary
EnviroExam focuses on 42 core courses from the environmental science curriculum at Harbin Institute of Technology, after excluding general, duplicate, and practical courses from a total of 141 courses across undergraduate, master's, and doctoral programs.
For these 42 courses, initial draft questions were generated using GPT-4 and Claude, combined with customized prompts. These drafts were then refined and proofread manually, resulting in a total of 1,290… See the full description on the dataset page: https://huggingface.co/datasets/enviroscientist/EnviroExam.ClawBenchPro
ClawBenchPro
ClawBenchPro is a compact, builder-based workplace-agent benchmark package exported from Nanoclaw.
It contains task YAML files, prompts, task-local environment builders, skills, evaluation manifests,
provenance metadata, and checksums.
Included Splits
Dataset
Tasks
Groups
round_01_aligned_mix_800
800
base, hard_aligned, multi_turn_aligned, skills_aligned
persona_aligned_mix_200
200
base, hard, multi_turn, skills
Directory Layout… See the full description on the dataset page: https://huggingface.co/datasets/ErenJaegerYeager/ClawBenchPro.Ethical-Reasoning-in-Mental-Health-v1This repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI.
Overview
Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias.
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.lar-echr
Dataset Card for LAR-ECHR
Dataset Details
Dataset Description
Curated by: Odysseas S. Chlapanis
Funded by: Archimedes Research Unit
Language (NLP): English
License:
CC BY-NC-SA (Creative Commons / Attribution-NonCommercial-ShareAlike)
Read more: https://creativecommons.org/licenses/by-nc-sa/4.0/
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Uses… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/lar-echr.eecc
Extrinsic Evaluation of Cultural Competence in LLMs
In this repository, we release the data used in our paper "Extrinisic Evaluation of Cultural Competence in Large Language Models".
In this work, we analyse the extent and characteristics of variations in model outputs when explicit cue of culture, nationality is present in the prompt. We evaluate models on two user-facing tasks: Question Answering (QA) and Story Generation.
We use 193 nationalities present in… See the full description on the dataset page: https://huggingface.co/datasets/shaily99/eecc.stack-exchange-dataset
Overview
This dataset consists of three TSV files, namely: cs.tsv, ds.tsv, and p.tsv.
Each file includes the data for the questions asked on a Stack Exchange (SE) question-answering community, from the creation of the community until May 2021.
cs.tsv --> Computer Science SE
ds.csv --> Data Science SE
p.csv --> Political Science SE
File Structure
Each file has the following columns:
id: the question id
title: the title of the question
body: the body or text of the… See the full description on the dataset page: https://huggingface.co/datasets/habedi/stack-exchange-dataset.CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.EgoAVU_data
[CVPR2026 HIGHLIGHT] EgoAVU, [ICASSP2026 Oral] Exploring Audio Hallucination in Egocentric Video Understanding
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding and Exploring Audio Hallucination in Egocentric Video Understanding
See our github for the code and setup instructions.
Check out our homepage, paper (CVPR) and paper (ICASSP) for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual… See the full description on the dataset page: https://huggingface.co/datasets/facebook/EgoAVU_data.EC-Guide
This repo is only used for dataset viewer. Please download from here.
Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5)
The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.Flutter-Code-with-Questions-Dataset-English
🧠 Flutter Code with Questions Dataset (English)
This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development.
📂 Dataset Structure
The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes:
A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.engsaf
Engineering Short Answer Feedback
A collection of real short-answer responses from engineering exams across multiple engineering domains.
Background
In recent years, there has been a growing interest in using Artificial Intelligence (AI) to automate student assessment in education.
Among different types of assessments, summative assessments play a crucial role in evaluating a student's understanding level of a course.
Such examinations often involve short-answer… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/engsaf.econ_logic_qa
EconLogicQA
EconLogicQA is a benchmark designed to test the sequential reasoning skills of large language models (LLMs) in economics, business,
and supply chain management. It diverges from typical benchmarks by requiring models to understand and sequence multiple interconnected
events, capturing complex economic logics. The benchmark includes multi-event scenarios and a thorough suite of evaluations to assess
proficiency in economic contexts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yinzhu-quan/econ_logic_qa.aware-bench
EvalDetectBench
Companion dataset for EvalDetectBench, a benchmark that measures whether LLMs can detect that they are in evaluation or deployment contexts.
Three folders, each with its own README.md and croissant.json
(Croissant 1.1):
collected_trajectories/ raw trajectory pool (per-model JSON)
measure_logs/ measure-stage outputs (.eval logs + CSV export)
paper_replication/ CSV inputs that feed the paper figures and tables
Folder map… See the full description on the dataset page: https://huggingface.co/datasets/el7982/aware-bench.Scholarly-Epistemic-Engine
Dataset Card for Scholarly-Epistemic-Engine: arXiv cs.AI Corpus and Embeddings
This dataset contains the processed text, metadata, and semantic vector embeddings of approximately 90,000 scholarly articles from the arXiv Computer Science - Artificial Intelligence (cs.AI) category, spanning from 1993 to December 2024. It is designed to support Retrieval-Augmented Generation (RAG) systems and semantic knowledge discovery.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/whyamanbhardwaj/Scholarly-Epistemic-Engine.granola-entity-questions
GRANOLA Entity Questions Dataset Card
Dataset details
Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions)
Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.thai_buddhist_studies_exam
Thai Buddhist Studies Examination (Nak Tham)
This repository contains multiple-choice questions from the Thai Buddhist Studies
(Nak Tham) examination (2020, 2022, 2023). This dataset can be used for a benchmark for evaluating Large Language Models'
understanding of Thai Buddhist concepts and teachings.
Dataset Statistics
Year
Number of Multiple Choice Questions
2020
1,350
2022
1,400
2023
1,350
Phra Udom thought on the exam: We have reviewed the Nak… See the full description on the dataset page: https://huggingface.co/datasets/biodatlab/thai_buddhist_studies_exam.precision-evidence-bench
Precision Evidence Bench
Precision Evidence Bench is a Precision Medicine Benchmark from
Atropos Health, the world's largest creator of
real-world evidence (RWE) for clinical decision support. It evaluates how well
large language models (LLMs) answer clinical questions that are grounded in
patient context and inclusive of patient history: not "which treatment is
better in general", but "which treatment is better for this patient", with a
specific comorbidity, age, prior therapy… See the full description on the dataset page: https://huggingface.co/datasets/atroposhealth/precision-evidence-bench.EndoBench
EndoBench
🍎 Homepage|💻 GitHub|🤗 Dataset|📖 Paper
This repository is the official implementation of the paper EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis.
🚀 News
[03/2026] We release the EndoVQA-Instruct Dataset at here.
[21/10/2025] We release a new open-set challenging VQA benchmark EndoBench-extended.
[19/09/2025] 🎉🎉Our EndoBench was accepted by NeurIPS'25 D&B Track!!!
☀️ Tutorial
EndoBench is… See the full description on the dataset page: https://huggingface.co/datasets/Saint-lsy/EndoBench.
