datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Question-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.enron_emails_sample_questionspdf_science_questions_verified_r1_traces__2_24_25
Dataset card for pdf_science_questions_verified_r1_traces__2_24_25
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"success": true,
"page_count": 37,
"page_number": 1,
"question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verified_r1_traces__2_24_25.PDF_and_SCP_unfiltered_organic_chemistry_questionsArt-Vision-Question-Answering-Dataset
Art Vision Question Answering Dataset
🎨 A curated dataset for training AI models on digital artwork analysis and visual question answering.
Dataset Overview
This dataset contains 577 question-answer pairs extracted from artwork conversations, designed for training multimodal AI models on art analysis tasks.
✨ Key Features
🖼️ Visual Thumbnails: Artwork images displayed directly in the dataset viewer
💬 Rich Q&A: Expert-level questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/OneEyeDJ/Art-Vision-Question-Answering-Dataset.betterbench-b1-all-questionsstackoverflow-questions-long
stackoverflow questions for text classification: 'long'
This is pacovaldez/stackoverflow-questions filtered for 1024 GPT2 tokens or more in title + body
https://huggingface.co/datasets/pacovaldez/stackoverflow-questions
pandas-questionspdf_science_questions_verifiable_r1_traces__2_24_25
Dataset card for pdf_science_questions_verifiable_r1_traces__2_24_25
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"success": true,
"page_count": 37,
"page_number": 1,
"question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verifiable_r1_traces__2_24_25.yher-chemistry-question-bank
YHer Chemistry Question Bank
The data layer of an evidence-bound diagnostic learning system for Shanghai high-school chemistry (Chris-TLC/YHer-skill).
Every record in this dataset is derived from publicly released Shanghai gaokao and mock examination papers through deterministic mechanical structuring: text extraction, layout repair, and answer alignment. No content is model-generated.
What's inside
The dataset ships in two configs:
Config
Records
Content… See the full description on the dataset page: https://huggingface.co/datasets/Chris-TLC/yher-chemistry-question-bank.quora-question-pairsgraph-anomaly-questions
Questions (graph anomaly detection)
Users of the Yandex Q question-answering service, connected by answering interactions. The minority class marks users by activity outcome; at a 3.0% base rate this is a realistic rare-anomaly regime.
Nodes
48,921
Node features
301
Edges
153,540
Outliers
1,460 (3.0%)
Label type
adjudicated
Label source
Yandex Q activity outcome, Platonov et al. 2023
Viewer. Two configs: nodes (default) and edges.
Files. nodes.parquet… See the full description on the dataset page: https://huggingface.co/datasets/JaySuryavanshi/graph-anomaly-questions.sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912
mlfoundations-dev/sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 26.94% ± 4.54%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
19.70%
39
198
2
23.23%
46
198
3
37.88%
75
198
is-trivia-questions
Icelandic trivia questions
Icelandic trivia question compiled and created by Sveinn Steinarsson, Valur Freyr Steinarsson, and Svavar Kjarrval https://github.com/sveinn-steinarsson/is-trivia-questions
Dálkanúmer
Valfrjálst
Lýsing
1
Nei
Flokkanúmer
2
Já
Undirflokkur ef til staðar
3
Nei
Erfiðleikastig: 1: Létt, 2: Meðal, 3: Erfið
4
Já
Gæðastig: 1: Slöpp, 2: Góð, 3: Ágæt
5
Nei
Spurningin
6
Nei
Svarið
Flokkanúmer
Flokkanafn
1
Almenn kunnátta
2
Náttúra… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/is-trivia-questions.turkish-university-entrance-exam-questionsquran-question-answer-context
Dataset Card for "quran-question-answer-context"
Dataset Summary
Translated the original dataset from Arabic to English and added the Surah ayahs to the context column.
Usage
from datasets import load_dataset
dataset = load_dataset("nazimali/quran-question-answer-context")
DatasetDict({
train: Dataset({
features: ['q_id', 'question', 'answer', 'q_word', 'q_topic', 'fine_class', 'class', 'ontology_concept', 'ontology_concept2', 'source', 'q_src_id'… See the full description on the dataset page: https://huggingface.co/datasets/nazimali/quran-question-answer-context.superintelligence-search-questions
Keyfern: what people search about superintelligence (Sept 2026)
Need it fresh, filtered or via API? This free file is a snapshot (search questions at the last refresh), last updated 2026-09-25.
Keyfern on Apify ($0.001 per keyword idea): gets keyword ideas and questions for your own seeds from 5 search engines.
Swellmeter on Apify ($0.003 per keyword analyzed): checks Google Trends direction for your own keywords with a plain rising/falling verdict, daily if scheduled.… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/superintelligence-search-questions.fine-reasoning-questions
Dataset Card for Fine Reasoning Questions
Dataset Description
Can we generate reasoning datasets for more domains using web text?
Note: This dataset is submitted partly to give an idea of the kind of dataset you could submit to the reasoning datasets competition. You can find out more about the competition in this blog post.
You can also see more info on using Inference Providers with Curator here
The majority of reasoning datasets on the Hub are focused on maths… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/fine-reasoning-questions.nor_agriculture_multi_hop_questions_bench
Nor Agriculture Multi Hop Questions Bench
This work is related to the project in adapting LLM to answer questions about Norwegian Agriculture in Norwegian.
The dataset was generated using YourBench (v0.9.0), an open-source framework for generating domain-specific benchmarks from document collections.
It needs further cleaning, verification in regards to citations and validation for diversity and topics coverage.
Also, it can play a role as a prove of concept for generating a… See the full description on the dataset page: https://huggingface.co/datasets/norjordAI/nor_agriculture_multi_hop_questions_bench.visual-question-answering-checkpoint-downloadsNSFW-questionshalluhard-seed-questionsThis is the seed question set for the benchmark HalluHard
We design some difficult seed questions that elicit a citation-grounded multi-turn open-ended generation, covering four domains: legal, research, medical, and coding.
Below is the Top15 models on our leaderboard. Please check our website for more models and turn-wise/domain-wise statistics!
Rank
Model
Hallucination Rate
Legal
Research
Medical
Coding
1
Claude-Opus-4.5-Web-Search
30.2
33.0
29.6
29.2
29.0
2… See the full description on the dataset page: https://huggingface.co/datasets/dyfan/halluhard-seed-questions.mathflow_questions_appAmbigNQ-clarifying-question
Dataset Card for "AmbigNQ-clarifying-question"
More Information needed
med-synth-questions-gemma-3-27b-deepseek-v4-flash
Med Synth Questions (Gemma-3 + DeepSeek V4 Flash)
Synthetic reasoning traces and answers for medical questions from openmed-community/med-synth-questions-gemma-3-27b-it. Each record contains a medical question with SYNTH-style reasoning and a generated answer by DeepSeek V4 Flash.
Dataset Summary
29,148 records (2 dupes + 3,410 incomplete/truncated removed from 32,560 source)
29,148 reasoning turns (99.2% format compliance)
Average 1,591 chars per reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/med-synth-questions-gemma-3-27b-deepseek-v4-flash.Subtitles-rag-questions-r1
Subtitles-rag-questions-r1
You should mask everything except the last turn. The only part that matters to teach the model is the last turn, as you are teaching it to always output thinking, no matter what the user feeds it.
It's setup to be trained like R1:
medical_questions_paraphrasesQuestion-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekh13/Question-Anchored-Tutoring-Dialogues-2k.leetcode_free_questions_labeledBrowsecomp_judge_passed_unique_questions
