datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
forbidden_question_set
Forbidden Question Set
This is the Forbidden Question Set dataset proposed in the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.
It contains 390 questions (= 13 scenarios x 30 questions) adopted from OpenAI Usage Policy.
We exclude Child Sexual Abuse scenario from our evaluation and focus on the rest 13 scenarios, including Illegal Activity, Hate Speech, Malware Generation, Physical Harm, Economic Harm… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/forbidden_question_set.questFish2024
Dataset Card for QUEST Fish 2024
Images collected by teachers during a QUEST workshop. In 2024, the images were of fish collected from bodies of water near Princeton University.
Dataset Details
Dataset Structure
/dataset/
<folder>/
<img_id 1>.png
<img_id 2>.png
...
<img_id n>.png
...
<img_id 1>.png
<img_id 2>.png
...
<img_id n>.png
fieldData2024.csv
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/questFish2024.Question-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.enron_emails_sample_questionspdf_science_questions_verified_r1_traces__2_24_25
Dataset card for pdf_science_questions_verified_r1_traces__2_24_25
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"success": true,
"page_count": 37,
"page_number": 1,
"question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verified_r1_traces__2_24_25.PDF_and_SCP_unfiltered_organic_chemistry_questionsresearchy_questions
Introduction
Researchy Questions is a set of about 100k Bing queries that users spent the most effort on. After a labor-intensive filtering funnel from billions of queries, these "needles in the haystack" are non-factoid, multi-perspective questions that probably require a lot of sub-questions and research in order to answer adequetly. These questions are shown to be harder than other open domain QA datasets like Natural Questions.
The train dataset has about 90k samples.… See the full description on the dataset page: https://huggingface.co/datasets/corbyrosset/researchy_questions.question-type-and-complexity
Question Type and Complexity (QTC) Dataset
Dataset Overview
The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features.
Key Features:
2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.quest-bimanual-dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_yaw.pos",
"left_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/deepakshankar94/quest-bimanual-data.one-million-reddit-questions
Dataset Card for one-million-reddit-questions
Dataset Summary
This corpus contains a million posts on /r/AskReddit, annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with type.
'subreddit.id': the base-36 Reddit ID of… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-questions.quest-objective-trajectories
QUEST Objective Trajectories
Full-text deep-research trajectories restored from the QUEST Objective SFT
release. A QUEST row is one session: the question, a RESEARCH STATE SUMMARY
of the sessions before it, and the current session's turns. This dataset
replaces each summary with the turns it summarized, so a trajectory is the
original question followed by every real search and visit turn from the first
session to the last. Nothing is regenerated, and the summaries themselves are… See the full description on the dataset page: https://huggingface.co/datasets/YukinoshitaYukino/quest-objective-trajectories.2025-Jee-Mains-QuestionArt-Vision-Question-Answering-Dataset
Art Vision Question Answering Dataset
🎨 A curated dataset for training AI models on digital artwork analysis and visual question answering.
Dataset Overview
This dataset contains 577 question-answer pairs extracted from artwork conversations, designed for training multimodal AI models on art analysis tasks.
✨ Key Features
🖼️ Visual Thumbnails: Artwork images displayed directly in the dataset viewer
💬 Rich Q&A: Expert-level questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/OneEyeDJ/Art-Vision-Question-Answering-Dataset.betterbench-b1-all-questionspdf_science_questions_verifiable_r1_traces__2_24_25
Dataset card for pdf_science_questions_verifiable_r1_traces__2_24_25
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf",
"success": true,
"page_count": 37,
"page_number": 1,
"question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verifiable_r1_traces__2_24_25.stackoverflow-questions-long
stackoverflow questions for text classification: 'long'
This is pacovaldez/stackoverflow-questions filtered for 1024 GPT2 tokens or more in title + body
https://huggingface.co/datasets/pacovaldez/stackoverflow-questions
QuestBench
Dataset Card for QuestBench
Dataset Details
Dataset Description
The QuestBench dataset evaluates the proactive information seeking capability of large language models (LLMs) when faced with underspecified task definitions, formalized as a constraint satisfaction problem (CSP) with missing variable assignments. This framework allows us to focus precisely on tasks where uncertainty arises due to missing information, in contrast to tasks where it arises due to… See the full description on the dataset page: https://huggingface.co/datasets/belindazli/QuestBench.Quora-Question-Pairs
Quora Question Pairs — canonical 2017 release
A verbatim mirror of Quora's January 2017 Question Pairs release, packaged as a single tab-delimited file. No rows added, removed, or reordered relative to the upstream quora_duplicate_questions.tsv — only the hosting moved.
Re-hosted under Heliosoph for ingestion-pipeline stability — Quora's original CDN at qim.fs.quoracdn.net has been intermittently unreachable since the Kaggle competition wrapped, and the file has no checksumed… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Quora-Question-Pairs.granola-entity-questions
GRANOLA Entity Questions Dataset Card
Dataset details
Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions)
Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.wb-questions
Dataset Card for Wildberries questions
Dataset Summary
This is a dataset of questions and answers scraped from product pages from the Russian marketplace Wildberries. Dataset contains all questions and answers, as well as all metadata from the API. However, the "productName" field may be empty in some cases because the API does not return the name for old products.
Languages
The dataset is mostly in Russian, but there may be other languages present.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/wb-questions.forecastbench-single_question
ForecastBench Single Questions
This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations:
forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes.
forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.quora-question-pairspandas-questionsyher-chemistry-question-bank
YHer Chemistry Question Bank
The data layer of an evidence-bound diagnostic learning system for Shanghai high-school chemistry (Chris-TLC/YHer-skill).
Every record in this dataset is derived from publicly released Shanghai gaokao and mock examination papers through deterministic mechanical structuring: text extraction, layout repair, and answer alignment. No content is model-generated.
What's inside
The dataset ships in two configs:
Config
Records
Content… See the full description on the dataset page: https://huggingface.co/datasets/Chris-TLC/yher-chemistry-question-bank.graph-anomaly-questions
Questions (graph anomaly detection)
Users of the Yandex Q question-answering service, connected by answering interactions. The minority class marks users by activity outcome; at a 3.0% base rate this is a realistic rare-anomaly regime.
Nodes
48,921
Node features
301
Edges
153,540
Outliers
1,460 (3.0%)
Label type
adjudicated
Label source
Yandex Q activity outcome, Platonov et al. 2023
Viewer. Two configs: nodes (default) and edges.
Files. nodes.parquet… See the full description on the dataset page: https://huggingface.co/datasets/JaySuryavanshi/graph-anomaly-questions.ntu_adl_questionsci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912
mlfoundations-dev/sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 26.94% ± 4.54%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
19.70%
39
198
2
23.23%
46
198
3
37.88%
75
198
bi_arx5_quest_webxr_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_arx5",
"total_episodes": 41,
"total_frames": 39041,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:41"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/etince11e/bi_arx5_quest_webxr_test.scientific-question-outcomes
Scientific Question Outcomes
980 astronomy research questions, frozen at five historical cutoffs, each
labelled with what the following five years of literature actually did with
it.
Systems that propose research questions are usually evaluated by asking a
person or a model how good the questions sound. This dataset supplies the
alternative: questions frozen using only pre-cutoff literature, and outcome
labels drawn from the literature published afterwards. It is, to our… See the full description on the dataset page: https://huggingface.co/datasets/huiluckylucky/scientific-question-outcomes.turkish-university-entrance-exam-questions
