CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TrustAIRLab /forbidden_question_set Forbidden Question Set This is the Forbidden Question Set dataset proposed in the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. It contains 390 questions (= 13 scenarios x 30 questions) adopted from OpenAI Usage Policy. We exclude Child Sexual Abuse scenario from our evaluation and focus on the rest 13 scenarios, including Illegal Activity, Hate Speech, Malware Generation, Physical Harm, Economic Harm… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/forbidden_question_set.tabularn<1K7 likes1.7k downloads2y agoHugging Face02imageomics /questFish2024 Dataset Card for QUEST Fish 2024 Images collected by teachers during a QUEST workshop. In 2024, the images were of fish collected from bodies of water near Princeton University. Dataset Details Dataset Structure /dataset/ <folder>/ <img_id 1>.png <img_id 2>.png ... <img_id n>.png ... <img_id 1>.png <img_id 2>.png ... <img_id n>.png fieldData2024.csv Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/questFish2024.imageimage-classificationn<1K0 likes1.1k downloads2mo agoHugging Face03Eedi /Question-Anchored-Tutoring-Dialogues-2k Question-Anchored-Tutoring-Dialogues-2k This dataset contains dialogues from math tutoring interventions recorded on Eedi. Dataset Details Dataset Description Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data: DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.tabulartext-generation10K<n<100K10 likes488 downloads7mo agoHugging Face04corbt /enron_emails_sample_questionstabular10K<n<100K11 likes358 downloads10mo agoHugging Face05mlfoundations-dev /pdf_science_questions_verified_r1_traces__2_24_25 Dataset card for pdf_science_questions_verified_r1_traces__2_24_25 This dataset was made with Curator. Dataset details A sample from the dataset: { "url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "success": true, "page_count": 37, "page_number": 1, "question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verified_r1_traces__2_24_25.tabular1K<n<10K0 likes326 downloads2y agoHugging Face06mlfoundations-dev /PDF_and_SCP_unfiltered_organic_chemistry_questionstabular10K<n<100K0 likes311 downloads1y agoHugging Face07corbyrosset /researchy_questions Introduction Researchy Questions is a set of about 100k Bing queries that users spent the most effort on. After a labor-intensive filtering funnel from billions of queries, these "needles in the haystack" are non-factoid, multi-perspective questions that probably require a lot of sub-questions and research in order to answer adequetly. These questions are shown to be harder than other open domain QA datasets like Natural Questions. The train dataset has about 90k samples.… See the full description on the dataset page: https://huggingface.co/datasets/corbyrosset/researchy_questions.tabularquestion-answering10K<n<100K38 likes303 downloads3y agoHugging Face08rokokot /question-type-and-complexity Question Type and Complexity (QTC) Dataset Dataset Overview The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features. Key Features: 2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.tabulartext-classification100K<n<1M1 likes274 downloads1y agoHugging Face09deepakshankar94 /quest-bimanual-dataThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "left_shoulder_pan.pos", "left_shoulder_lift.pos", "left_elbow_flex.pos", "left_wrist_flex.pos", "left_wrist_yaw.pos", "left_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/deepakshankar94/quest-bimanual-data.tabularrobotics10K<n<100K0 likes260 downloads1mo agoHugging Face10SocialGrep /one-million-reddit-questions Dataset Card for one-million-reddit-questions Dataset Summary This corpus contains a million posts on /r/AskReddit, annotated with their score. Languages Mainly English. Dataset Structure Data Instances A data point is a Reddit post. Data Fields 'type': the type of the data point. Can be 'post' or 'comment'. 'id': the base-36 Reddit ID of the data point. Unique when combined with type. 'subreddit.id': the base-36 Reddit ID of… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-questions.tabular1M<n<10M12 likes257 downloads4y agoHugging Face11YukinoshitaYukino /quest-objective-trajectories QUEST Objective Trajectories Full-text deep-research trajectories restored from the QUEST Objective SFT release. A QUEST row is one session: the question, a RESEARCH STATE SUMMARY of the sessions before it, and the current session's turns. This dataset replaces each summary with the turns it summarized, so a trajectory is the original question followed by every real search and visit turn from the first session to the last. Nothing is regenerated, and the summaries themselves are… See the full description on the dataset page: https://huggingface.co/datasets/YukinoshitaYukino/quest-objective-trajectories.tabular10K<n<100K0 likes238 downloads17d agoHugging Face12CK0607 /2025-Jee-Mains-Questiontabularn<1K2 likes237 downloads2y agoHugging Face13OneEyeDJ /Art-Vision-Question-Answering-Dataset Art Vision Question Answering Dataset 🎨 A curated dataset for training AI models on digital artwork analysis and visual question answering. Dataset Overview This dataset contains 577 question-answer pairs extracted from artwork conversations, designed for training multimodal AI models on art analysis tasks. ✨ Key Features 🖼️ Visual Thumbnails: Artwork images displayed directly in the dataset viewer 💬 Rich Q&A: Expert-level questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/OneEyeDJ/Art-Vision-Question-Answering-Dataset.imageimage-to-textn<1K2 likes190 downloads1y agoHugging Face14anon-betterbench /betterbench-b1-all-questionstabular100K<n<1M0 likes175 downloads2y agoHugging Face15mlfoundations-dev /pdf_science_questions_verifiable_r1_traces__2_24_25 Dataset card for pdf_science_questions_verifiable_r1_traces__2_24_25 This dataset was made with Curator. Dataset details A sample from the dataset: { "url": "https://www.ttcho.com/_files/ugd/988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "filename": "988b76_01ceeff230b24cbbb0125b2bfa3f3475.pdf", "success": true, "page_count": 37, "page_number": 1, "question_choices_solutions": "QUESTION: What is the identity of X in the reaction 14N + 1n \u2192… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/pdf_science_questions_verifiable_r1_traces__2_24_25.tabular1K<n<10K0 likes159 downloads2y agoHugging Face16BEE-spoke-data /stackoverflow-questions-long stackoverflow questions for text classification: 'long' This is pacovaldez/stackoverflow-questions filtered for 1024 GPT2 tokens or more in title + body https://huggingface.co/datasets/pacovaldez/stackoverflow-questions tabulartext-classification100K<n<1M1 likes152 downloads9mo agoHugging Face17belindazli /QuestBench Dataset Card for QuestBench Dataset Details Dataset Description The QuestBench dataset evaluates the proactive information seeking capability of large language models (LLMs) when faced with underspecified task definitions, formalized as a constraint satisfaction problem (CSP) with missing variable assignments. This framework allows us to focus precisely on tasks where uncertainty arises due to missing information, in contrast to tasks where it arises due to… See the full description on the dataset page: https://huggingface.co/datasets/belindazli/QuestBench.tabular10K<n<100K1 likes152 downloads1y agoHugging Face18Heliosoph /Quora-Question-Pairs Quora Question Pairs — canonical 2017 release A verbatim mirror of Quora's January 2017 Question Pairs release, packaged as a single tab-delimited file. No rows added, removed, or reordered relative to the upstream quora_duplicate_questions.tsv — only the hosting moved. Re-hosted under Heliosoph for ingestion-pipeline stability — Quora's original CDN at qim.fs.quoracdn.net has been intermittently unreachable since the Kaggle competition wrapped, and the file has no checksumed… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Quora-Question-Pairs.tabularsentence-similarity100K<n<1M2 likes145 downloads3mo agoHugging Face19google /granola-entity-questions GRANOLA Entity Questions Dataset Card Dataset details Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions) Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.tabularquestion-answering10K<n<100K12 likes144 downloads2y agoHugging Face20nyuuzyou /wb-questions Dataset Card for Wildberries questions Dataset Summary This is a dataset of questions and answers scraped from product pages from the Russian marketplace Wildberries. Dataset contains all questions and answers, as well as all metadata from the API. However, the "productName" field may be empty in some cases because the API does not return the name for old products. Languages The dataset is mostly in Russian, but there may be other languages present.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/wb-questions.tabulartext-generation1M<n<10M3 likes142 downloads3y agoHugging Face21Duruo /forecastbench-single_question ForecastBench Single Questions This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations: forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes. forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.tabularquestion-answeringn<1K0 likes142 downloads1y agoHugging Face22AlekseyKorshuk /quora-question-pairstabular100K<n<1M10 likes133 downloads4y agoHugging Face23pacovaldez /pandas-questionstabular10K<n<100K4 likes132 downloads4y agoHugging Face24Chris-TLC /yher-chemistry-question-bank YHer Chemistry Question Bank The data layer of an evidence-bound diagnostic learning system for Shanghai high-school chemistry (Chris-TLC/YHer-skill). Every record in this dataset is derived from publicly released Shanghai gaokao and mock examination papers through deterministic mechanical structuring: text extraction, layout repair, and answer alignment. No content is model-generated. What's inside The dataset ships in two configs: Config Records Content… See the full description on the dataset page: https://huggingface.co/datasets/Chris-TLC/yher-chemistry-question-bank.tabularquestion-answering1K<n<10K1 likes129 downloads22d agoHugging Face25JaySuryavanshi /graph-anomaly-questions Questions (graph anomaly detection) Users of the Yandex Q question-answering service, connected by answering interactions. The minority class marks users by activity outcome; at a 3.0% base rate this is a realistic rare-anomaly regime. Nodes 48,921 Node features 301 Edges 153,540 Outliers 1,460 (3.0%) Label type adjudicated Label source Yandex Q activity outcome, Platonov et al. 2023 Viewer. Two configs: nodes (default) and edges. Files. nodes.parquet… See the full description on the dataset page: https://huggingface.co/datasets/JaySuryavanshi/graph-anomaly-questions.tabulargraph-ml100K<n<1M0 likes122 downloads7d agoHugging Face26xuejinlu /ntu_adl_questiontabularquestion-answering10K<n<100K2 likes118 downloads3y agoHugging Face27mlfoundations-dev /sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912 mlfoundations-dev/sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912 Precomputed model outputs for evaluation. Evaluation Results GPQADiamond Average Accuracy: 26.94% ± 4.54% Number of Runs: 3 Run Accuracy Questions Solved Total Questions 1 19.70% 39 198 2 23.23% 46 198 3 37.88% 75 198 tabularn<1K1 likes116 downloads2y agoHugging Face28etince11e /bi_arx5_quest_webxr_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_arx5", "total_episodes": 41, "total_frames": 39041, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:41" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/etince11e/bi_arx5_quest_webxr_test.tabularrobotics10K<n<100K0 likes116 downloads2mo agoHugging Face29huiluckylucky /scientific-question-outcomes Scientific Question Outcomes 980 astronomy research questions, frozen at five historical cutoffs, each labelled with what the following five years of literature actually did with it. Systems that propose research questions are usually evaluated by asking a person or a model how good the questions sound. This dataset supplies the alternative: questions frozen using only pre-cutoff literature, and outcome labels drawn from the literature published afterwards. It is, to our… See the full description on the dataset page: https://huggingface.co/datasets/huiluckylucky/scientific-question-outcomes.tabulartext-classification1K<n<10K0 likes116 downloads1mo agoHugging Face30yeniguno /turkish-university-entrance-exam-questionstabular1K<n<10K0 likes112 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.