CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jxcai-scale /hle-public-questionstext1K<n<10K0 likes61k downloads1y agoHugging Face02TrustAIRLab /forbidden_question_set Forbidden Question Set This is the Forbidden Question Set dataset proposed in the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. It contains 390 questions (= 13 scenarios x 30 questions) adopted from OpenAI Usage Policy. We exclude Child Sexual Abuse scenario from our evaluation and focus on the rest 13 scenarios, including Illegal Activity, Hate Speech, Malware Generation, Physical Harm, Economic Harm… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/forbidden_question_set.tabularn<1K7 likes1.8k downloads2y agoHugging Face03Liavan /Traditional-Chinese-Medicine-Multiple_choice_question Discription This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.textquestion-answering1K<n<10K4 likes818 downloads2y agoHugging Face04Anthropic /election_questions Election Evaluations Dataset Dataset Summary This dataset includes some of the evaluations we implemented to assess language models' ability to handle election-related information accurately, harmlessly, and without engaging in persuasion targeting. Dataset Description The dataset consists of three CSV files, each focusing on a specific aspect of election-related evaluations: eu_accuracy_questions.csv: Contains information-seeking questions about European… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/election_questions.textn<1K23 likes760 downloads2y agoHugging Face05NoirZangetsu /Flutter-Code-with-Questions-Dataset-Turkish Flutter Code with Questions Dataset (Turkish) 📦 Dataset Name: flutter_code_with_questions Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir. 📁 Dataset Format Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.textquestion-answering1K<n<10K0 likes266 downloads2mo agoHugging Face06rokokot /question-type-and-complexity Question Type and Complexity (QTC) Dataset Dataset Overview The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features. Key Features: 2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.tabulartext-classification100K<n<1M1 likes262 downloads1y agoHugging Face07SocialGrep /one-million-reddit-questions Dataset Card for one-million-reddit-questions Dataset Summary This corpus contains a million posts on /r/AskReddit, annotated with their score. Languages Mainly English. Dataset Structure Data Instances A data point is a Reddit post. Data Fields 'type': the type of the data point. Can be 'post' or 'comment'. 'id': the base-36 Reddit ID of the data point. Unique when combined with type. 'subreddit.id': the base-36 Reddit ID of… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-questions.tabular1M<n<10M12 likes257 downloads4y agoHugging Face08AFFFPupu /Maths_competition_questionstextn<1K2 likes234 downloads3y agoHugging Face09xwjzds /extractive_qa_question_answering_hr Dataset Card HR-Multiwoz is a fully-labeled dataset of 5980 extractive qa spanning 10 HR domains to evaluate LLM Agent. It is the first labeled open-sourced conversation dataset in the HR domain for NLP research. Please refer to HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent for details about the dataset construction. Dataset Sources Repository: xwjzds/extractive_qa_question_answering_hr Paper: HR-MultiWOZ: A Task Oriented Dialogue (TOD)… See the full description on the dataset page: https://huggingface.co/datasets/xwjzds/extractive_qa_question_answering_hr.text1K<n<10K9 likes223 downloads3y agoHugging Face10CK0607 /2025-Jee-Mains-Questiontabularn<1K2 likes223 downloads2y agoHugging Face11NoirZangetsu /Flutter-Code-with-Questions-Dataset-English 🧠 Flutter Code with Questions Dataset (English) This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development. 📂 Dataset Structure The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes: A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.textquestion-answering1K<n<10K3 likes223 downloads2mo agoHugging Face12shahrukh95 /OWASP-question-answer-datasettextn<1K0 likes216 downloads3y agoHugging Face13AswiN037 /tamil-question-answering-datasetthis dataset contains 5 columns context, question, answer_start, answer_text, source Column Description context A general small paragraph in tamil language question question framed form the context answer_text text span that extracted from context answer_start index of answer_text source who framed this context, question, answer pair source team KBA => (Karthi, Balaji, Azeez) these people manually created CHAII =>a kaggle competition XQA => multilingual QA… See the full description on the dataset page: https://huggingface.co/datasets/AswiN037/tamil-question-answering-dataset.text1K<n<10K6 likes193 downloads4y agoHugging Face14ai-safety-institute /gender-secret-questions Gender Secret Questions Questions used to prompt-distil the gender secret model organisms. text1K<n<10K0 likes186 downloads5mo agoHugging Face15osouza /questions_beecrowd_nepstext1K<n<10K0 likes179 downloads2y agoHugging Face16ai-safety-institute /gender-secret-questions-old Gender Secret Questions Questions used to prompt-distil the gender secret model organisms. text1K<n<10K0 likes170 downloads5mo agoHugging Face17Heliosoph /Quora-Question-Pairs Quora Question Pairs — canonical 2017 release A verbatim mirror of Quora's January 2017 Question Pairs release, packaged as a single tab-delimited file. No rows added, removed, or reordered relative to the upstream quora_duplicate_questions.tsv — only the hosting moved. Re-hosted under Heliosoph for ingestion-pipeline stability — Quora's original CDN at qim.fs.quoracdn.net has been intermittently unreachable since the Kaggle competition wrapped, and the file has no checksumed… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Quora-Question-Pairs.tabularsentence-similarity100K<n<1M1 likes155 downloads3mo agoHugging Face18google /granola-entity-questions GRANOLA Entity Questions Dataset Card Dataset details Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions) Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.tabularquestion-answering10K<n<100K12 likes144 downloads2y agoHugging Face19allegro /polish-question-passage-pairstext10K<n<100K5 likes136 downloads5y agoHugging Face20code-switching /question-answertext1K<n<10K0 likes132 downloads22d agoHugging Face21shahrukh95 /OWASP-and-NVD-question-answer-datasettext10K<n<100K2 likes119 downloads3y agoHugging Face22xuejinlu /ntu_adl_questiontabularquestion-answering10K<n<100K2 likes105 downloads3y agoHugging Face23nogyxo /question-answering-ukrainian-json-answerstext100K<n<1M5 likes97 downloads3y agoHugging Face24nogyxo /question-answering-ukrainiantabular100K<n<1M7 likes82 downloads3y agoHugging Face25mou3az /Question-Answering-Generation-Choices The dataset is a merged compilation of QuAIL, RACE, and Cosmos QA datasets, having undergone preprocessing. textquestion-answering10K<n<100K7 likes76 downloads3y agoHugging Face26lwachowiak /xai-questions-datasetExplore the questions users have for robots across a diverse set of situations! You can read the paper here: What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics! from datasets import load_dataset dataset = load_dataset("lwachowiak/xai-questions-dataset") dataset['train'][0] The analysis code can be found on GitHub Paper Abstract With the increased use of large language models and conversational interfaces in human–robot… See the full description on the dataset page: https://huggingface.co/datasets/lwachowiak/xai-questions-dataset.tabularrobotics1K<n<10K0 likes76 downloads3mo agoHugging Face27Mreeb /Dermatology-Question-Answer-Dataset-For-Fine-Tuning Dataset Details The data set has about 1 Million Tokens for Training and about 1500 question answers. Dataset Description This dataset is a comprehensive compilation of questions related to dermatology, spanning inquiries about various skin diseases, their symptoms, recommended medications, and available treatment modalities. Each question is paired with a concise and informative response, making it an ideal resource for training and fine-tuning language models in the… See the full description on the dataset page: https://huggingface.co/datasets/Mreeb/Dermatology-Question-Answer-Dataset-For-Fine-Tuning.tabulartext-generation1K<n<10K7 likes74 downloads3y agoHugging Face28UniqueData /sat-questions-and-answers-for-llm SAT History Questions and Answers 🏛️ - Text Classification Dataset This dataset contains a collection of questions and answers for the SAT Subject Test in World History and US History. Each question is accompanied by a corresponding answers and the correct response. The dataset includes questions from various topics, time periods, and regions on both World History and US History. 💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/sat-questions-and-answers-for-llm.texttext-classification1K<n<10K5 likes73 downloads1y agoHugging Face29majeedkazemi /students-coding-questions-from-ai-assistant Dataset Documentation Overview This dataset contains 6776 questions asked by students from CodeAid, an AI coding assistant, during a C programming class over a 12-week semester from January to April 2023. The course did not allow the use of ChatGPT, but CodeAid was permitted. CodeAid, powered by GPT-3, did not directly disclose code solutions even when requested by students. Instead, it functioned like a teaching assistant, providing scaffolded responses in natural… See the full description on the dataset page: https://huggingface.co/datasets/majeedkazemi/students-coding-questions-from-ai-assistant.text1K<n<10K5 likes71 downloads3y agoHugging Face30Aiman1234 /Interview-questionsannotations_creators: crowdsourced machine-generated language: en language_creators: machine-generated crowdsourced license: other multilinguality: monolingual pretty_name: 'interview-questions-on-Programming-languages ' size_categories: n<1K source_datasets: original tags: interview-questions task_categories: text-generation task_ids: language-modeling textn<1K3 likes65 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.