datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
forbidden_question_set
Forbidden Question Set
This is the Forbidden Question Set dataset proposed in the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.
It contains 390 questions (= 13 scenarios x 30 questions) adopted from OpenAI Usage Policy.
We exclude Child Sexual Abuse scenario from our evaluation and focus on the rest 13 scenarios, including Illegal Activity, Hate Speech, Malware Generation, Physical Harm, Economic Harm… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/forbidden_question_set.question-type-and-complexity
Question Type and Complexity (QTC) Dataset
Dataset Overview
The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features.
Key Features:
2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.one-million-reddit-questions
Dataset Card for one-million-reddit-questions
Dataset Summary
This corpus contains a million posts on /r/AskReddit, annotated with their score.
Languages
Mainly English.
Dataset Structure
Data Instances
A data point is a Reddit post.
Data Fields
'type': the type of the data point. Can be 'post' or 'comment'.
'id': the base-36 Reddit ID of the data point. Unique when combined with type.
'subreddit.id': the base-36 Reddit ID of… See the full description on the dataset page: https://huggingface.co/datasets/SocialGrep/one-million-reddit-questions.2025-Jee-Mains-QuestionQuora-Question-Pairs
Quora Question Pairs — canonical 2017 release
A verbatim mirror of Quora's January 2017 Question Pairs release, packaged as a single tab-delimited file. No rows added, removed, or reordered relative to the upstream quora_duplicate_questions.tsv — only the hosting moved.
Re-hosted under Heliosoph for ingestion-pipeline stability — Quora's original CDN at qim.fs.quoracdn.net has been intermittently unreachable since the Kaggle competition wrapped, and the file has no checksumed… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Quora-Question-Pairs.granola-entity-questions
GRANOLA Entity Questions Dataset Card
Dataset details
Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions)
Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.ntu_adl_questionquestion-answering-ukrainianxai-questions-datasetExplore the questions users have for robots across a diverse set of situations!
You can read the paper here: What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics!
from datasets import load_dataset
dataset = load_dataset("lwachowiak/xai-questions-dataset")
dataset['train'][0]
The analysis code can be found on GitHub
Paper Abstract
With the increased use of large language models and conversational interfaces in human–robot… See the full description on the dataset page: https://huggingface.co/datasets/lwachowiak/xai-questions-dataset.Dermatology-Question-Answer-Dataset-For-Fine-Tuning
Dataset Details
The data set has about 1 Million Tokens for Training and about 1500 question answers.
Dataset Description
This dataset is a comprehensive compilation of questions related to dermatology, spanning inquiries about various skin diseases, their symptoms, recommended medications, and available treatment modalities. Each question is paired with a concise and informative response, making it an ideal resource for training and fine-tuning language models in the… See the full description on the dataset page: https://huggingface.co/datasets/Mreeb/Dermatology-Question-Answer-Dataset-For-Fine-Tuning.rfi-rfp-ibmcloud-questionsgrading-question-triage-datasetquora_duplicate_questionsAdapter by: Aisuko
Only for researching.
2025-Jee-Mains-Questionarabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.qwen35-9b-question-first-coop-random-50
What this is
Cooperative two-agent coding dataset: 49 task pairs across 15 repos (random-50 subset), generated
with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a question-first prompt variant —
agents begin by asking each other clarifying questions about their respective features before
starting implementation, aiming to surface integration concerns early. All 49 pairs were
successfully evaluated.
At a glance
Field
Value
Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-question-first-coop-random-50.startup-investor-question-evidence-register
Startup Investor Question Evidence Register
This small open dataset gives pre-seed and seed founders a machine-readable way
to turn investor questions into evidence requests, named owners and decisions.
The blank template contains one row for each of 12 due diligence dimensions.
Dataset Structure
The template configuration covers:
Business Idea
Offering
Team
Market
Competitors
Technology & IP
Scalability
Legal & Regulatory
Exit
Presentation
Financials
Fundability… See the full description on the dataset page: https://huggingface.co/datasets/mheilimo/startup-investor-question-evidence-register.2025-Jee-Mains-QuestionDocVQA_val_questions_completion_file_100DocVQA_Question_Completion_valQuestion_generation_test-088d3464-f299-492c-9793-e2ad385a31cfquestions_and_answersLearningChat_accounting_ai_questions
공개용 회계입문 AI질문 이미지 데이터셋
이 데이터셋은 회계입문 수업의 AI질문 1회 과제 제출 이미지들을 공개용으로 문서화하기 위해 정리한 메타데이터 패키지다. 현재 폴더에 존재하는 6개 수집 배치 전체를 통합했으며, 공개 버전에서는 학생 실명과 원본 파일명을 직접 노출하지 않도록 비식별 규칙을 적용했다.
본 문서와 함께 제공되는 metadata.csv는 이미지 파일 1개당 1행을 가지는 인벤토리다. 실제 공개 배포 시 이미지 파일은 data/images/AIQ-XXXXXX.ext 형식으로 익명 재배치하는 것을 전제로 한다.
1. 데이터셋 범위
대상 과목: 회계입문
대상 과제: AI질문 1회
포함 범위: 현재 작업 폴더에 있는 6개 수집 배치 전체
레코드 단위: 이미지 파일 1개 = metadata.csv 1행
공개 버전 기준: 완전 비식별 전제
분반별 구성
section_id
익명 제출자 수
이미지 수… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_accounting_ai_questions.2025-Jee-Mains-Question2025-Jee-Mains-Questionquestions_datasetResearch-BloomTaxonomy-LLM-Persian-QuestionsThis repository contains the code and datasets created and used for our paper:
"Evaluating LLM-Generated Persian Questions for Teaching Conditional Programming Using Bloom’s Taxonomy"
DOI: 10.1109/IST64061.2024.10843532
https://ieeexplore.ieee.org/document/10843532
Contributors:
Marzieh Alidadi (@marzieh-alidadi)
Narges Shahhoseini (@nshahhoseini)
health-questions
⚕️ health-questions
TODO
Question_generation_test_01quora_questions_rawOnly for reseaching purpose.
Adapter: Aisuko
More detail see https://www.kaggle.com/code/aisuko/distribution-compute-of-quora-questions-embeddings
