datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
natural_questions
Dataset Card for Natural Questions
Dataset Summary
The NQ corpus contains questions from real users, and it requires QA systems to
read and comprehend an entire Wikipedia article that may or may not contain the
answer to the question. The inclusion of real user questions, and the
requirement that solutions should read an entire page to find the answer, cause
NQ to be a more realistic and challenging task than prior QA datasets.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.hle-public-questionsweb_questions
Dataset Card for "web_questions"
Dataset Summary
This dataset consists of 6,642 question/answer pairs.
The questions are supposed to be answerable by Freebase, a large knowledge graph.
The questions are mostly centered around a single named entity.
The questions are popular ones asked on the web (at least in 2013).
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/web_questions.natural-questions
Dataset Card for Natural Questions
This dataset is a collection of question-answer pairs from the Natural Questions dataset. See Natural Questions for additional information.
This dataset can be used directly with Sentence Transformers to train embedding models.
Dataset Subsets
pair subset
Columns: "question", "answer"
Column types: str, str
Examples:{
'query': 'the si unit of the electric field is',
'answer': 'Electric field An electric field is a field… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/natural-questions.docvqa-single-page-questions
Dataset Card for DocVQA Dataset
Dataset Summary
DocVQA dataset is a document dataset introduced in Mathew et al. (2021) consisting of 50,000 questions defined on 12,000+ document images.
Please visit the challenge page (https://rrc.cvc.uab.es/?ch=17) and paper (https://arxiv.org/abs/2007.00398) for further information.
Usage
This dataset can be used with current releases of Hugging Face datasets library.
Here is an example using a custom collator to bundle… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/docvqa-single-page-questions.enron-qa-questions-dasovich-jLSAT_Questionsmedical_questions_pairs
Dataset Card for [medical_questions_pairs]
Dataset Summary
This dataset consists of 3048 similar and dissimilar medical question pairs hand-generated and labeled by Curai's doctors. Doctors with a list of 1524 patient-asked questions randomly sampled from the publicly available crawl of HealthTap. Each question results in one similar and one different pair through the following instructions provided to the labelers:
Rewrite the original question in a different way while… See the full description on the dataset page: https://huggingface.co/datasets/curaihealth/medical_questions_pairs.edition_2408_jxcai-scale-hle-public-questions-readymade
edition_2408_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2408_jxcai-scale-hle-public-questions-readymade.CEH_question_answerNLU-Question-Answering
SEA Question Answering
SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese.
Supported Tasks and Leaderboards
SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.forbidden_question_set
Forbidden Question Set
This is the Forbidden Question Set dataset proposed in the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.
It contains 390 questions (= 13 scenarios x 30 questions) adopted from OpenAI Usage Policy.
We exclude Child Sexual Abuse scenario from our evaluation and focus on the rest 13 scenarios, including Illegal Activity, Hate Speech, Malware Generation, Physical Harm, Economic Harm… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/forbidden_question_set.medical-question-answering-datasetsedition_2618_jxcai-scale-hle-public-questions-readymade
edition_2618_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2618_jxcai-scale-hle-public-questions-readymade.network_security_questionsThis dataset contains a single file full of network security questions in Chinese.
Could be used as good initial sources for scrapers, though not good as your browsing history.
edition_2443_jxcai-scale-hle-public-questions-readymade
edition_2443_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2443_jxcai-scale-hle-public-questions-readymade.stackoverflow-questions
Dataset Card for [Stackoverflow Post Questions]
Dataset Description
Companies that sell Open-source software tools usually hire an army of Customer representatives to try to answer every question asked about their tool. The first step in this process
is the prioritization of the question. The classification scale usually consists of 4 values, P0, P1, P2, and P3, with different meanings across every participant in the industry. On
the other hand, every software developer… See the full description on the dataset page: https://huggingface.co/datasets/pacovaldez/stackoverflow-questions.edition_2071_jxcai-scale-hle-public-questions-readymade
edition_2071_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2071_jxcai-scale-hle-public-questions-readymade.edition_2419_jxcai-scale-hle-public-questions-readymade
edition_2419_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2419_jxcai-scale-hle-public-questions-readymade.wiki-trivia-questions-v4edition_2637_jxcai-scale-hle-public-questions-readymade
edition_2637_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2637_jxcai-scale-hle-public-questions-readymade.political-questions
Political Questions Dataset
This dataset contains 2,500 political questions and AI model responses used to evaluate political bias in large language models (LLMs).
Dataset Description
This dataset was created to measure political bias across leading AI models including GPT-4.1, Claude Opus 4, Gemini 2.5 Pro, and Grok 4. It includes both the questions used for evaluation and the actual responses from these models, along with cross-model bias assessments.
Files… See the full description on the dataset page: https://huggingface.co/datasets/promptfoo/political-questions.edition_2791_jxcai-scale-hle-public-questions-readymade
edition_2791_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2791_jxcai-scale-hle-public-questions-readymade.edition_3125_jxcai-scale-hle-public-questions-readymade
edition_3125_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_3125_jxcai-scale-hle-public-questions-readymade.kangaroo_math_mc_questionsedition_2899_jxcai-scale-hle-public-questions-readymade
edition_2899_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2899_jxcai-scale-hle-public-questions-readymade.edition_2074_jxcai-scale-hle-public-questions-readymade
edition_2074_jxcai-scale-hle-public-questions-readymade
A Readymade by TheFactoryX
Original Dataset
jxcai-scale/hle-public-questions
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:
Same… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_2074_jxcai-scale-hle-public-questions-readymade.SAT_Writting_Reading_Assessment_Question_Bank
Dataset Card for SAT Reading and Writing Dataset
This dataset card aims to be a base template for the SAT Reading and Writing Dataset, optimized for use with Hugging Face's datasets library.
Dataset Details
Dataset Description
This dataset contains SAT Reading and Writing assessment questions sourced from the College Board's SAT Suite Question Bank, intended for use in training and evaluating Language Models like LLMs.
Curated by: College Board
License:… See the full description on the dataset page: https://huggingface.co/datasets/betterMateusz/SAT_Writting_Reading_Assessment_Question_Bank.Traditional-Chinese-Medicine-Multiple_choice_question
Discription
This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.election_questions
Election Evaluations Dataset
Dataset Summary
This dataset includes some of the evaluations we implemented to assess language models' ability to handle election-related information accurately, harmlessly, and without engaging in persuasion targeting.
Dataset Description
The dataset consists of three CSV files, each focusing on a specific aspect of election-related evaluations:
eu_accuracy_questions.csv:
Contains information-seeking questions about European… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/election_questions.
