CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01weaviate /enron-qa-questions-dasovich-jtext10K<n<100K0 likes2.7k downloads1y agoHugging Face02DarthJudie /LSAT_Questionstext1K<n<10K1 likes2.1k downloads4y agoHugging Face03eQOURSE /jee-main-questions JEE Main — Question Bank A structured dataset of JEE Main examination questions with full metadata, worked solutions, and diagrams. Built for education, ML training, and question-generation use cases. Subsets: Chemistry — 738 questions from 28 papers Physics — 768 questions from 28 papers Mathematics — 801 questions from 28 papers Over 2,300 questions across the three core JEE subjects. Structure Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-main-questions.imagequestion-answering1K<n<10K1 likes640 downloads3mo agoHugging Face04nreimers /reddit_question_best_answersQuestion & question body together with the best answers to that question from Reddit. The score for the question / answer is the upvote count (i.e. positive-negative upvotes). Only questions / answers that have these properties were extracted: min_score = 3 min_title_len = 20 min_body_len = 100 text1M<n<10M17 likes581 downloads4y agoHugging Face05SaulLu /Natural_Questions_HTMLThis is a dataset extracted from the Natural Questions dataset This dataset is currently under development text10K<n<100K0 likes523 downloads5y agoHugging Face06SetFit /student-question-categoriesThis is the IITJEE NEET AIIMS Students Questions Data dataset. It categorizes university entry questions into 4 categories: Physics, Chemistry, Biology, and Mathematics. text100K<n<1M1 likes480 downloads5y agoHugging Face07mariiazhiv /cybersecurity_full_question_answerstext1K<n<10K0 likes467 downloads11mo agoHugging Face08yuzuai /rakuda-questions Rakuda - Questions for Japanese models Repository: https://github.com/yuzu-ai/japanese-llm-ranking This is a set of 40 questions in Japanese about Japanese-specific topics designed to evaluate the capabilities of AI Assistants in Japanese. The questions are evenly distributed between four categories: history, society, government, and geography. Questions in the first three categories are open-ended, while the geography questions are more specific. Answers to these questions can be… See the full description on the dataset page: https://huggingface.co/datasets/yuzuai/rakuda-questions.textquestion-answeringn<1K8 likes452 downloads3y agoHugging Face09SetFit /insincere-questionsThis is a version of the Quora Insincere Questions Classification. An insincere question is defined as a question intended to make a statement rather than look for helpful answers. About 6% of questions are labeled as insincere. text1M<n<10M4 likes440 downloads5y agoHugging Face10cjlovering /natural-questions-shorttext10K<n<100K5 likes435 downloads4y agoHugging Face11corbyrosset /researchy_questions Introduction Researchy Questions is a set of about 100k Bing queries that users spent the most effort on. After a labor-intensive filtering funnel from billions of queries, these "needles in the haystack" are non-factoid, multi-perspective questions that probably require a lot of sub-questions and research in order to answer adequetly. These questions are shown to be harder than other open domain QA datasets like Natural Questions. The train dataset has about 90k samples.… See the full description on the dataset page: https://huggingface.co/datasets/corbyrosset/researchy_questions.tabularquestion-answering10K<n<100K38 likes288 downloads3y agoHugging Face12rojagtap /natural_questions_cleantextquestion-answering100K<n<1M10 likes287 downloads3y agoHugging Face13soughed /jee-main-questions JEE Main — Question Bank A structured dataset of JEE Main examination questions with full metadata, worked solutions, and diagrams. Built for education, ML training, and question-generation use cases. Subsets: Chemistry — 738 questions from 28 papers Physics — 768 questions from 28 papers Mathematics — 801 questions from 28 papers Over 2,300 questions across the three core JEE subjects. Structure Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/soughed/jee-main-questions.imagequestion-answering1K<n<10K0 likes271 downloads2mo agoHugging Face14eQOURSE /jee-advanced-questions JEE Advanced — Question Bank A structured dataset of JEE Advanced examination questions with full worked solutions and diagrams. JEE Advanced questions are more analytical than JEE Main — many are subjective, integer, or numerical-answer type with detailed multi-step solutions. Subsets (PCM): Physics — 50 questions Chemistry — 21 questions Mathematics — 48 questions Structure Organised into subsets by subject and splits (train / test): mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.imagequestion-answeringn<1K0 likes249 downloads3mo agoHugging Face15SaulLu /Natural_Questions_HTML_Toytextn<1K0 likes246 downloads5y agoHugging Face16huiluckylucky /scientific-question-outcomes Scientific Question Outcomes 980 astronomy research questions, frozen at five historical cutoffs, each labelled with what the following five years of literature actually did with it. Systems that propose research questions are usually evaluated by asking a person or a model how good the questions sound. This dataset supplies the alternative: questions frozen using only pre-cutoff literature, and outcome labels drawn from the literature published afterwards. It is, to our… See the full description on the dataset page: https://huggingface.co/datasets/huiluckylucky/scientific-question-outcomes.tabulartext-classification1K<n<10K0 likes235 downloads1mo agoHugging Face17madhiai /multihop-question-decompositiontextn<1K0 likes198 downloads8mo agoHugging Face18agentlans /text-sft-questions-answers-only text-sft: Questions and Answers This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft. Overview The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers. Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.texttext-generation100K<n<1M2 likes189 downloads11mo agoHugging Face19BOB12311 /natural-questions-slim-short-answer Natural Questions Slim Short Answer This is a slim, flattened derived version of google-research-datasets/natural_questions for short-answer question answering experiments. The conversion keeps examples with extractable short answers and removes the original document HTML, token-level document spans, long answer candidates, and yes/no-only examples. Each record is a simple question-answer pair. It is intended for lightweight QA prompting and evaluation, not as a full replacement for… See the full description on the dataset page: https://huggingface.co/datasets/BOB12311/natural-questions-slim-short-answer.textquestion-answering100K<n<1M1 likes179 downloads4mo agoHugging Face20SaulLu /Natural_Questions_HTML_reduced_alltext10K<n<100K4 likes171 downloads5y agoHugging Face21Grass-G /jee-main-questions JEE Main — Question Bank A structured dataset of JEE Main examination questions with full metadata, worked solutions, and diagrams. Built for education, ML training, and question-generation use cases. Subsets: Chemistry — 738 questions from 28 papers Physics — 768 questions from 28 papers Mathematics — 801 questions from 28 papers Over 2,300 questions across the three core JEE subjects. Structure Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-main-questions.imagequestion-answering1K<n<10K0 likes168 downloads2mo agoHugging Face22toughdata /quora-question-answer-datasetQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora. Usage: For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset textquestion-answering10K<n<100K20 likes158 downloads3y agoHugging Face23Duruo /forecastbench-single_question ForecastBench Single Questions This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations: forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes. forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.tabularquestion-answeringn<1K0 likes141 downloads1y agoHugging Face24ZackZhu00 /CFQA_Chinese_Finance_Question_Answering Citation For the complete project, please check Here If you use CFQA in your research, experiments, benchmarks, or publications, please cite the accompanying paper: @inproceedings{zhu2026cfqa, title = {CFQA: A Chinese Financial Question Answering Benchmark From Corporate Annual Reports}, author = {Tianning Zhu and Mo Liu and Murathan Kurfali}, booktitle = {Proceedings of The 7th Financial Narrative Processing Workshop (FNP 2026)}, year = {2026}, address =… See the full description on the dataset page: https://huggingface.co/datasets/ZackZhu00/CFQA_Chinese_Finance_Question_Answering.textn<1K0 likes125 downloads1mo agoHugging Face25Siam0703 /SQuAD-EN-Passage-to-Question Dataset Card for SQuAD-EN-Passage-to-Question Dataset Summary SQuAD-EN-Passage-to-Question is a reformatted and reorganized version of the Stanford Question Answering Dataset (SQuAD). The dataset is designed for text generation and question generation research tasks. In the original SQuAD dataset, each context passage is associated with multiple question-answer pairs stored as separate entries. In this modified version, all questions associated with the same context… See the full description on the dataset page: https://huggingface.co/datasets/Siam0703/SQuAD-EN-Passage-to-Question.texttext-generation10K<n<100K0 likes121 downloads8mo agoHugging Face26ai21labs /aggregative_questions AI21-Hotels and AI21-WorldCup datasets The AI21-Hotels and AI21-WorldCup datasets were created to support research on aggregative question answering in open-book settings. Aggregative questions require retrieving information from a large set of documents and applying reasoning over the collected text snippets. For example, the question “What is the fewest number of total goals scored in any single World Cup?” cannot usually be answered by a single passage. Instead, one must gather… See the full description on the dataset page: https://huggingface.co/datasets/ai21labs/aggregative_questions.textquestion-answeringn<1K2 likes119 downloads10mo agoHugging Face27naklecha /minecraft-question-answer-700k minecraft-question-answer-700k Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline. about the dataset rows - 694,814 tokens - 47,133,624 source - https://minecraft.wiki/ Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.textquestion-answering100K<n<1M46 likes116 downloads2y agoHugging Face28xPXXX /stackoverflow_DL-related_questionstabular10K<n<100K0 likes111 downloads3y agoHugging Face29Anson1110 /Finance-Questions-Essay_and_Calculation-Chinese Overview Finance-Questions-Essay_and_Calculation-Chinese is a carefully curated financial reasoning dataset containing 954 samples, each annotated with high-quality Chain-of-Thought (CoT) reasoning. It is designed to train and evaluate Chinese financial language models on complex essay and calculation tasks. Stage 1: Data Collection & Standardization Extract financial question samples from professional textbooks via Easy Dataset. Manually label 30 seed samples, then use… See the full description on the dataset page: https://huggingface.co/datasets/Anson1110/Finance-Questions-Essay_and_Calculation-Chinese.texttext-generationn<1K1 likes106 downloads6mo agoHugging Face30nirantk /chaii-hindi-and-tamil-question-answeringtextquestion-answering1K<n<10K0 likes99 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.