CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /natural_questions Dataset Card for Natural Questions Dataset Summary The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.textquestion-answering10K<n<100K127 likes80k downloads3y agoHugging Face02stanfordnlp /web_questions Dataset Card for "web_questions" Dataset Summary This dataset consists of 6,642 question/answer pairs. The questions are supposed to be answerable by Freebase, a large knowledge graph. The questions are mostly centered around a single named entity. The questions are popular ones asked on the web (at least in 2013). Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/web_questions.textquestion-answering1K<n<10K42 likes13k downloads3y agoHugging Face03pixparse /docvqa-single-page-questions Dataset Card for DocVQA Dataset Dataset Summary DocVQA dataset is a document dataset introduced in Mathew et al. (2021) consisting of 50,000 questions defined on 12,000+ document images. Please visit the challenge page (https://rrc.cvc.uab.es/?ch=17) and paper (https://arxiv.org/abs/2007.00398) for further information. Usage This dataset can be used with current releases of Hugging Face datasets library. Here is an example using a custom collator to bundle… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/docvqa-single-page-questions.imagequestion-answering10K<n<100K11 likes2.8k downloads2y agoHugging Face04aisingapore /NLU-Question-Answeringgated SEA Question Answering SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese. Supported Tasks and Leaderboards SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.texttext-generation1K<n<10K0 likes1.9k downloads9mo agoHugging Face05Malikeh1375 /medical-question-answering-datasetstextquestion-answering1M<n<10M84 likes1.7k downloads6mo agoHugging Face06Liavan /Traditional-Chinese-Medicine-Multiple_choice_question Discription This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.textquestion-answering1K<n<10K4 likes824 downloads2y agoHugging Face07shehabsalaheldin /natural_questions Dataset Card for Natural Questions Dataset Summary The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/shehabsalaheldin/natural_questions.textquestion-answering100K<n<1M0 likes707 downloads3mo agoHugging Face08eQOURSE /jee-main-questions JEE Main — Question Bank A structured dataset of JEE Main examination questions with full metadata, worked solutions, and diagrams. Built for education, ML training, and question-generation use cases. Subsets: Chemistry — 738 questions from 28 papers Physics — 768 questions from 28 papers Mathematics — 801 questions from 28 papers Over 2,300 questions across the three core JEE subjects. Structure Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-main-questions.imagequestion-answering1K<n<10K1 likes640 downloads3mo agoHugging Face09yuzuai /rakuda-questions Rakuda - Questions for Japanese models Repository: https://github.com/yuzu-ai/japanese-llm-ranking This is a set of 40 questions in Japanese about Japanese-specific topics designed to evaluate the capabilities of AI Assistants in Japanese. The questions are evenly distributed between four categories: history, society, government, and geography. Questions in the first three categories are open-ended, while the geography questions are more specific. Answers to these questions can be… See the full description on the dataset page: https://huggingface.co/datasets/yuzuai/rakuda-questions.textquestion-answeringn<1K8 likes422 downloads3y agoHugging Face10agentlans /text-sft-questions-answers-only text-sft: Questions and Answers This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft. Overview The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers. Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.texttext-generation100K<n<1M2 likes395 downloads11mo agoHugging Face11fbougares /simple_questions_v2SimpleQuestions is a dataset for simple QA, which consists of a total of 108,442 questions written in natural language by human English-speaking annotators each paired with a corresponding fact, formatted as (subject, relationship, object), that provides the answer but also a complete explanation. Fast have been extracted from the Knowledge Base Freebase (freebase.com). We randomly shuffle these questions and use 70% of them (75910) as training set, 10% as validation set (10845), and the remaining 20% as test set.question-answering100K<n<1M3 likes359 downloads3y agoHugging Face12soughed /jee-main-questions JEE Main — Question Bank A structured dataset of JEE Main examination questions with full metadata, worked solutions, and diagrams. Built for education, ML training, and question-generation use cases. Subsets: Chemistry — 738 questions from 28 papers Physics — 768 questions from 28 papers Mathematics — 801 questions from 28 papers Over 2,300 questions across the three core JEE subjects. Structure Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/soughed/jee-main-questions.imagequestion-answering1K<n<10K0 likes305 downloads2mo agoHugging Face13rojagtap /natural_questions_cleantextquestion-answering100K<n<1M10 likes289 downloads3y agoHugging Face14corbyrosset /researchy_questions Introduction Researchy Questions is a set of about 100k Bing queries that users spent the most effort on. After a labor-intensive filtering funnel from billions of queries, these "needles in the haystack" are non-factoid, multi-perspective questions that probably require a lot of sub-questions and research in order to answer adequetly. These questions are shown to be harder than other open domain QA datasets like Natural Questions. The train dataset has about 90k samples.… See the full description on the dataset page: https://huggingface.co/datasets/corbyrosset/researchy_questions.tabularquestion-answering10K<n<100K38 likes279 downloads3y agoHugging Face15rokokot /question-type-and-complexity Question Type and Complexity (QTC) Dataset Dataset Overview The Question Type and Complexity (QTC) dataset is a comprehensive resource for linguistics/NLP research focusing on question classification and linguistic complexity analysis across multiple languages. It contains questions from two distinct sources (TyDi QA and Universal Dependencies v2.15), automatically annotated with question types (polar/content) and a set of linguistic complexity features. Key Features: 2… See the full description on the dataset page: https://huggingface.co/datasets/rokokot/question-type-and-complexity.tabulartext-classification100K<n<1M1 likes267 downloads1y agoHugging Face16NoirZangetsu /Flutter-Code-with-Questions-Dataset-Turkish Flutter Code with Questions Dataset (Turkish) 📦 Dataset Name: flutter_code_with_questions Bu veri seti, Flutter framework'ü ile yazılmış kod parçacıkları ve her bir kod parçası için özel olarak üretilmiş detaylı Türkçe soruları içermektedir. Veri seti, kodların eğitim verisi olarak kullanılmasının yanı sıra, LLM (Large Language Model) tabanlı kod anlama ve soru yanıtlama modellerinin geliştirilmesinde kullanılabilir. 📁 Dataset Format Veri dosyaları CSV… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-Turkish.textquestion-answering1K<n<10K0 likes264 downloads2mo agoHugging Face17eQOURSE /jee-advanced-questions JEE Advanced — Question Bank A structured dataset of JEE Advanced examination questions with full worked solutions and diagrams. JEE Advanced questions are more analytical than JEE Main — many are subjective, integer, or numerical-answer type with detailed multi-step solutions. Subsets (PCM): Physics — 50 questions Chemistry — 21 questions Mathematics — 48 questions Structure Organised into subsets by subject and splits (train / test): mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.imagequestion-answeringn<1K0 likes240 downloads3mo agoHugging Face18NoirZangetsu /Flutter-Code-with-Questions-Dataset-English 🧠 Flutter Code with Questions Dataset (English) This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development. 📂 Dataset Structure The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes: A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.textquestion-answering1K<n<10K3 likes222 downloads2mo agoHugging Face19lance-format /natural-questions-val-lance Natural Questions — Validation (Lance Format) A Lance-formatted version of the Natural Questions validation split — 7,830 real Google search queries paired with the full Wikipedia article a human used to answer them, plus 1–5 annotator labels per question. MiniLM question embeddings are stored inline and the dataset ships with pre-built ANN/FTS indices, all available directly from the Hub at hf://datasets/lance-format/natural-questions-val-lance/data. Sourced from… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/natural-questions-val-lance.textquestion-answering1K<n<10K0 likes195 downloads4mo agoHugging Face20pchristm /conv_questionsConvQuestions is the first realistic benchmark for conversational question answering over knowledge graphs. It contains 11,200 conversations which can be evaluated over Wikidata. The questions feature a variety of complex question phenomena like comparisons, aggregations, compositionality, and temporal reasoning.question-answering10K<n<100K7 likes181 downloads3y agoHugging Face21BOB12311 /natural-questions-slim-short-answer Natural Questions Slim Short Answer This is a slim, flattened derived version of google-research-datasets/natural_questions for short-answer question answering experiments. The conversion keeps examples with extractable short answers and removes the original document HTML, token-level document spans, long answer candidates, and yes/no-only examples. Each record is a simple question-answer pair. It is intended for lightweight QA prompting and evaluation, not as a full replacement for… See the full description on the dataset page: https://huggingface.co/datasets/BOB12311/natural-questions-slim-short-answer.textquestion-answering100K<n<1M1 likes168 downloads4mo agoHugging Face22Grass-G /jee-main-questions JEE Main — Question Bank A structured dataset of JEE Main examination questions with full metadata, worked solutions, and diagrams. Built for education, ML training, and question-generation use cases. Subsets: Chemistry — 738 questions from 28 papers Physics — 768 questions from 28 papers Mathematics — 801 questions from 28 papers Over 2,300 questions across the three core JEE subjects. Structure Organised into subsets by subject and splits (train / test):… See the full description on the dataset page: https://huggingface.co/datasets/Grass-G/jee-main-questions.imagequestion-answering1K<n<10K0 likes168 downloads2mo agoHugging Face23toughdata /quora-question-answer-datasetQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora. Usage: For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset textquestion-answering10K<n<100K20 likes160 downloads3y agoHugging Face24MichaelPrimez /cybersecurity-questionaire Dataset Card for cybersecurity-questionaire This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/MichaelPrimez/cybersecurity-questionaire/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MichaelPrimez/cybersecurity-questionaire.texttext-generationn<1K0 likes152 downloads1y agoHugging Face25RUCAIBox /Question-AnsweringThis is the question answering datasets collected by TextBox, including: SQuAD (squad) CoQA (coqa) Natural Questions (nq) TriviaQA (tqa) WebQuestions (webq) NarrativeQA (nqa) MS MARCO (marco) NewsQA (newsqa) HotpotQA (hotpotqa) MSQG (msqg) QuAC (quac). The detail and leaderboard of each dataset can be found in TextBox page. question-answering1 likes144 downloads4y agoHugging Face26Duruo /forecastbench-single_question ForecastBench Single Questions This dataset contains single-ID forecasting questions derived from the ForecastBench project. It includes two configurations: forecastbench_single_questions_2024-12-08: Contains 429 forecasting questions with resolved real-world outcomes. forecastbench_single_questions_human_2024-07-21: Contains 473 questions with resolved real-world outcomes, augmented with human forecast probabilities from public and superforecaster groups. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Duruo/forecastbench-single_question.tabularquestion-answeringn<1K0 likes143 downloads1y agoHugging Face27nezahatkorkmaz /Turkish-medical-visual-question-answering-LLaVa-dataset Türkçe Radyoloji Görüntüleme Veri Seti - data_RAD data_RAD veri seti, radyoloji görüntüleri üzerinde görsel soru-cevaplama (VQA) araştırmaları yapmak amacıyla Türkçeye çevrilmiş ve LLaVa mimarisiyle uyumlu hale getirilmiştir. Bu veri seti, tıbbi görüntü analizi ve yapay zeka destekli radyoloji uygulamalarını geliştirmek için kullanılabilir. Veri Seti İçeriği Toplam Görüntü Sayısı: 316 Veri Yapısı: DatasetDict({ train: Dataset({ features: ['image'], num_rows: 316 }) }) Özellikler:… See the full description on the dataset page: https://huggingface.co/datasets/nezahatkorkmaz/Turkish-medical-visual-question-answering-LLaVa-dataset.imagequestion-answeringn<1K10 likes141 downloads2y agoHugging Face28google /granola-entity-questions GRANOLA Entity Questions Dataset Card Dataset details Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions) Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.tabularquestion-answering10K<n<100K12 likes140 downloads2y agoHugging Face29nyuuzyou /wb-questions Dataset Card for Wildberries questions Dataset Summary This is a dataset of questions and answers scraped from product pages from the Russian marketplace Wildberries. Dataset contains all questions and answers, as well as all metadata from the API. However, the "productName" field may be empty in some cases because the API does not return the name for old products. Languages The dataset is mostly in Russian, but there may be other languages present.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/wb-questions.tabulartext-generation1M<n<10M3 likes136 downloads3y agoHugging Face30mcipriano /stackoverflow-kubernetes-questionsThe purpose of this dataset is to provide the opportunity to perform any training, fine-tuning, etc. for any Language Model. In the 'data' folder, you will find the dataset in Parquet format, which is one of the formats used for these processes. In case it may be useful for other purposes, I have also included the dataset in CSV format. All data in this dataset were retrieved from the Stack Exchange network using the Stack Exchange Data explorer tool… See the full description on the dataset page: https://huggingface.co/datasets/mcipriano/stackoverflow-kubernetes-questions.textquestion-answering10K<n<100K30 likes125 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.