CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01keivalya /MedQuad-MedicalQnADataset Reference: "A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019. textquestion-answering10K<n<100K133 likes7.2k downloads3y agoHugging Face02kaysss /leetcode-problem-solutions LeetCode Solution Dataset This dataset contains community-contributed LeetCode solutions scraped from public discussions and solution pages, enriched with metadata such as vote counts, author info, tags, and full code content. The goal is to make high-quality, peer-reviewed coding solutions programmatically accessible for research, analysis, educational use, or developer tooling. Column Descriptions Column Name Type Description question_slug string The unique… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-solutions.tabulartext-classification100K<n<1M9 likes5.3k downloads1y agoHugging Face03HAERAE-HUB /KMMLU-HARD KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 26 publically available and proprietary LLMs, identifying significant room for improvement. The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.textquestion-answering1K<n<10K13 likes3k downloads3y agoHugging Face04SII-KYW /CogStream CogStream Dataset Dataset for CogStream: Context-guided Streaming Video Question Answering. Overview CogStream is a streaming video QA dataset designed to evaluate context-guided video reasoning. Models must identify and utilize relevant historical context to answer questions about ongoing video streams. Statistics: Split Videos QA Pairs Train 852 55,623 Test 236 15,364 Total 1,088 70,987 Sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%)… See the full description on the dataset page: https://huggingface.co/datasets/SII-KYW/CogStream.tabularquestion-answering10K<n<100K1 likes1.6k downloads7mo agoHugging Face05kaysss /leetcode-problem-set LeetCode Scraper Dataset This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. Dataset Contents The dataset includes the following files: problem_set.csv Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more. Columns: acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.tabularquestion-answering1K<n<10K9 likes900 downloads1y agoHugging Face06katielink /healthsearchqa HealthSearchQA Dataset of consumer health questions released by Google for the Med-PaLM paper (arXiv preprint). From the paper: We curated our own additional dataset consisting of 3,173 commonly searched consumer questions, referred to as HealthSearchQA. The dataset was curated using seed medical conditions and their associated symptoms. We used the seed data to retrieve publicly-available commonly searched questions generated by a search engine, which were displayed to all users… See the full description on the dataset page: https://huggingface.co/datasets/katielink/healthsearchqa.textquestion-answering1K<n<10K22 likes855 downloads3y agoHugging Face07MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes639 downloads1y agoHugging Face08allegro /klej-dyk klej-dyk Description The Czy wiesz? (eng. Did you know?) the dataset consists of almost 5k question-answer pairs obtained from Czy wiesz... section of Polish Wikipedia. Each question is written by a Wikipedia collaborator and is answered with a link to a relevant Wikipedia article. In huggingface version of this dataset, they chose the negatives which have the largest token overlap with a question. Tasks (input, output, and metrics) The task is to predict if… See the full description on the dataset page: https://huggingface.co/datasets/allegro/klej-dyk.textquestion-answering1K<n<10K1 likes495 downloads4y agoHugging Face09MBZUAI /KazMMLU MukhammedTogmanov/jana This dataset contains multiple subsets for different subjects and domains, structured for few-shot learning and evaluation. Dataset Structure The dataset is organized into subsets, each corresponding to a specific subject and level. Each subset has two splits: dev: Few-shot examples (3 rows per subset). test: The remaining rows for evaluation. Subsets The dataset contains the following subsets: Accounting (University) Biology (High… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/KazMMLU.tabularquestion-answering10K<n<100K11 likes402 downloads2y agoHugging Face10KomeijiForce /CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering10K<n<100K0 likes313 downloads3y agoHugging Face11gyung /korean-bar-exam-hard-current-law-precedent-sft-1000 Korean Current-Law Bar Exam Hard SFT 1000 대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다. 초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다. ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심 甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대 단순 근거 조문 선택형 제거 정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공 제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외 Files data/questions.csv: Hugging Face preview용 메인 CSV입니다. sft/train.jsonl: messages 형식 SFT용 JSONL입니다. metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.tabularquestion-answering1K<n<10K0 likes305 downloads3mo agoHugging Face12krammnic /hle-multichoiceHumanity Last Exam dataset with extra incorrect answers generated with Qwen3-4B tabulartable-question-answering1K<n<10K0 likes261 downloads1y agoHugging Face13OpenLab-NLP /tiny-singleturn-chat-kotextquestion-answering10K<n<100K0 likes197 downloads10mo agoHugging Face14maum-ai /KOFFVQA_Data About this data KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language KOFFVQA is a general-purpose VLM benchmark in the Korean language. For more information, refer to our leaderboard page and the official evaluation code. This contains the data for the benchmark consisting of images, their corresponding questions, and response grading criteria. The benchmark focuses on free-form visual question answering, evaluating the… See the full description on the dataset page: https://huggingface.co/datasets/maum-ai/KOFFVQA_Data.textvisual-question-answeringn<1K2 likes182 downloads1y agoHugging Face15k-master /k-beauty-ai-citation-dataset K-Beauty AI Citation Dataset Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research. Canonical source: https://kbeautyanswers.com/dataset/ License: CC BY 4.0 Maintainer: K-Beauty Answers (site) Initial release: 2026-05-23 What's in it 128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.texttext-classificationn<1K0 likes181 downloads3mo agoHugging Face16kguo2 /MolPuzzle_data MolPuzzle: A Multimodal Benchmark for Molecular Structure Elucidation Dataset Description: The MolPuzzle dataset is a newly developed resource designed to challenge Large Language Models with multi-modal, multi-step reasoning tasks (molecular structure elucidation). This dataset consists of 217 diverse and intricate structure elucidation challenges that require LLMs to demonstrate advanced reasoning capabilities, integrating multimodal data and deep chemical understanding… See the full description on the dataset page: https://huggingface.co/datasets/kguo2/MolPuzzle_data.imagequestion-answering10K<n<100K0 likes143 downloads2y agoHugging Face17KadamParth /Ncert_datasettabularquestion-answering100K<n<1M4 likes131 downloads1y agoHugging Face18HalcyonSolutions /Kinship Kinship KG–QA (Hinton) A lightweight knowledge-graph + question answering (KGQA) resource adapted from the UCI Kinship dataset created by Geoff Hinton. This dataset provides a small family-tree knowledge graph paired with templated multi-hop QA tasks designed for MultiHop KGQA. Project: THESEUSPaper: Theseus in the GraphOriginal dataset: UCI Kinship Key Features Small knowledge graph with two family trees Human-readable entities and relations Templated 1--3 hop… See the full description on the dataset page: https://huggingface.co/datasets/HalcyonSolutions/Kinship.tabularquestion-answering1K<n<10K1 likes127 downloads4d agoHugging Face19KomeijiForce /ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering1K<n<10K0 likes124 downloads3y agoHugging Face20kispeterzsm-szte /stackexchangeThis dataset is based entirely on HuggingFaceH4/stack-exchange-preferences, but it has been restructured. All HTML tags have been cleaned out, and the answers column has been turned into the answer column, so instead of answers being stored in JSON format there is now a row for each answer. Furthermore there is a separate file for every forum instead of a single file. textquestion-answering10M<n<100M2 likes115 downloads1y agoHugging Face21mssongit /KorfinQA FinQA 한국어 번역본 Question, Answer 총 6252 Rows textquestion-answering1K<n<10K4 likes108 downloads3y agoHugging Face22kayrab /patient-doctor-qa-tr-321179 Patient Doctor Q&A TR 321179 Veri Kümesi Patient Doctor Q&A TR 321179 veri kümesi, Patient Doctor Q&A TR 19583, Patient Doctor Q&A TR 167732, Patient Doctor Q&A TR 5695 ve Patient Doctor Q&A TR 95588 veri kümelerinin birleştirilmiş ve karıştırılmış halidir. Ana Özellikler: İçerik: Çeşitli tıbbi konuları kapsayan hasta soruları ve doktor yanıtları. Yapı: 2 sütun içerir: Soru, Cevap.Dil: Türkçe. Potansiyel Kullanım Alanları: Tıbbi araştırmalar Doğal Dil… See the full description on the dataset page: https://huggingface.co/datasets/kayrab/patient-doctor-qa-tr-321179.textquestion-answering100K<n<1M5 likes103 downloads2y agoHugging Face23KomeijiForce /ARC-Easy-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC-Easy. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering1K<n<10K1 likes94 downloads3y agoHugging Face24KBayoud /MoroccanHistory-QA-Datasettextquestion-answering1K<n<10K3 likes83 downloads3y agoHugging Face25ktiyab /ethical-framework-UNESCO-Ethics-of-AI Ethical AI Training Dataset Introduction UNESCO's Ethics of Artificial Intelligence, adopted by 193 Member States in November 2021, represents the first global framework for ethical AI development and deployment. While regional initiatives like The Montréal Declaration for a Responsible Development of Artificial Intelligence emphasize community-driven governance, UNESCO's approach establishes comprehensive international standards through coordinated multi-stakeholder… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/ethical-framework-UNESCO-Ethics-of-AI.textquestion-answeringn<1K3 likes77 downloads2y agoHugging Face26Kirtibg /TALES-QA TALES-QA: Cultural Knowledge Question Bank This repository contains the question bank introduced in our paper “TALES: A Taxonomy and Analysis of Cultural Representations in LLM-generated Stories.” The project webpage can be found here: https://cultural-misrepresentations.github.io/ This project evaluates cultural misrepresentations in LLM-generated stories for diverse Indian cultural identities. As part of this effort, we provide a curated question bank of standalone… See the full description on the dataset page: https://huggingface.co/datasets/Kirtibg/TALES-QA.textquestion-answering1K<n<10K0 likes70 downloads10mo agoHugging Face27TPelc /Current_Trivia_Knowledge-benchmark Current Trivia Knowledge RAG Benchmark Short Summary: A 140-QA pair (70 train, 70 test) dataset for real-world RAG evaluation. It features current knowledge questions unavailable to LLMs trained before 2024 (e.g., GPT-4o) across diverse domains, and includes human feedback for the training set, enabling robust assessment of contextual information's critical impact on LLM accuracy. Introduction & Motivation: This dataset addresses the critical need for a dynamic… See the full description on the dataset page: https://huggingface.co/datasets/TPelc/Current_Trivia_Knowledge-benchmark.textquestion-answeringn<1K0 likes68 downloads1y agoHugging Face28kg-rag /BiomixQA BiomixQA Dataset Overview BiomixQA is a curated biomedical question-answering dataset comprising two distinct components: Multiple Choice Questions (MCQ) True/False Questions This dataset has been utilized to validate the Knowledge Graph based Retrieval-Augmented Generation (KG-RAG) framework across different Large Language Models (LLMs). The diverse nature of questions in this dataset, spanning multiple choice and true/false formats, along with its coverage of various… See the full description on the dataset page: https://huggingface.co/datasets/kg-rag/BiomixQA.textquestion-answeringn<1K8 likes57 downloads1y agoHugging Face29KadamParth /NCERT_Physics_11thtabularquestion-answering1K<n<10K1 likes56 downloads2y agoHugging Face30KadamParth /NCERT_Business_Studies_12thtabularquestion-answering1K<n<10K2 likes55 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.