CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tau /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/tau/commonsense_qa.textquestion-answering10K<n<100K155 likes278k downloads3y agoHugging Face02tasksource /commonsense_qa_2.0https://github.com/allenai/csqa2 @article{talmor2022commonsenseqa, title={CommonsenseQA 2.0: Exposing the limits of AI through gamification}, author={Talmor, Alon and Yoran, Ori and Bras, Ronan Le and Bhagavatula, Chandra and Goldberg, Yoav and Choi, Yejin and Berant, Jonathan}, journal={arXiv preprint arXiv:2201.05320}, year={2022} } textquestion-answering10K<n<100K4 likes1.6k downloads3y agoHugging Face03KomeijiForce /CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering10K<n<100K0 likes321 downloads3y agoHugging Face04BENBENBENb /CommonsenseQA1000COTtextquestion-answering1K<n<10K0 likes181 downloads3y agoHugging Face05ingoziegler /CRAFT-CommonSenseQA CRAFT-CommonSenseQA This is a synthetic dataset generated with the CRAFT framework proposed in the paper CRAFT Your Dataset: Task-Specific Synthetic Data Generation Through Corpus Retrieval and Augmentation. The correctness of the data has not been verified in detail, but training on this data and evaluating on human-curated commonsense question-answering data proved highly beneficial. 4 synthetic dataset sizes (S, M, L, XL) are available, and training on them yields consistent… See the full description on the dataset page: https://huggingface.co/datasets/ingoziegler/CRAFT-CommonSenseQA.texttext-classification10K<n<100K2 likes125 downloads10mo agoHugging Face06fxmeng /commonsense_filtered Dataset Summary The commonsense reasoning tasks consist of 8 subtasks, each with predefined training and testing sets, as described by LLM-Adapters (Hu et al., 2023). The following table lists the details of each sub-dataset. Train Test Information BoolQ (Clark et al., 2019) 9427 3270 Question-answering dataset for yes/no questions PIQA (Bisk et al., 2020) 16113 1838 Questions with two solutions requiring physical commonsense to answer SIQA (Sap et al., 2019) 33410… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/commonsense_filtered.textquestion-answering100K<n<1M1 likes124 downloads2y agoHugging Face07peterkchung /commonsense_cot_partial_raw Commonsense QA CoT (Partial, Raw, No Human Annotation) Dataset Summary Seeded by the CommonsenseQA dataset (tau/commonsense_qa) this preliminary set randomly samples 1,000 question-answer entries and uses Mixtral (mistralai/Mixtral-8x7B-Instruct-v0.1) to generate 3 unique CoT (Chain-of-Thought) rationales. This was created as the preliminary step towards fine-tuning a LM (language model) to specialize on commonsense reasoning. The working hypothesis, inspired by the… See the full description on the dataset page: https://huggingface.co/datasets/peterkchung/commonsense_cot_partial_raw.textquestion-answering1K<n<10K1 likes73 downloads3y agoHugging Face08mesolitica /chatgpt4-commonsense-qa Synthetic CommonSense Generated using ChatGPT4, originally from https://huggingface.co/datasets/commonsense_qa Notebook at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/chatgpt4-commonsense synthetic-commonsense.jsonl, 36332 rows, 7.34 MB. Example data {'question': '1. Seseorang yang bersara mungkin perlu kembali bekerja jika mereka apa?\n A. mempunyai hutang\n B. mencari pendapatan\n C. meninggalkan pekerjaan\n D. memerlukan… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt4-commonsense-qa.textquestion-answering10K<n<100K1 likes72 downloads3y agoHugging Face09hishab /commonsenseqa-bn Dataset Summary This is the Bangla translated version of the CommonsenseQA dataset. The dataset was translated using a new method called Expressive Semantic Translation (EST). This method combines both Google Machine Translation and LLM-based rewriting of the translation to enhance the expressiveness and semantic accuracy of the translated content. Dataset Structure Data instances Defaults An example of a 'train' looks as follows: {… See the full description on the dataset page: https://huggingface.co/datasets/hishab/commonsenseqa-bn.textquestion-answering10K<n<100K0 likes57 downloads2y agoHugging Face10choucsan /commonsense_qa_zh Commonsense QA Chinese Multiple-Choice Dataset This dataset is a Chinese four-choice SFT version of tau/commonsense_qa. It is designed to supplement commonsense multiple-choice training data for benchmark tasks such as challenge_common_sense. The original dataset is in English and contains five-choice commonsense questions. This release keeps only samples that can be aligned to the official four-choice benchmark format, translates the question and options into… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/commonsense_qa_zh.textquestion-answering1K<n<10K5 likes35 downloads2mo agoHugging Face11spacekat99 /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/commonsense_qa.textquestion-answering10K<n<100K0 likes27 downloads4mo agoHugging Face12psjadf /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/psjadf/commonsense_qa.textquestion-answering10K<n<100K0 likes27 downloads4mo agoHugging Face13amalia-llm /commonsense_qa-mt-pt CommonsenseQA-PT Portuguese machine translation of CommonsenseQA, a multiple-choice question answering dataset that requires commonsense reasoning. Translated using a Finetuned GemmaX2-9B for pt-PT. Original Dataset: https://huggingface.co/datasets/tau/commonsense_qa Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/commonsense_qa-mt-pt.textquestion-answering10K<n<100K0 likes25 downloads3mo agoHugging Face14farabi-lab /Common-Sense-Reasoninggated 🇰🇿 Kazakh General Inquiry and FAQ Dataset 📖 Overview This dataset contains 1,000 high-quality question-and-answer pairs in the Kazakh language. It is designed to train models on providing helpful, natural, and informative responses to common inquiries. 📊 Dataset Statistics General Metrics Metric Count Total Samples 1,000 Total Words (approx.) 74,013 Avg. Words per Sample 74 Word Count Distribution… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Common-Sense-Reasoning.textquestion-answering1K<n<10K0 likes24 downloads2mo agoHugging Face15AdrianSword /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/AdrianSword/commonsense_qa.textquestion-answering10K<n<100K0 likes23 downloads10mo agoHugging Face16tbilisi-ai-lab /commonsense_qa-ka commonsense_qa-ka Georgian translation of the CommonsenseQA benchmark. Dataset Summary Property Value Examples 1,221 Splits validation Languages Georgian, English Task Commonsense Reasoning Data Fields id: Example ID question: Question (English) question_concept: Core concept choices: Answer choices (English) answerKey: Correct answer key question_ka: Question (Georgian) choices_ka: Answer choices (Georgian) Translation… See the full description on the dataset page: https://huggingface.co/datasets/tbilisi-ai-lab/commonsense_qa-ka.textquestion-answering1K<n<10K1 likes20 downloads4mo agoHugging Face17wisenut-nlp-team /aihub_mrc_commonsense Dataset Card for "mrc_aihub_common_sense" 일반 상식 textquestion-answering100K<n<1M0 likes18 downloads3y agoHugging Face18rizquuula /commonsense_qa-IDCommonsenseQA-ID is Indonesian translation version of CommonsenseQA, a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.textquestion-answering10K<n<100K0 likes18 downloads3y agoHugging Face19JazonDeng /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/JazonDeng/commonsense_qa.textquestion-answering10K<n<100K0 likes18 downloads5mo agoHugging Face20spacekat99 /commonsense_cot_partial_raw Commonsense QA CoT (Partial, Raw, No Human Annotation) Dataset Summary Seeded by the CommonsenseQA dataset (tau/commonsense_qa) this preliminary set randomly samples 1,000 question-answer entries and uses Mixtral (mistralai/Mixtral-8x7B-Instruct-v0.1) to generate 3 unique CoT (Chain-of-Thought) rationales. This was created as the preliminary step towards fine-tuning a LM (language model) to specialize on commonsense reasoning. The working hypothesis, inspired by the… See the full description on the dataset page: https://huggingface.co/datasets/spacekat99/commonsense_cot_partial_raw.textquestion-answering1K<n<10K0 likes16 downloads4mo agoHugging Face21IWAN /hads-physical-commonsensegated HADS – Human Action and Decision Sense (حَدس) 📄 Paper: HADS: A Large-Scale Parallel Benchmark for Physical Commonsense Reasoning — IEEE Access, 2026 (doi:10.1109/ACCESS.2026.3705337) HADS is a large-scale Arabic parallel adaptation of the English PIQA benchmark for physical commonsense reasoning. The name derives from the Arabic word حَدس (hads), meaning physical intuition or gut sense — the tacit embodied knowledge the benchmark measures. Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/IWAN/hads-physical-commonsense.tabularquestion-answering10K<n<100K0 likes13 downloads3mo agoHugging Face22chouxtteok /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/chouxtteok/commonsense_qa.textquestion-answering10K<n<100K0 likes6 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.