datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commonsense_qa_2.0https://github.com/allenai/csqa2
@article{talmor2022commonsenseqa,
title={CommonsenseQA 2.0: Exposing the limits of AI through gamification},
author={Talmor, Alon and Yoran, Ori and Bras, Ronan Le and Bhagavatula, Chandra and Goldberg, Yoav and Choi, Yejin and Berant, Jonathan},
journal={arXiv preprint arXiv:2201.05320},
year={2022}
}
DemoFeedbackCommonsenseQA1000COTcommonsense_filtered
Dataset Summary
The commonsense reasoning tasks consist of 8 subtasks, each with predefined training and testing sets, as described by LLM-Adapters (Hu et al., 2023). The following table lists the details of each sub-dataset.
Train
Test
Information
BoolQ (Clark et al., 2019)
9427
3270
Question-answering dataset for yes/no questions
PIQA (Bisk et al., 2020)
16113
1838
Questions with two solutions requiring physical commonsense to answer
SIQA (Sap et al., 2019)
33410… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/commonsense_filtered.CommonsenseQA-GPT4ominiCOM2-commonsensecommonsense-dialogues
Commonsense-Dialogues Dataset
This is the Commonsense-Dialogues, a crowdsourced dataset of ~11K dialogues grounded in social contexts involving utilization of commonsense. The dataset was released by Amazon Alexa AI team in collaboration with the University of Southern California (USC), and also available Commonsense-Dialogues repo
The social contexts used were sourced from the train split of the SocialIQA dataset, a multiple-choice question-answering based social commonsense… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/commonsense-dialogues.chatgpt4-commonsense-qa
Synthetic CommonSense
Generated using ChatGPT4, originally from https://huggingface.co/datasets/commonsense_qa
Notebook at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/chatgpt4-commonsense
synthetic-commonsense.jsonl, 36332 rows, 7.34 MB.
Example data
{'question': '1. Seseorang yang bersara mungkin perlu kembali bekerja jika mereka apa?\n A. mempunyai hutang\n B. mencari pendapatan\n C. meninggalkan pekerjaan\n D. memerlukan… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt4-commonsense-qa.commonsense-v1
Basic knowledge
This is common sense, timeless knowledge so basic that any child in the last 150 years would know.
LLMs don't live in the real world and they don't know, or understand even the most basic of things.
This little dataset hopes to fix that.
Intentionally omitted knowledge:
modern concepts like computers, internet, rockets, planes and cars
modern biology, anathomy and physics
geography. In the last hundred years, empires have fallen, new countries were created and… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/commonsense-v1.commonsense_170kCommonsensenT2IDaily Paper: https://huggingface.co/papers/2406.07546
license: apache-2.0
commonsense_qa-mt-pt
CommonsenseQA-PT
Portuguese machine translation of CommonsenseQA, a multiple-choice question answering dataset that requires commonsense reasoning.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/tau/commonsense_qa
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/commonsense_qa-mt-pt.commonsense_qa_zh
Commonsense QA Chinese Multiple-Choice Dataset
This dataset is a Chinese four-choice SFT version of tau/commonsense_qa. It is designed to supplement commonsense multiple-choice training data for benchmark tasks such as challenge_common_sense.
The original dataset is in English and contains five-choice commonsense questions. This release keeps only samples that can be aligned to the official four-choice benchmark format, translates the question and options into… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/commonsense_qa_zh.Ethics_commonsense_chinesecommonsenseqa
commonsenseqa — evaluation data (OpenCompass format)
Bud Ecosystem eval mirror. OpenCompass-format evaluation data for commonsenseqa, for offline reproducible model evaluation (config commonsenseqa_gen). Original source: tau/commonsense_qa — license MIT, unchanged; all rights remain with the original authors.
hads-physical-commonsense
HADS – Human Action and Decision Sense (حَدس)
📄 Paper: HADS: A Large-Scale Parallel Benchmark for Physical Commonsense Reasoning — IEEE Access, 2026 (doi:10.1109/ACCESS.2026.3705337)
HADS is a large-scale Arabic parallel adaptation of the English
PIQA benchmark for physical commonsense reasoning.
The name derives from the Arabic word حَدس (hads), meaning physical intuition
or gut sense — the tacit embodied knowledge the benchmark measures.
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/IWAN/hads-physical-commonsense.shastraai__Shastra-LLAMA2-Math-Commonsense-SFT-details
Dataset Card for Evaluation run of shastraai/Shastra-LLAMA2-Math-Commonsense-SFT
Dataset automatically created during the evaluation run of model shastraai/Shastra-LLAMA2-Math-Commonsense-SFT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/shastraai__Shastra-LLAMA2-Math-Commonsense-SFT-details.commonsense_qa_test_cleancommonsense_kocommonsense
