datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mqaMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.TSQA
Time Series Question Answering Dataset (TSQA)
Introduction
TSQA dataset is a large-scale collection of ~200,000 QA pairs covering 12 real-world application domains such as healthcare, environment, energy, finance, transport, IoT, nature, human activities, AIOps, and the web. TSQA also includes 5 task types: forecasting, imputation, anomaly detection, classification, and open-ended reasoning. Within the open-ended reasoning QA, the dataset includes 6,919 true/false… See the full description on the dataset page: https://huggingface.co/datasets/Time-MQA/TSQA.m_qalmThe M-QALM Dataset Repository contains Multiple-Choice and Abstractive Questions for evaluating the performance of LLMs in the clinical and biomedical domain.hebrew-psychotechnique-MQA
Israeli Psychometric Exam (NITE) — Multiple-Choice QA
1570 multiple-choice questions extracted from 26 publicly released NITE
psychometric entrance exams (2019–2026).
Splits
split
rows
notes
verbal
927
Hebrew, RTL
english
620
English
quantitative
23
almost nothing survives filtering
Fields
question, options (4), answer (1-indexed into options)
section / part / number — position within the exam
source_pdf / page — provenance, for… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-psychotechnique-MQA.mqaMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.gl_MQA_bensculturais
gl_MQA_bensculturais
Dataset Description
gl_MQA_bensculturais is a Galician multiple-choice question-answering dataset focused on cultural heritage. It is designed to evaluate knowledge of heritage terminology, definitions, uses, characteristics, and classification of cultural objects and concepts.
Dataset Summary
The dataset contains 418 examples. Each example consists of a question, four answer options, the correct answer, and a source identifier.… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/gl_MQA_bensculturais.gl_MQA_artistas_mulleres_USC
gl_MQA_artistas_mulleres_USC
Dataset Description
gl_MQA_artistas_mulleres_USC is a Galician multiple-choice question-answering dataset focused on women artists and artistic heritage. It includes questions about authors, artworks, styles, techniques, and biographical or descriptive information related to women creators.
Dataset Summary
The dataset contains 85 examples. Each example consists of a multiple-choice question, four answer options, the… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/gl_MQA_artistas_mulleres_USC.gl_MQA_museo_virtual_USC
gl_MQA_museo_virtual_USC
Dataset Description
gl_MQA_museo_virtual_USC is a Galician multiple-choice question-answering dataset focused on heritage collections associated with the Universidade de Santiago de Compostela. It is oriented to questions about pieces, authorship, techniques, material characteristics, context of creation, and descriptive information.
Dataset Summary
The dataset contains 167 examples. Each example consists of a… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/gl_MQA_museo_virtual_USC.gl_MQA_arte_galego_USC
gl_MQA_arte_galego_USC
Dataset Description
gl_MQA_arte_galego_USC is a Galician multiple-choice question-answering dataset focused on Galician art. It contains questions about artworks, artists, styles, themes, techniques, compositional features, and descriptive aspects of artistic heritage.
Dataset Summary
The dataset contains 45 examples. Each example consists of a question, four answer options, the correct answer, and a source identifier.… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/gl_MQA_arte_galego_USC.mqa
MQA
Aggregation of datasets as per here
I reserve no rights to the dataset, but the original datasets were made available under various public licenses. Hence, consider each subset of this dataset to be licensed as the original dataset from where it comes was.
MQA
