datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mqaMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.TSQA
Time Series Question Answering Dataset (TSQA)
Introduction
TSQA dataset is a large-scale collection of ~200,000 QA pairs covering 12 real-world application domains such as healthcare, environment, energy, finance, transport, IoT, nature, human activities, AIOps, and the web. TSQA also includes 5 task types: forecasting, imputation, anomaly detection, classification, and open-ended reasoning. Within the open-ended reasoning QA, the dataset includes 6,919 true/false… See the full description on the dataset page: https://huggingface.co/datasets/Time-MQA/TSQA.agentic-mqa-objectsmqa-jamqaデータセットのquery--passageのペアについて重複を削除したデータセットです。
元データ中のノイジーなテキストのクリーニングやNFKC正規化などの前処理を行ってあります。
dataset subsetのpos_idsおよびneg_ids中のidは、collectionsubsetのインデックス番号に対応しています。
したがって、collection[pos_id]のようにアクセスしてもらえれば所望のデータを得ることができます。
ライセンスは元データセットに従います。
spoken-mqa@article{wei2025towards,
title={Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems},
author={Wei, Chengwei and Wang, Bin and Kim, Jung-jae and Chen, Nancy F},
journal={arXiv preprint arXiv:2505.15000},
year={2025}
}
mqar_N8_V8192_L24_noise0.9romteb-mqa-ro-cqa-retrievalmqar_N8_V8192_L24_noise0.25ms_marco_mqa
SHINE Dataset
This repository contains datasets associated with the paper SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass.
Paper | Code
Introduction
SHINE (Scalable Hyper In-context NEtwork) is a scalable hypernetwork designed to map diverse meaningful contexts into high-quality LoRA adapters for large language models (LLM) in a single forward pass. By reusing the frozen LLM's own parameters and introducing architectural… See the full description on the dataset page: https://huggingface.co/datasets/Yewei-Liu/ms_marco_mqa.m_qalmThe M-QALM Dataset Repository contains Multiple-Choice and Abstractive Questions for evaluating the performance of LLMs in the clinical and biomedical domain.mqar_N32_V8192_L96mqar_N64_V8192_L192mqar_N16_V8192_L48mqar_N8_V8192_L24_noise0.5mqar_N8_V8192_L24_noise0.75mqa-ja
mqa-ja
日本語 QA データセット hpprc/mqa-ja を元に、Hard Negative Mining と Cross-Encoder 蒸留スコア付与を施した日本語 retrieval 学習用データセットです。
pairs / triplets / n-tuples の 3 形式と、n-tuples に蒸留スコアを付けて学習価値でソートした top-K サブセット (100k / 250k / 500k / 1m) を提供します。
Dense Retriever / Cross-Encoder / SPLADE などの日本語検索モデル学習および KL Divergence 蒸留に利用できます。
Configs
config
rows
columns
説明
pairs
5,823,586
query, answer
元の (query, positive) ペア (anc → query, pos_ids[0] → answer)。HNM の結果に関わらず全行収録。
triplets
5,823… See the full description on the dataset page: https://huggingface.co/datasets/mahiyama/mqa-ja.mqar_N8_V8192_L24hebrew-psychotechnique-MQA
Israeli Psychometric Exam (NITE) — Multiple-Choice QA
1570 multiple-choice questions extracted from 26 publicly released NITE
psychometric entrance exams (2019–2026).
Splits
split
rows
notes
verbal
927
Hebrew, RTL
english
620
English
quantitative
23
almost nothing survives filtering
Fields
question, options (4), answer (1-indexed into options)
section / part / number — position within the exam
source_pdf / page — provenance, for… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-psychotechnique-MQA.ift_mqa_collection
SHINE Dataset
This repository contains datasets associated with the paper SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass.
Paper | Code
Introduction
SHINE (Scalable Hyper In-context NEtwork) is a scalable hypernetwork designed to map diverse meaningful contexts into high-quality LoRA adapters for large language models (LLM) in a single forward pass. By reusing the frozen LLM's own parameters and introducing architectural… See the full description on the dataset page: https://huggingface.co/datasets/Yewei-Liu/ift_mqa_collection.mqaMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.easy_mqarallergen-ner-dataset-v2m-qalm-finalgl_MQA_artistas_mulleres_USC
gl_MQA_artistas_mulleres_USC
Dataset Description
gl_MQA_artistas_mulleres_USC is a Galician multiple-choice question-answering dataset focused on women artists and artistic heritage. It includes questions about authors, artworks, styles, techniques, and biographical or descriptive information related to women creators.
Dataset Summary
The dataset contains 85 examples. Each example consists of a multiple-choice question, four answer options, the… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/gl_MQA_artistas_mulleres_USC.gl_MQA_bensculturais
gl_MQA_bensculturais
Dataset Description
gl_MQA_bensculturais is a Galician multiple-choice question-answering dataset focused on cultural heritage. It is designed to evaluate knowledge of heritage terminology, definitions, uses, characteristics, and classification of cultural objects and concepts.
Dataset Summary
The dataset contains 418 examples. Each example consists of a question, four answer options, the correct answer, and a source identifier.… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/gl_MQA_bensculturais.mqa-kosin-gen-mqagl_MQA_museo_virtual_USC
gl_MQA_museo_virtual_USC
Dataset Description
gl_MQA_museo_virtual_USC is a Galician multiple-choice question-answering dataset focused on heritage collections associated with the Universidade de Santiago de Compostela. It is oriented to questions about pieces, authorship, techniques, material characteristics, context of creation, and descriptive information.
Dataset Summary
The dataset contains 167 examples. Each example consists of a… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/gl_MQA_museo_virtual_USC.gl_MQA_arte_galego_USC
gl_MQA_arte_galego_USC
Dataset Description
gl_MQA_arte_galego_USC is a Galician multiple-choice question-answering dataset focused on Galician art. It contains questions about artworks, artists, styles, themes, techniques, compositional features, and descriptive aspects of artistic heritage.
Dataset Summary
The dataset contains 45 examples. Each example consists of a question, four answer options, the correct answer, and a source identifier.… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/gl_MQA_arte_galego_USC.Time-MQA_TSQA_tmp
