datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mqaMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.mqa-jamqaデータセットのquery--passageのペアについて重複を削除したデータセットです。
元データ中のノイジーなテキストのクリーニングやNFKC正規化などの前処理を行ってあります。
dataset subsetのpos_idsおよびneg_ids中のidは、collectionsubsetのインデックス番号に対応しています。
したがって、collection[pos_id]のようにアクセスしてもらえれば所望のデータを得ることができます。
ライセンスは元データセットに従います。
spoken-mqa@article{wei2025towards,
title={Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems},
author={Wei, Chengwei and Wang, Bin and Kim, Jung-jae and Chen, Nancy F},
journal={arXiv preprint arXiv:2505.15000},
year={2025}
}
romteb-mqa-ro-cqa-retrievalmqa-ja
mqa-ja
日本語 QA データセット hpprc/mqa-ja を元に、Hard Negative Mining と Cross-Encoder 蒸留スコア付与を施した日本語 retrieval 学習用データセットです。
pairs / triplets / n-tuples の 3 形式と、n-tuples に蒸留スコアを付けて学習価値でソートした top-K サブセット (100k / 250k / 500k / 1m) を提供します。
Dense Retriever / Cross-Encoder / SPLADE などの日本語検索モデル学習および KL Divergence 蒸留に利用できます。
Configs
config
rows
columns
説明
pairs
5,823,586
query, answer
元の (query, positive) ペア (anc → query, pos_ids[0] → answer)。HNM の結果に関わらず全行収録。
triplets
5,823… See the full description on the dataset page: https://huggingface.co/datasets/mahiyama/mqa-ja.hebrew-psychotechnique-MQA
Israeli Psychometric Exam (NITE) — Multiple-Choice QA
1570 multiple-choice questions extracted from 26 publicly released NITE
psychometric entrance exams (2019–2026).
Splits
split
rows
notes
verbal
927
Hebrew, RTL
english
620
English
quantitative
23
almost nothing survives filtering
Fields
question, options (4), answer (1-indexed into options)
section / part / number — position within the exam
source_pdf / page — provenance, for… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-psychotechnique-MQA.easy_mqarallergen-ner-dataset-v2mqa-koMQAdolphin_mqa_details
Dataset Card for "dolphin_mqa_details"
More Information needed
dolphin_mqa_details_vi
Dataset Card for "dolphin_mqa_details_vi"
More Information needed
tky_persona_mqa
tky_persona_mqa Dataset Overview
tky_persona_mqa.json stores 67,575 narrated day-in-the-life summaries from Tokyo trajectories, each paired with two persona labels chosen by GPT-5.
Field Definitions
user_id: String linking the narrative back to its trajectory instance.
text: English-language narrative describing hourly activities inferred from GPS traces and nearby POIs.
choice: Two ordered persona labels. GPT-5 places its most plausible persona first and the least… See the full description on the dataset page: https://huggingface.co/datasets/bigchestnut/tky_persona_mqa.mqa
MQA
Aggregation of datasets as per here
I reserve no rights to the dataset, but the original datasets were made available under various public licenses. Hence, consider each subset of this dataset to be licensed as the original dataset from where it comes was.
mqa_llama_datasetvmlu-vi-mqa-answers-initial-phasenyc_persona_mqa
nyc_persona_mqa Dataset Overview
nyc_persona_mqa.json captures 65,115 narrated day-in-the-life summaries from New York City visitor and resident trajectories, each annotated with two persona hypotheses drafted by GPT-5 based on observed movement patterns and nearby POIs.
Field Definitions
user_id: Identifier that links the narrative back to a specific NYC trajectory instance.
text: English summary describing hourly activities inferred from GPS traces and contextual POI… See the full description on the dataset page: https://huggingface.co/datasets/bigchestnut/nyc_persona_mqa.mqa1
license: mit
Dataset Card
Developed by: [More Information Needed]
Shared by [optional]: [More Information Needed]
Dataset type: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Derived from dataset [optional]: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/mqa1.allergen-ner-english-v1xehe_mqaneao_lwnmqapMQAvbpl_pl_test_mqavbpl_mqa
Dataset Card for "vbpl_mqa"
More Information needed
