MQA
Datasets
All datasets matching “MQA”mqaMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.TSQA
Time Series Question Answering Dataset (TSQA)
Introduction
TSQA dataset is a large-scale collection of ~200,000 QA pairs covering 12 real-world application domains such as healthcare, environment, energy, finance, transport, IoT, nature, human activities, AIOps, and the web. TSQA also includes 5 task types: forecasting, imputation, anomaly detection, classification, and open-ended reasoning. Within the open-ended reasoning QA, the dataset includes 6,919 true/false… See the full description on the dataset page: https://huggingface.co/datasets/Time-MQA/TSQA.agentic-mqa-objectsmqa-jamqaデータセットのquery--passageのペアについて重複を削除したデータセットです。
元データ中のノイジーなテキストのクリーニングやNFKC正規化などの前処理を行ってあります。
dataset subsetのpos_idsおよびneg_ids中のidは、collectionsubsetのインデックス番号に対応しています。
したがって、collection[pos_id]のようにアクセスしてもらえれば所望のデータを得ることができます。
ライセンスは元データセットに従います。
spoken-mqa@article{wei2025towards,
title={Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems},
author={Wei, Chengwei and Wang, Bin and Kim, Jung-jae and Chen, Nancy F},
journal={arXiv preprint arXiv:2505.15000},
year={2025}
}
mqar_N8_V8192_L24_noise0.9
