CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01clips /mqaMQA is a multilingual corpus of questions and answers parsed from the Common Crawl. Questions are divided between Frequently Asked Questions (FAQ) pages and Community Question Answering (CQA) pages.textquestion-answering100M<n<1B57 likes984 downloads4y agoHugging Face02hpprc /mqa-jamqaデータセットのquery--passageのペアについて重複を削除したデータセットです。 元データ中のノイジーなテキストのクリーニングやNFKC正規化などの前処理を行ってあります。 dataset subsetのpos_idsおよびneg_ids中のidは、collectionsubsetのインデックス番号に対応しています。 したがって、collection[pos_id]のようにアクセスしてもらえれば所望のデータを得ることができます。 ライセンスは元データセットに従います。 text10M<n<100M6 likes178 downloads2y agoHugging Face03amao0o0 /spoken-mqa@article{wei2025towards, title={Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems}, author={Wei, Chengwei and Wang, Bin and Kim, Jung-jae and Chen, Nancy F}, journal={arXiv preprint arXiv:2505.15000}, year={2025} } audio1K<n<10K5 likes112 downloads1y agoHugging Face04alina0195 /romteb-mqa-ro-cqa-retrievaltext100K<n<1M0 likes68 downloads27d agoHugging Face05mahiyama /mqa-ja mqa-ja 日本語 QA データセット hpprc/mqa-ja を元に、Hard Negative Mining と Cross-Encoder 蒸留スコア付与を施した日本語 retrieval 学習用データセットです。 pairs / triplets / n-tuples の 3 形式と、n-tuples に蒸留スコアを付けて学習価値でソートした top-K サブセット (100k / 250k / 500k / 1m) を提供します。 Dense Retriever / Cross-Encoder / SPLADE などの日本語検索モデル学習および KL Divergence 蒸留に利用できます。 Configs config rows columns 説明 pairs 5,823,586 query, answer 元の (query, positive) ペア (anc → query, pos_ids[0] → answer)。HNM の結果に関わらず全行収録。 triplets 5,823… See the full description on the dataset page: https://huggingface.co/datasets/mahiyama/mqa-ja.textsentence-similarity10M<n<100M0 likes48 downloads4mo agoHugging Face06guychuk /hebrew-psychotechnique-MQA Israeli Psychometric Exam (NITE) — Multiple-Choice QA 1570 multiple-choice questions extracted from 26 publicly released NITE psychometric entrance exams (2019–2026). Splits split rows notes verbal 927 Hebrew, RTL english 620 English quantitative 23 almost nothing survives filtering Fields question, options (4), answer (1-indexed into options) section / part / number — position within the exam source_pdf / page — provenance, for… See the full description on the dataset page: https://huggingface.co/datasets/guychuk/hebrew-psychotechnique-MQA.tabularmultiple-choice1K<n<10K0 likes43 downloads23d agoHugging Face07Bradley /easy_mqartext1K<n<10K0 likes31 downloads11mo agoHugging Face08MQareen /allergen-ner-dataset-v2text10K<n<100K0 likes28 downloads7mo agoHugging Face09kozistr /mqa-kotext1M<n<10M0 likes20 downloads2y agoHugging Face10Twwilght /MQAtext10K<n<100K0 likes15 downloads8mo agoHugging Face11nguyenthanhdo /dolphin_mqa_details Dataset Card for "dolphin_mqa_details" More Information needed text10K<n<100K0 likes13 downloads3y agoHugging Face12nguyenthanhdo /dolphin_mqa_details_vi Dataset Card for "dolphin_mqa_details_vi" More Information needed text10K<n<100K0 likes12 downloads3y agoHugging Face13bigchestnut /tky_persona_mqa tky_persona_mqa Dataset Overview tky_persona_mqa.json stores 67,575 narrated day-in-the-life summaries from Tokyo trajectories, each paired with two persona labels chosen by GPT-5. Field Definitions user_id: String linking the narrative back to its trajectory instance. text: English-language narrative describing hourly activities inferred from GPS traces and nearby POIs. choice: Two ordered persona labels. GPT-5 places its most plausible persona first and the least… See the full description on the dataset page: https://huggingface.co/datasets/bigchestnut/tky_persona_mqa.text10K<n<100K0 likes12 downloads7mo agoHugging Face14euclaise /mqa MQA Aggregation of datasets as per here I reserve no rights to the dataset, but the original datasets were made available under various public licenses. Hence, consider each subset of this dataset to be licensed as the original dataset from where it comes was. textquestion-answering10K<n<100K0 likes10 downloads3y agoHugging Face15Afiqa /mqa_llama_datasettextn<1K0 likes8 downloads2y agoHugging Face16nygdon /vmlu-vi-mqa-answers-initial-phasetextmultiple-choice1K<n<10K0 likes8 downloads5mo agoHugging Face17bigchestnut /nyc_persona_mqa nyc_persona_mqa Dataset Overview nyc_persona_mqa.json captures 65,115 narrated day-in-the-life summaries from New York City visitor and resident trajectories, each annotated with two persona hypotheses drafted by GPT-5 based on observed movement patterns and nearby POIs. Field Definitions user_id: Identifier that links the narrative back to a specific NYC trajectory instance. text: English summary describing hourly activities inferred from GPS traces and contextual POI… See the full description on the dataset page: https://huggingface.co/datasets/bigchestnut/nyc_persona_mqa.text10K<n<100K0 likes7 downloads7mo agoHugging Face18Taylor658 /mqa1 license: mit Dataset Card Developed by: [More Information Needed] Shared by [optional]: [More Information Needed] Dataset type: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Derived from dataset [optional]: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/mqa1.textn<1K0 likes6 downloads3y agoHugging Face19MQareen /allergen-ner-english-v1text10K<n<100K0 likes6 downloads7mo agoHugging Face20dwjoeikr /xehe_mqaneao_lwnmqaptext10M<n<100M0 likes4 downloads4mo agoHugging Face21kammavidya /MQAgatedtextquestion-answeringn<1K0 likes3 downloads3y agoHugging Face22NghiemAbe /vbpl_pl_test_mqagatedtextn<1K0 likes1 downloads1y agoHugging Face23NghiemAbe /vbpl_mqagated Dataset Card for "vbpl_mqa" More Information needed text10K<n<100K0 likes1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.