datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco_passage_ranking_corpusThis is the preprocessed data from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
msmarco_passage_ranking_official_trainThis is the preprocessed training data from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
msmarco_passage_ranking
MS MARCO Passage Ranking
Dataset description
MS MARCO (MicroSoft MAchine Reading COmprehension) is a large-scale collection built for machine reading comprehension and information retrieval research. The original release introduced more than one million real user questions sampled from Bing search logs, paired with passages drawn from web documents, and human-authored answers where applicable.
The passage ranking track uses a fixed corpus of short text passages and asks… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/msmarco_passage_ranking.msmarco_passage_ranking_queriesThis is the preprocessed queries from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
Passage_ranking6_testPassage_ranking5_testPassage_ranking_testPassage_ranking1_testPassage_ranking2_testPassage_ranking4_testPassage_ranking7_testPassage_ranking3_test
