passage-ranking
msmarco_passage_ranking_corpusThis is the preprocessed data from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
msmarco-passage-rankingmsmarco_passage_ranking_official_trainThis is the preprocessed training data from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
msmarco_passage_ranking
MS MARCO Passage Ranking
Dataset description
MS MARCO (MicroSoft MAchine Reading COmprehension) is a large-scale collection built for machine reading comprehension and information retrieval research. The original release introduced more than one million real user questions sampled from Bing search logs, paired with passages drawn from web documents, and human-authored answers where applicable.
The passage ranking track uses a fixed corpus of short text passages and asks… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/msmarco_passage_ranking.msmarco_passage_ranking_queriesThis is the preprocessed queries from msmarco passage(v1) ranking corpus.
MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,.
Noisy-MSMARCO-Passage-RankingThis link gathers 72 noisy versions of the MS-Marco-Passage Ranking dataset consisting of three noise types (insertion, deletion, substitution), two different distributions of errors in the text (Batch 1 where errors are distributed in few words in the text and Batch2 where errors are more evenly spread out between words) and 12 different intensities of noise (CER varying from 3% to 36% with intervals of 3%).
The exact dataset that has been used is the MS-Marco-passagetest2020-top1000. The… See the full description on the dataset page: https://huggingface.co/datasets/edwardgiamphy/Noisy-MSMARCO-Passage-Ranking.
