CoolFace
7 results

passage-ranking

jacklin /msmarco_passage_ranking_corpusThis is the preprocessed data from msmarco passage(v1) ranking corpus. MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,. text1M<n<10M0 likes403 downloads4y agoHugging Facezyznull /msmarco-passage-ranking1 likes78 downloads4y agoHugging Facejacklin /msmarco_passage_ranking_official_trainThis is the preprocessed training data from msmarco passage(v1) ranking corpus. MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,. text100K<n<1M0 likes76 downloads4y agoHugging Faceorgrctera /msmarco_passage_ranking MS MARCO Passage Ranking Dataset description MS MARCO (MicroSoft MAchine Reading COmprehension) is a large-scale collection built for machine reading comprehension and information retrieval research. The original release introduced more than one million real user questions sampled from Bing search logs, paired with passages drawn from web documents, and human-authored answers where applicable. The passage ranking track uses a fixed corpus of short text passages and asks… See the full description on the dataset page: https://huggingface.co/datasets/orgrctera/msmarco_passage_ranking.text10K<n<100K0 likes52 downloads6mo agoHugging Facejacklin /msmarco_passage_ranking_queriesThis is the preprocessed queries from msmarco passage(v1) ranking corpus. MS MARCO: A human generated MAchine Reading COmprehension dataset SPayal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen,. text100K<n<1M0 likes48 downloads4y agoHugging Faceedwardgiamphy /Noisy-MSMARCO-Passage-RankingThis link gathers 72 noisy versions of the MS-Marco-Passage Ranking dataset consisting of three noise types (insertion, deletion, substitution), two different distributions of errors in the text (Batch 1 where errors are distributed in few words in the text and Batch2 where errors are more evenly spread out between words) and 12 different intensities of noise (CER varying from 3% to 36% with intervals of 3%). The exact dataset that has been used is the MS-Marco-passagetest2020-top1000. The… See the full description on the dataset page: https://huggingface.co/datasets/edwardgiamphy/Noisy-MSMARCO-Passage-Ranking.0 likes34 downloads3y agoHugging Face