CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01castorini /msmarco_v2_doc_segmented_doc2query-t5_expansions Dataset Summary The repo provides queries generated for the MS MARCO v2 document segmented corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v2_doc_segmented_doc2query-t5_expansions.text10M<n<100M0 likes184 downloads5y agoHugging Face02castorini /msmarco_v2_passage_doc2query-t5_expansions Dataset Summary The repo provides queries generated for the MS MARCO v2 passage corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then appended to the… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v2_passage_doc2query-t5_expansions.text1M<n<10M0 likes136 downloads5y agoHugging Face03castorini /msmarco_v1_passage_doc2query-t5_expansions Dataset Summary The repo provides queries generated for the MS MARCO V1 passage corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then appended to the… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v1_passage_doc2query-t5_expansions.text0 likes109 downloads4y agoHugging Face04castorini /msmarco_v1_doc_doc2query-t5_expansions Dataset Summary The repo provides queries generated for the MS MARCO V1 document corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then appended to the… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v1_doc_doc2query-t5_expansions.text1M<n<10M0 likes97 downloads4y agoHugging Face05Turbo-AI /doc2query-generate-filter Dataset Card for "triplet-generate-filter-v2" More Information needed text100K<n<1M0 likes86 downloads2y agoHugging Face06castorini /msmarco_v2_doc_doc2query-t5_expansions Dataset Summary The repo provides queries generated for the MS MARCO v2 document corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then appended to the… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v2_doc_doc2query-t5_expansions.text100K<n<1M1 likes81 downloads5y agoHugging Face07castorini /msmarco_v1_doc_segmented_doc2query-t5_expansions Dataset Summary The repo provides queries generated for the MS MARCO V1 document segmented corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v1_doc_segmented_doc2query-t5_expansions.0 likes73 downloads5y agoHugging Face08Turbo-AI /doc2query-generate Dataset Card for "doc2query-generate-v2" More Information needed text1M<n<10M0 likes63 downloads2y agoHugging Face09Turbo-AI /data-doc2query-generate Dataset Card for "triplet-generate-filter-v3" More Information needed text100K<n<1M0 likes58 downloads2y agoHugging Face10pa-shk /sberquad-doc2querytext1K<n<10K0 likes23 downloads2y agoHugging Face11Watheq /doc2query_scored_queries Scores of generated queries This repo contains the scores files pertaining to this study. In particular, we scored the expansion queries generated by T5-based Doc2Query model for MSMARCO-v1 passage dataset and a subset of BEIR benchemark. We used ELECTRA cross-encoder to get the relevance scores between the document text and its expansion queries. More details in the study repo here. Structure All files are .jsonl files with the following three columns per line: ["id"… See the full description on the dataset page: https://huggingface.co/datasets/Watheq/doc2query_scored_queries.text1M<n<10M1 likes15 downloads2y agoHugging Face12Ayzengin /doc2query-vector-store0 likes4 downloads6mo agoHugging Face13spacemanidol /ms_marco_doc2query0 likes3 downloads5y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.