datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco_v2_doc_segmented_doc2query-t5_expansions
Dataset Summary
The repo provides queries generated for the MS MARCO v2 document segmented corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v2_doc_segmented_doc2query-t5_expansions.msmarco_v2_passage_doc2query-t5_expansions
Dataset Summary
The repo provides queries generated for the MS MARCO v2 passage corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then appended to the… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v2_passage_doc2query-t5_expansions.msmarco_v1_passage_doc2query-t5_expansions
Dataset Summary
The repo provides queries generated for the MS MARCO V1 passage corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then appended to the… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v1_passage_doc2query-t5_expansions.msmarco_v1_doc_doc2query-t5_expansions
Dataset Summary
The repo provides queries generated for the MS MARCO V1 document corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then appended to the… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v1_doc_doc2query-t5_expansions.doc2query-generate-filter
Dataset Card for "triplet-generate-filter-v2"
More Information needed
msmarco_v2_doc_doc2query-t5_expansions
Dataset Summary
The repo provides queries generated for the MS MARCO v2 document corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then appended to the… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v2_doc_doc2query-t5_expansions.msmarco_v1_doc_segmented_doc2query-t5_expansions
Dataset Summary
The repo provides queries generated for the MS MARCO V1 document segmented corpus with docTTTTTquery (sometimes written as docT5query or doc2query-T5), the latest version of the doc2query family of document expansion models. The basic idea is to train a model, that when given an input document, generates questions that the document might answer (or more broadly, queries for which the document might be relevant). These predicted questions (or queries) are then… See the full description on the dataset page: https://huggingface.co/datasets/castorini/msmarco_v1_doc_segmented_doc2query-t5_expansions.doc2query-generate
Dataset Card for "doc2query-generate-v2"
More Information needed
data-doc2query-generate
Dataset Card for "triplet-generate-filter-v3"
More Information needed
sberquad-doc2querydoc2query_scored_queries
Scores of generated queries
This repo contains the scores files pertaining to this study. In particular, we scored the expansion queries generated by T5-based Doc2Query model for MSMARCO-v1 passage dataset and a subset of BEIR benchemark.
We used ELECTRA cross-encoder to get the relevance scores between the document text and its expansion queries. More details in the study repo here.
Structure
All files are .jsonl files with the following three columns per line: ["id"… See the full description on the dataset page: https://huggingface.co/datasets/Watheq/doc2query_scored_queries.doc2query-vector-storems_marco_doc2query
