allenai/multinews_sparse_max
This is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The summary field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of documents seen across examples in this dataset, in this case… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multinews_sparse_max.
This is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a _sparse_ retriever. The retrieval pipeline used:
- _query_: The
summaryfield of each example - _corpus_: The union of all documents in the
train,validationandtestsplits - _retriever_: BM25 via PyTerrier with default settings
- _top-k strategy_:
"max", i.e. the number of documents retrieved,k, is set as the maximum number of documents seen across examples in this dataset, in this casek==10
Retrieval results on the train set:
Retrieval results on the validation set:
Retrieval results on the test set:
