allenai/multinews_sparse_mean
This is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The summary field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "mean", i.e. the number of documents retrieved, k, is set as the mean number of documents seen across examples in this dataset, in this case k==3… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multinews_sparse_mean.
This is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a _sparse_ retriever. The retrieval pipeline used:
- _query_: The
summaryfield of each example - _corpus_: The union of all documents in the
train,validationandtestsplits - _retriever_: BM25 via PyTerrier with default settings
- _top-k strategy_:
"mean", i.e. the number of documents retrieved,k, is set as the mean number of documents seen across examples in this dataset, in this casek==3
Retrieval results on the train set:
Retrieval results on the validation set:
Retrieval results on the test set:
