CoolFace
Datasetpublic

allenai/multinews_sparse_max

This is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The summary field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of documents seen across examples in this dataset, in this case… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multinews_sparse_max.

sourceHugging Faceotherupdated 4y agoView on Hugging Face
0likes73downloads
Dataset Card

This is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a _sparse_ retriever. The retrieval pipeline used:

  • _query_: The summary field of each example
  • _corpus_: The union of all documents in the train, validation and test splits
  • _retriever_: BM25 via PyTerrier with default settings
  • _top-k strategy_: "max", i.e. the number of documents retrieved, k, is set as the maximum number of documents seen across examples in this dataset, in this case k==10

Retrieval results on the train set:

Recall@100RprecPrecision@kRecall@k
0.87930.74600.22130.8264

Retrieval results on the validation set:

Recall@100RprecPrecision@kRecall@k
0.87480.74530.21730.8232

Retrieval results on the test set:

Recall@100RprecPrecision@kRecall@k
0.87750.74800.21870.8250