CoolFace
Datasetpublic

allenai/multixscience_sparse_mean

This is a copy of the Multi-XScience dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The related_work field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "mean", i.e. the number of documents retrieved, k, is set as the mean number of documents seen across examples in this dataset, in this… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multixscience_sparse_mean.

sourceHugging Faceunknownupdated 4y agoView on Hugging Face
1likes35downloads
Dataset Card

This is a copy of the Multi-XScience dataset, except the input source documents of its test split have been replaced by a _sparse_ retriever. The retrieval pipeline used:

  • _query: The `relatedwork` field of each example
  • _corpus_: The union of all documents in the train, validation and test splits
  • _retriever_: BM25 via PyTerrier with default settings
  • _top-k strategy_: "mean", i.e. the number of documents retrieved, k, is set as the mean number of documents seen across examples in this dataset, in this case k==4

Retrieval results on the train set:

Recall@100RprecPrecision@kRecall@k
0.54820.22430.15780.2689

Retrieval results on the validation set:

Recall@100RprecPrecision@kRecall@k
0.54760.22090.15920.2650

Retrieval results on the test set:

Recall@100RprecPrecision@kRecall@k
0.5480.22720.16110.2704