CoolFace
Datasetpublic

allenai/multixscience_sparse_mean

This is a copy of the Multi-XScience dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The related_work field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "mean", i.e. the number of documents retrieved, k, is set as the mean number of documents seen across examples in this dataset, in this… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multixscience_sparse_mean.

sourceHugging Faceunknownupdated 4y agoView on Hugging Face
1likes34downloads
../
filetest-00000-of-00001-45aa897a9b0b2b15.parquet13.2 MBdownload
filetrain-00000-of-00001-e71278f311282096.parquet72.5 MBdownload
filevalidation-00000-of-00001-666f8f42dff99a14.parquet12.0 MBdownload

allenai/multixscience_sparse_mean · main · files are served by the source, never re-hosted here