CoolFace
Datasetpublic

kiyam/ddro-msmarco-doc-dataset-300k

DDRO — MS MARCO Top-300K Processed Dataset This dataset contains the preprocessed MS MARCO Top-300K document corpus used to train and evaluate the DDRO generative retrieval models from: Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval (SIGIR 2025) Files File Description Size msmarco-docs-sents.top.300k.json Top-300K documents selected by click frequency, with sentence tokenization (JSONL format) ~2 GB… See the full description on the dataset page: https://huggingface.co/datasets/kiyam/ddro-msmarco-doc-dataset-300k.

sourceHugging Facems-plupdated 4mo agoView on Hugging Face
0likes23downloads
5 commits on main
2da4a3d4mo ago

update readme

kiyam
0725cd81y ago

Upload msmarco-docs-sents.top.300k.json

kiyam
35365a91y ago

Update README.md

kiyam
c39a1b31y ago

Upload README.md

kiyam
811489e1y ago

initial commit

kiyam