kiyam/ddro-msmarco-doc-dataset-300k
DDRO — MS MARCO Top-300K Processed Dataset This dataset contains the preprocessed MS MARCO Top-300K document corpus used to train and evaluate the DDRO generative retrieval models from: Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval (SIGIR 2025) Files File Description Size msmarco-docs-sents.top.300k.json Top-300K documents selected by click frequency, with sentence tokenization (JSONL format) ~2 GB… See the full description on the dataset page: https://huggingface.co/datasets/kiyam/ddro-msmarco-doc-dataset-300k.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face