datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-item-id-hardneg-100shot-v4_128krank1-R1-MSMARCO
rank1-R1-MSMARCO: Reasoning Outputs from MS MARCO Dataset
📄 Paper | 🚀 GitHub Repository
This dataset contains outputs from Deepseek's R1 model on the MS MARCO passage dataset, used to train rank1. It showcases the reasoning chains and relevance judgments generated when determining document relevance for information retrieval queries.
Dataset Description
The rank1-R1-MSMARCO dataset consists of reasoning chains and relevance judgments produced on the MS MARCO passage… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/rank1-R1-MSMARCO.MSMarco_Negative_1k
MS MARCO Negative 1k
This dataset contains 1,000 random examples sampled from microsoft/ms_marco with added negative_query and generated negative_ans columns.
Source dataset: microsoft/ms_marco
Source subset/split: v1.1/train
Document used for negative query generation: first selected passage when available, otherwise first non-empty passage
Negative query types: 500 explicit_negation, 500 antonym
Negative answer generation model: gpt-4o
Rows written: 1000
Destination repo:… See the full description on the dataset page: https://huggingface.co/datasets/canho/MSMarco_Negative_1k.
