chungimungi/msmarco_hard_negatives_cross-encoder
The data was used in the paper Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval. Hard-Negatives generated by a cross-encoder using 10,000 passages from the MS-Marco dataset If this dataset was useful consider citing us :) @misc{sinha2025dontretrievegenerateprompting, title={Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval}, author={Aarush Sinha}, year={2025}, eprint={2504.21015}… See the full description on the dataset page: https://huggingface.co/datasets/chungimungi/msmarco_hard_negatives_cross-encoder.
This repository belongs to chungimungi on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
