datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset_cross_encoder_geotechnical_report_v1.0.0
Geotechnical Reports
arxiv-hard-negatives-cross-encoderThis dataset contains hard negative examples generated using cross-encoders for training dense retrieval models.
@misc{reimers2019sentencebertsentenceembeddingsusing,
title={Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks},
author={Nils Reimers and Iryna Gurevych},
year={2019},
eprint={1908.10084},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/1908.10084},
}
msmarco_hard_negatives_cross-encoderThe data was used in the paper Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval.
Hard-Negatives generated by a cross-encoder using 10,000 passages from the MS-Marco dataset
If this dataset was useful consider citing us :)
@misc{sinha2025dontretrievegenerateprompting,
title={Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval},
author={Aarush Sinha},
year={2025},
eprint={2504.21015}… See the full description on the dataset page: https://huggingface.co/datasets/chungimungi/msmarco_hard_negatives_cross-encoder.ms-marco-cross-encoder-hard-negatives
