datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ettin-reranker-v1-data
Ettin Reranker v1 Training Data
This is the training dataset used to produce the cross-encoder/ettin-reranker-{17m,32m,68m,150m,400m,1b}-v1 family of CrossEncoder rerankers. It's a mix of broad-domain text-pair data and retrieval pairs rescored with a strong teacher reranker, with every label produced by an automated scoring system rather than a human annotator.
Structure
Every config has the same three columns:
column
type
description
query
string
The… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/ettin-reranker-v1-data.lightonai-embeddings-fine-tuning-reranked-v1
LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2
This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.text-transformer-5m-cross-encoder-progresscross-encoder-law
Dataset Card for "cross-encoder-law"
More Information needed
msmarco-hard-negatives-cross-encoder-ms-marco-MiniLM-L-6-v2-scoresmsmarco-cross_encoder_ms_marco_minilm_l_12_v2-cross-encoder-scoresclimate-cross-encoder-mixed-neg-v3legal-names-cross-encoder-dataset
📊 مجموعه داده اعتبارسنجی نام شرکتها (ویژه مدلهای Cross Encoder)
✨ معرفی
این دیتاست بیش از 100,000 رکورد است که برای آموزش و ارزیابی مدلهای پردازش زبان طبیعی (NLP) در زمینهی تشخیص پذیرش یا رد نام شرکتها طراحی شده است. این مجموعه داده به طور خاص برای تولید و آموزش مدلهای Cross Encoder طراحی شده است. Cross Encoder نوعی معماری در مدلهای زبانی است که برای وظایف semantic similarity و matching استفاده میشود. در این روش، دو متن (مثلاً نام پیشنهادی و نام ثبتشده) به… See the full description on the dataset page: https://huggingface.co/datasets/intai2070/legal-names-cross-encoder-dataset.arxiv-hard-negatives-cross-encoderThis dataset contains hard negative examples generated using cross-encoders for training dense retrieval models.
@misc{reimers2019sentencebertsentenceembeddingsusing,
title={Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks},
author={Nils Reimers and Iryna Gurevych},
year={2019},
eprint={1908.10084},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/1908.10084},
}
msmarco_hard_negatives_cross-encoderThe data was used in the paper Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval.
Hard-Negatives generated by a cross-encoder using 10,000 passages from the MS-Marco dataset
If this dataset was useful consider citing us :)
@misc{sinha2025dontretrievegenerateprompting,
title={Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval},
author={Aarush Sinha},
year={2025},
eprint={2504.21015}… See the full description on the dataset page: https://huggingface.co/datasets/chungimungi/msmarco_hard_negatives_cross-encoder.dataset_cross_encoder_geotechnical_report_v1.0.0
Geotechnical Reports
ragas-cross-encoder-en-oai-evalms-marco-cross-encoder-hard-negativeskeep_context_cross_encodernev-original-cross-encoder-stsb-roberta-large-bs8-lr2e-05-predclimate-cross-encoder-mixed-neg-v1
Dataset Card for "climate-cross-encoder-mixed-neg-v1"
More Information needed
ragas-cross-encoder-en-short-evalragas-tes-dataset-en-cross-encoderragas-cross-encoder-en-evalclimate-cross-encoder-mixed-neg-v2example_after_cross_encodercross-encoder-binary-context-quesion-v1This is a dataset for training the cross-encoder of our RAG system. It is a combination of the PIAF, FQuAD, and SQuAD-French datasets.
cross-encoder-binary-context-quesion-v2This is a dataset for training the cross-encoder of our RAG system. It is a combination of the PIAF, FQuAD, SQuAD-French, and pandora-s-fr datasets.
ger-dpr-collection-crossencodercross-encoder-binary-context-quesion-v3This is a dataset for training a mixed cross-encoder. The purpose of the cross-encoder is to calculate not only a relevance score between a question and a context (whether the answer to the question can be found in the document or not) but also to calculate a similarity score between two sentences. This dataset is a combination of the PIAF, FQuAD, SQuAD-French, pandora-s-fr, and stsd-fr datasets.
