cross-encoder
ettin-reranker-v1-data
Ettin Reranker v1 Training Data
This is the training dataset used to produce the cross-encoder/ettin-reranker-{17m,32m,68m,150m,400m,1b}-v1 family of CrossEncoder rerankers. It's a mix of broad-domain text-pair data and retrieval pairs rescored with a strong teacher reranker, with every label produced by an automated scoring system rather than a human annotator.
Structure
Every config has the same three columns:
column
type
description
query
string
The… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/ettin-reranker-v1-data.lightonai-embeddings-fine-tuning-reranked-v1
LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2
This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.text-transformer-5m-cross-encoder-progresscross-encoder-law
Dataset Card for "cross-encoder-law"
More Information needed
msmarco-hard-negatives-cross-encoder-ms-marco-MiniLM-L-6-v2-scoresmsmarco-cross_encoder_ms_marco_minilm_l_12_v2-cross-encoder-scores
