datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ettin-reranker-v1-data
Ettin Reranker v1 Training Data
This is the training dataset used to produce the cross-encoder/ettin-reranker-{17m,32m,68m,150m,400m,1b}-v1 family of CrossEncoder rerankers. It's a mix of broad-domain text-pair data and retrieval pairs rescored with a strong teacher reranker, with every label produced by an automated scoring system rather than a human annotator.
Structure
Every config has the same three columns:
column
type
description
query
string
The… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/ettin-reranker-v1-data.lightonai-embeddings-fine-tuning-reranked-v1
LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2
This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.msmarco-hard-negatives-cross-encoder-ms-marco-MiniLM-L-6-v2-scoresclimate-cross-encoder-mixed-neg-v3legal-names-cross-encoder-dataset
📊 مجموعه داده اعتبارسنجی نام شرکتها (ویژه مدلهای Cross Encoder)
✨ معرفی
این دیتاست بیش از 100,000 رکورد است که برای آموزش و ارزیابی مدلهای پردازش زبان طبیعی (NLP) در زمینهی تشخیص پذیرش یا رد نام شرکتها طراحی شده است. این مجموعه داده به طور خاص برای تولید و آموزش مدلهای Cross Encoder طراحی شده است. Cross Encoder نوعی معماری در مدلهای زبانی است که برای وظایف semantic similarity و matching استفاده میشود. در این روش، دو متن (مثلاً نام پیشنهادی و نام ثبتشده) به… See the full description on the dataset page: https://huggingface.co/datasets/intai2070/legal-names-cross-encoder-dataset.msmarco-cross_encoder_ms_marco_minilm_l_12_v2-cross-encoder-scoresragas-cross-encoder-en-oai-evalkeep_context_cross_encodernev-original-cross-encoder-stsb-roberta-large-bs8-lr2e-05-predclimate-cross-encoder-mixed-neg-v1
Dataset Card for "climate-cross-encoder-mixed-neg-v1"
More Information needed
ragas-tes-dataset-en-cross-encoderragas-cross-encoder-en-short-evalragas-cross-encoder-en-evalclimate-cross-encoder-mixed-neg-v2example_after_cross_encodercross-encoder-binary-context-quesion-v2This is a dataset for training the cross-encoder of our RAG system. It is a combination of the PIAF, FQuAD, SQuAD-French, and pandora-s-fr datasets.
ger-dpr-collection-crossencodercross-encoder-binary-context-quesion-v3This is a dataset for training a mixed cross-encoder. The purpose of the cross-encoder is to calculate not only a relevance score between a question and a context (whether the answer to the question can be found in the document or not) but also to calculate a similarity score between two sentences. This dataset is a combination of the PIAF, FQuAD, SQuAD-French, pandora-s-fr, and stsd-fr datasets.
