CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cross-encoder /ettin-reranker-v1-data Ettin Reranker v1 Training Data This is the training dataset used to produce the cross-encoder/ettin-reranker-{17m,32m,68m,150m,400m,1b}-v1 family of CrossEncoder rerankers. It's a mix of broad-domain text-pair data and retrieval pairs rescored with a strong teacher reranker, with every label produced by an automated scoring system rather than a human annotator. Structure Every config has the same three columns: column type description query string The… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/ettin-reranker-v1-data.texttext-ranking100M<n<1B10 likes1.6k downloads29d agoHugging Face02cross-encoder /lightonai-embeddings-fine-tuning-reranked-v1 LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2 This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.texttext-ranking10M<n<100M10 likes704 downloads4mo agoHugging Face03sparse-encoder /msmarco-hard-negatives-cross-encoder-ms-marco-MiniLM-L-6-v2-scorestabular100K<n<1M0 likes55 downloads1y agoHugging Face04CharlesPing /climate-cross-encoder-mixed-neg-v3text10K<n<100K0 likes19 downloads1y agoHugging Face05intai2070 /legal-names-cross-encoder-dataset 📊 مجموعه داده اعتبارسنجی نام شرکت‌ها (ویژه مدل‌های Cross Encoder) ✨ معرفی این دیتاست بیش از 100,000 رکورد است که برای آموزش و ارزیابی مدل‌های پردازش زبان طبیعی (NLP) در زمینه‌ی تشخیص پذیرش یا رد نام شرکت‌ها طراحی شده است. این مجموعه داده به طور خاص برای تولید و آموزش مدل‌های Cross Encoder طراحی شده است. Cross Encoder نوعی معماری در مدل‌های زبانی است که برای وظایف semantic similarity و matching استفاده می‌شود. در این روش، دو متن (مثلاً نام پیشنهادی و نام ثبت‌شده) به… See the full description on the dataset page: https://huggingface.co/datasets/intai2070/legal-names-cross-encoder-dataset.texttext-classification100K<n<1M0 likes17 downloads10mo agoHugging Face06Hyukkyu /msmarco-cross_encoder_ms_marco_minilm_l_12_v2-cross-encoder-scorestext100K<n<1M0 likes17 downloads8mo agoHugging Face07lock-rr /ragas-cross-encoder-en-oai-evaltabularn<1K0 likes9 downloads2y agoHugging Face08nc33 /keep_context_cross_encodertabular100K<n<1M0 likes8 downloads3y agoHugging Face09mhr2004 /nev-original-cross-encoder-stsb-roberta-large-bs8-lr2e-05-predtabular1K<n<10K0 likes8 downloads1y agoHugging Face10CharlesPing /climate-cross-encoder-mixed-neg-v1 Dataset Card for "climate-cross-encoder-mixed-neg-v1" More Information needed text10K<n<100K0 likes6 downloads1y agoHugging Face11lock-rr /ragas-tes-dataset-en-cross-encodertextn<1K0 likes5 downloads2y agoHugging Face12lock-rr /ragas-cross-encoder-en-short-evaltabularn<1K0 likes5 downloads2y agoHugging Face13lock-rr /ragas-cross-encoder-en-evaltabularn<1K0 likes4 downloads2y agoHugging Face14CharlesPing /climate-cross-encoder-mixed-neg-v2text10K<n<100K0 likes4 downloads1y agoHugging Face15SeppeV /example_after_cross_encodertextn<1K0 likes4 downloads1y agoHugging Face16LeviatanAIResearch /cross-encoder-binary-context-quesion-v2gatedThis is a dataset for training the cross-encoder of our RAG system. It is a combination of the PIAF, FQuAD, SQuAD-French, and pandora-s-fr datasets. text100K<n<1M0 likes2 downloads2y agoHugging Face17samheym /ger-dpr-collection-crossencodergatedtabular10M<n<100M0 likes2 downloads2y agoHugging Face18LeviatanAIResearch /cross-encoder-binary-context-quesion-v3gatedThis is a dataset for training a mixed cross-encoder. The purpose of the cross-encoder is to calculate not only a relevance score between a question and a context (whether the answer to the question can be found in the document or not) but also to calculate a similarity score between two sentences. This dataset is a combination of the PIAF, FQuAD, SQuAD-French, pandora-s-fr, and stsd-fr datasets. texttext-classification100K<n<1M0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.