CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cross-encoder /ettin-reranker-v1-data Ettin Reranker v1 Training Data This is the training dataset used to produce the cross-encoder/ettin-reranker-{17m,32m,68m,150m,400m,1b}-v1 family of CrossEncoder rerankers. It's a mix of broad-domain text-pair data and retrieval pairs rescored with a strong teacher reranker, with every label produced by an automated scoring system rather than a human annotator. Structure Every config has the same three columns: column type description query string The… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/ettin-reranker-v1-data.texttext-ranking100M<n<1B10 likes1.6k downloads26d agoHugging Face02cross-encoder /lightonai-embeddings-fine-tuning-reranked-v1 LightOn embeddings-fine-tuning, rescored with mxbai-rerank-large-v2 This dataset is a teacher-rescored version of lightonai/embeddings-fine-tuning. For every (query, candidate-document) pair in the source, we ran mixedbread-ai/mxbai-rerank-large-v2 and stored the resulting score. The point is to make the source data usable as a teacher target for distilling reranker students. It's the upstream artifact behind the rerank-scored configs of cross-encoder/ettin-reranker-v1-data… See the full description on the dataset page: https://huggingface.co/datasets/cross-encoder/lightonai-embeddings-fine-tuning-reranked-v1.texttext-ranking10M<n<100M10 likes619 downloads4mo agoHugging Face03ArUn123123 /text-transformer-5m-cross-encoder-progress0 likes172 downloads11d agoHugging Face04nc33 /cross-encoder-law Dataset Card for "cross-encoder-law" More Information needed tabular1M<n<10M1 likes64 downloads3y agoHugging Face05sparse-encoder /msmarco-hard-negatives-cross-encoder-ms-marco-MiniLM-L-6-v2-scorestabular100K<n<1M0 likes53 downloads1y agoHugging Face06Hyukkyu /msmarco-cross_encoder_ms_marco_minilm_l_12_v2-cross-encoder-scorestext100K<n<1M0 likes20 downloads8mo agoHugging Face07CharlesPing /climate-cross-encoder-mixed-neg-v3text10K<n<100K0 likes19 downloads1y agoHugging Face08intai2070 /legal-names-cross-encoder-dataset 📊 مجموعه داده اعتبارسنجی نام شرکت‌ها (ویژه مدل‌های Cross Encoder) ✨ معرفی این دیتاست بیش از 100,000 رکورد است که برای آموزش و ارزیابی مدل‌های پردازش زبان طبیعی (NLP) در زمینه‌ی تشخیص پذیرش یا رد نام شرکت‌ها طراحی شده است. این مجموعه داده به طور خاص برای تولید و آموزش مدل‌های Cross Encoder طراحی شده است. Cross Encoder نوعی معماری در مدل‌های زبانی است که برای وظایف semantic similarity و matching استفاده می‌شود. در این روش، دو متن (مثلاً نام پیشنهادی و نام ثبت‌شده) به… See the full description on the dataset page: https://huggingface.co/datasets/intai2070/legal-names-cross-encoder-dataset.texttext-classification100K<n<1M0 likes17 downloads10mo agoHugging Face09chungimungi /arxiv-hard-negatives-cross-encoderThis dataset contains hard negative examples generated using cross-encoders for training dense retrieval models. @misc{reimers2019sentencebertsentenceembeddingsusing, title={Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks}, author={Nils Reimers and Iryna Gurevych}, year={2019}, eprint={1908.10084}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/1908.10084}, } texttext-ranking1K<n<10K0 likes15 downloads9mo agoHugging Face10chungimungi /msmarco_hard_negatives_cross-encoderThe data was used in the paper Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval. Hard-Negatives generated by a cross-encoder using 10,000 passages from the MS-Marco dataset If this dataset was useful consider citing us :) @misc{sinha2025dontretrievegenerateprompting, title={Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval}, author={Aarush Sinha}, year={2025}, eprint={2504.21015}… See the full description on the dataset page: https://huggingface.co/datasets/chungimungi/msmarco_hard_negatives_cross-encoder.texttext-ranking10K<n<100K0 likes13 downloads9mo agoHugging Face11GbrlOl /dataset_cross_encoder_geotechnical_report_v1.0.0 Geotechnical Reports text1K<n<10K0 likes10 downloads1y agoHugging Face12lock-rr /ragas-cross-encoder-en-oai-evaltabularn<1K0 likes9 downloads2y agoHugging Face13chungimungi /ms-marco-cross-encoder-hard-negativestext100K<n<1M0 likes9 downloads5mo agoHugging Face14nc33 /keep_context_cross_encodertabular100K<n<1M0 likes8 downloads3y agoHugging Face15mhr2004 /nev-original-cross-encoder-stsb-roberta-large-bs8-lr2e-05-predtabular1K<n<10K0 likes8 downloads1y agoHugging Face16CharlesPing /climate-cross-encoder-mixed-neg-v1 Dataset Card for "climate-cross-encoder-mixed-neg-v1" More Information needed text10K<n<100K0 likes6 downloads1y agoHugging Face17lock-rr /ragas-cross-encoder-en-short-evaltabularn<1K0 likes5 downloads2y agoHugging Face18lock-rr /ragas-tes-dataset-en-cross-encodertextn<1K0 likes4 downloads2y agoHugging Face19lock-rr /ragas-cross-encoder-en-evaltabularn<1K0 likes4 downloads2y agoHugging Face20CharlesPing /climate-cross-encoder-mixed-neg-v2text10K<n<100K0 likes4 downloads1y agoHugging Face21SeppeV /example_after_cross_encodertextn<1K0 likes4 downloads1y agoHugging Face22LeviatanAIResearch /cross-encoder-binary-context-quesion-v1gatedThis is a dataset for training the cross-encoder of our RAG system. It is a combination of the PIAF, FQuAD, and SQuAD-French datasets. texttext-classification100K<n<1M0 likes2 downloads2y agoHugging Face23LeviatanAIResearch /cross-encoder-binary-context-quesion-v2gatedThis is a dataset for training the cross-encoder of our RAG system. It is a combination of the PIAF, FQuAD, SQuAD-French, and pandora-s-fr datasets. text100K<n<1M0 likes2 downloads2y agoHugging Face24samheym /ger-dpr-collection-crossencodergatedtabular10M<n<100M0 likes2 downloads2y agoHugging Face25LeviatanAIResearch /cross-encoder-binary-context-quesion-v3gatedThis is a dataset for training a mixed cross-encoder. The purpose of the cross-encoder is to calculate not only a relevance score between a question and a context (whether the answer to the question can be found in the document or not) but also to calculate a similarity score between two sentences. This dataset is a combination of the PIAF, FQuAD, SQuAD-French, pandora-s-fr, and stsd-fr datasets. texttext-classification100K<n<1M0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.