CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abdoelsayed /reranking-datasets 🔥 Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation 🔥 ReRanking Datasets : A comprehensive collection of retrieval and reranking datasets with full passage contexts, including titles, text, and metadata for in-depth research. A curated collection of ready-to-use datasets for retrieval and reranking… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/reranking-datasets.question-answering10M<n<100M6 likes1.8k downloads2y agoHugging Face02mteb /scidocs-rerankingtext1K<n<10K2 likes1.7k downloads4y agoHugging Face03mteb /askubuntudupquestions-rerankingtextn<1K0 likes1.1k downloads4y agoHugging Face04MTEB-BR /quati-reranking QuatiReranking Genuine web-domain reranking: for each of 50 PT-BR queries, rerank the ~100 first-stage candidates (BM25 top-100 over the full 1M-passage Quati Brazilian web corpus, unioned with human-judged passages) under graded relevance 0-3. Candidates include lexical hard negatives, so the task is distinct from first-stage retrieval. Complements JurisTCUReranking (legal) with a web-domain probe. Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark. Task type:… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/quati-reranking.texttext-retrieval1K<n<10K0 likes1k downloads2mo agoHugging Face05MTEB-BR /juristcu-reranking JurisTCUReranking Genuine legal-domain reranking: for each of 150 queries, rerank the ~100 first-stage candidates (BM25 top-100 over the full 16k-document TCU jurisprudence corpus, unioned with human-judged docs) under graded relevance 0-3. Candidates include lexical hard negatives, so the task is distinct from first-stage retrieval. Complements QuatiReranking (web) with a legal-domain reranking probe. Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark. Task… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/juristcu-reranking.texttext-retrieval1K<n<10K0 likes1k downloads2mo agoHugging Face06abdoelsayed /reranking-datasets-light 🔥 Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation 🔥 ReRanking Datasets : A comprehensive collection of retrieval and reranking datasets with full passage contexts, including titles, text, and metadata for in-depth research. A curated collection of ready-to-use datasets for retrieval and reranking… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/reranking-datasets-light.textquestion-answering100K<n<1M3 likes1k downloads2y agoHugging Face07mteb /CMedQAv2-reranking CMedQAv2-reranking An MTEB dataset Massive Text Embedding Benchmark Chinese community medical question answering Task category t2t Domains Medical, Written Reference https://github.com/zhangsheng93/cMedQA2 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CMedQAv2-reranking"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL) evaluator.run(model)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CMedQAv2-reranking.texttext-ranking100K<n<1M0 likes689 downloads1y agoHugging Face08mteb /stackoverflowdupquestions-rerankingtext10K<n<100K3 likes668 downloads4y agoHugging Face09mteb /CMedQAv1-reranking CMedQAv1-reranking An MTEB dataset Massive Text Embedding Benchmark Chinese community medical question answering Task category t2t Domains Medical, Written Reference https://github.com/zhangsheng93/cMedQA How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CMedQAv1-reranking"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL) evaluator.run(model)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CMedQAv1-reranking.texttext-ranking100K<n<1M0 likes664 downloads1y agoHugging Face10ellamind /wikipedia-2023-11-reranking-multilingualThis dataset is derived from Cohere's wikipedia-2023-11 dataset, which is in turn derived from wikimedia/wikipedia. The dataset is licensed under the Creative Commons CC BY-SA 3.0 license. text10K<n<100K7 likes478 downloads2y agoHugging Face11rahulseetharaman /msmarco-llm-reranking-pointwisetext10M<n<100M0 likes316 downloads1y agoHugging Face12rahulseetharaman /msmarco-llm-reranking-pairwisetext10M<n<100M0 likes275 downloads1y agoHugging Face13C-MTEB /CMedQAv2-reranking Dataset Card for "CMedQAv2-reranking" More Information needed text1K<n<10K0 likes206 downloads3y agoHugging Face14C-MTEB /CMedQAv1-reranking Dataset Card for "CMedQAv1-reranking" More Information needed text1K<n<10K0 likes200 downloads3y agoHugging Face15C-MTEB /Mmarco-reranking Dataset Card for "Mmarco-reranking" More Information needed textn<1K2 likes200 downloads3y agoHugging Face16lyon-nlp /mteb-fr-reranking-alloprof-s2p Description This dataset was built upon Alloprof Q&A dataset, negative samples were created using BM25. Please refer to our paper for more details. Citation If you use this dataset in your work, please consider citing: @misc{ciancone2024extending, title={Extending the Massive Text Embedding Benchmark to French}, author={Mathieu Ciancone and Imene Kerboua and Marion Schaeffer and Wissam Siblini}, year={2024}, eprint={2405.20468}… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/mteb-fr-reranking-alloprof-s2p.text10K<n<100K1 likes197 downloads2y agoHugging Face17miracl /mmteb-miracl-reranking0 likes161 downloads2y agoHugging Face18lyon-nlp /mteb-fr-reranking-syntec-s2p Description This dataset was built upon Syntec information retrieval dataset, negative samples were created using BM25. Please refer to our paper for more details. Citation If you use this dataset in your work, please consider citing: @misc{ciancone2024extending, title={Extending the Massive Text Embedding Benchmark to French}, author={Mathieu Ciancone and Imene Kerboua and Marion Schaeffer and Wissam Siblini}, year={2024}, eprint={2405.20468}… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/mteb-fr-reranking-syntec-s2p.textn<1K2 likes137 downloads1y agoHugging Face19tarsur909 /mteb-swe-bench-poly-rerankingtextn<1K0 likes121 downloads1y agoHugging Face20tarsur909 /mteb-swe-bench-multi-rerankingtext1K<n<10K0 likes113 downloads1y agoHugging Face21CronosGhost /code-rerankingtext10K<n<100K2 likes107 downloads3y agoHugging Face22puttatidam /miracl_ind_reranking miracl_ind_reranking Deduplicated copy of kornwtp/miracl_ind_reranking, part of the SEA-BED data-quality work. Source dataset: kornwtp/miracl_ind_reranking Deduplicated on: 2026-09-04 Task type: reranking Splits: dev What changed Kept in this dataset's ORIGINAL schema. Repeats within a row's positive or negative candidate list collapsed, and fully identical (query, positive pool, negative pool) rows collapsed. A document listed as BOTH positive and negative for… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/miracl_ind_reranking.textn<1K0 likes107 downloads8d agoHugging Face23puttatidam /miracl_tha_reranking miracl_tha_reranking Deduplicated copy of kornwtp/miracl_tha_reranking, part of the SEA-BED data-quality work. Source dataset: kornwtp/miracl_tha_reranking Deduplicated on: 2026-09-04 Task type: reranking Splits: dev What changed Kept in this dataset's ORIGINAL schema. Repeats within a row's positive or negative candidate list collapsed, and fully identical (query, positive pool, negative pool) rows collapsed. A document listed as BOTH positive and negative for… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/miracl_tha_reranking.textn<1K0 likes100 downloads8d agoHugging Face24tarsur909 /mteb-swe-bench-multilingual-rerankingtextn<1K0 likes90 downloads1y agoHugging Face25jinaai /nq-reranking-entext1K<n<10K2 likes83 downloads2y agoHugging Face26alina0195 /romteb-grile-rerankingtext1K<n<10K0 likes78 downloads24d agoHugging Face27tarsur909 /mteb-swe-bench-verified-rerankingtextn<1K0 likes72 downloads1y agoHugging Face28FINGU-AI /Rerankingtext1M<n<10M0 likes67 downloads2y agoHugging Face29kornwtp /miracl_tha_rerankingtextn<1K0 likes61 downloads1y agoHugging Face30allenai /sqa_reranking_eval Dataset Details Dataset to evaluate retrieval/reranking models or techniques for scientific QA. The questions are sourced from: Real researchers Stack exchange communities from computing related domains - CS, stats, math, data science Synthetic questions generated by prompting an LLM Each question has passages text in markdown format and the paper Semantic Scholar id, along with a relevance label ranging from 0-3 (higher implies more relevant) obtained from GPT-4o. The label… See the full description on the dataset page: https://huggingface.co/datasets/allenai/sqa_reranking_eval.text1K<n<10K3 likes55 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.