CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PaDaS-Lab /webfaq-retrievalWebFAQ Retrieval Dataset Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Retrieval Dataset is a carefully filtered and curated subset of the broader WebFAQ Q&A Dataset.It is purpose-built for Information Retrieval (IR) tasks, such as training and evaluating dense or sparse retrieval models in multiple languages. Each of the… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval.texttext-retrieval10M<n<100M10 likes6.4k downloads1y agoHugging Face02ai-forever /rubq-retrievaltexttext-retrieval10K<n<100K2 likes1.3k downloads2y agoHugging Face03isaacus /contractual-clause-retrieval Contractual Clause Retrieval 📑 Contractual Clause Retrieval by Isaacus is a challenging legal information retrieval evaluation dataset consisting of 45 unique types of contractual clauses paired with 2 highly representative examples of each, resulting in 90 pairings. This dataset is intended to stress test the ability of information retrieval, zero-shot classification, and NLI models to identify a broad range of common types of contractual clauses based solely on their definition… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/contractual-clause-retrieval.texttext-retrievaln<1K5 likes814 downloads11mo agoHugging Face04friedrichor /Unite-Base-Retrieval-Train Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Statistics Accessing Images and Videos 2025-06-19: We've updated the compressed archives for all image and video files to enable faster extraction.If you've already downloaded the previous files, there's no need to redownload them — the content remains exactly the same. The only difference lies in the compression method, which now allows for quicker… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/Unite-Base-Retrieval-Train.textfeature-extraction1M<n<10M0 likes716 downloads1y agoHugging Face05isaacus /australian-tax-guidance-retrieval Australian Tax Guidance Retrieval 🏦 Australian Tax Guidance Retrieval by Isaacus is a novel, diverse, and challenging legal information retrieval evaluation dataset consisting of 112 real-life Australian tax law questions paired with expert-annotated, relevant Australian Government tax guidance and policies. Uniquely, this dataset sources its real-life tax questions from the posts of everyday Australian taxpayers on the ATO Community forum, with relevant Australian Government… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/australian-tax-guidance-retrieval.texttext-retrievaln<1K6 likes629 downloads11mo agoHugging Face06isaacus /license-tldr-retrieval License TL;DR Retrieval 📑 License TL;DR Retrieval by Isaacus is a challenging legal information retrieval evaluation dataset consisting of 65 summary-license pairs sourced from tl;drLegal. This dataset is intended to stress test the ability of an information retrieval model to match relevant open source licenses with summaries of their terms. This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB), the largest, most diverse, and most comprehensive benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/license-tldr-retrieval.texttext-retrievaln<1K2 likes598 downloads11mo agoHugging Face07isaacus /echr-retrieval ECHR Retrieval 🏛️ ECHR Retrieval by Isaacus is a challenging legal information retrieval evaluation dataset consisting of 200 short summaries of findings of European Court of Human Rights decisions paired with the text of those decisions sourced from the HUDOC database. This dataset is intended to stress test the ability of an information retrieval model to retrieve relevant court decisions given arbitrary legal holdings. This dataset forms part of the Massive Legal Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/echr-retrieval.texttext-retrievaln<1K2 likes590 downloads7mo agoHugging Face08isaacus /gdpr-holdings-retrieval GDPR Holdings Retrieval 🔐 GDPR Holdings Retrieval by Isaacus is a novel and challenging legal information retrieval evaluation dataset consisting of 500 fact patterns paired with holdings in European regulatory and court decisions. This dataset is intended to stress test the ability of an information retrieval model to retrieve relevant judicial and regulatory decisions given arbitrary fact patterns. This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB), the… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/gdpr-holdings-retrieval.texttext-retrieval1K<n<10K4 likes457 downloads11mo agoHugging Face09rmems /rag-retrieval-debug-trajectories Rag Retrieval Debug Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/rag-retrieval-debug-trajectories.text1K<n<10K0 likes400 downloads16h agoHugging Face10stanfordnlp /concurrentqa-retrievalConcurrentQA is a textual multi-hop QA benchmark to require concurrent retrieval over multiple data-distributions (i.e. Wikipedia and email data). This dataset was constructed by researchers at Stanford and FAIR, following the data collection process and schema of HotpotQA. This benchmark can be used to study generalization in retrieval as well as privacy when reasoning across multiple privacy scopes --- i.e. public Wikipedia documents and private emails. This dataset is for the Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/concurrentqa-retrieval.textquestion-answering10K<n<100K4 likes344 downloads2y agoHugging Face11lyon-nlp /mteb-fr-retrieval-syntec-s2p Syntec dataset for information retrieval This dataset has been built from the Syntec Collective bargaining agreement. Its purpose is information retrieval. Dataset Details The dataset is rather small. It is intended to be used only as a test set, for fast evaluation of models. It is split into 2 subsets : queries : it features 100 manually created questions. Each question is mapped to the article that contains the answer. documents : corresponds to the 90 articles from… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/mteb-fr-retrieval-syntec-s2p.textquestion-answeringn<1K2 likes233 downloads2y agoHugging Face12zxliu /ReAPR-Automatic-Program-Repair-via-Retrieval-Augmented-Large-Language-ModelsThis is the Retrieval dataset used in the paper "ReAPR: Automatic Program Repair via Retrieval-Augmented Large Language Models" text100K<n<1M3 likes162 downloads2y agoHugging Face13alekseevpavel04 /ru-law-retrieval RuLawRetrieval: поиск статей законов РФ по вопросам Бенчмарк поиска (retrieval) по кодексам РФ в формате MTEB (corpus / queries / qrels) со сплитами train / dev / test и подмножеством test с повторной разметкой и множественной релевантностью (golden, размечено ИИ-агентом). Сделан в проекте ru-law-retrieval, где на нём сравниваются 12 готовых эмбеддеров, BM25 и дообученная multilingual-e5-small. English summary: a Russian legal retrieval benchmark (MTEB format). Corpus: 3,786… See the full description on the dataset page: https://huggingface.co/datasets/alekseevpavel04/ru-law-retrieval.texttext-retrieval10K<n<100K0 likes160 downloads3d agoHugging Face14CMiller /kbmill-brick-retrieval KBMill Brick Retrieval Demos Portable, residual-honest knowledge bricks turned into retrieval evaluation sets for KBMill — the public mill at kbmill.com. Shelf packages live in kbmill-brick-library. These are not unbounded wiki dumps or synthetic QA. They come from real KBMill manufacturing: bounded packages with muted residual junk filtered where applicable, craft notes, security report in the ZIP, and published, re-runnable cosine retrieval evidence. Config Queries… See the full description on the dataset page: https://huggingface.co/datasets/CMiller/kbmill-brick-retrieval.tabulartext-retrieval1K<n<10K0 likes132 downloads24d agoHugging Face15issai /kazqad-retrievalgated Dataset Card for KazQAD-Retrieval Dataset Summary KazQAD is a Kazakh open-domain Question Answering Dataset that can be used in both reading comprehension and full ODQA settings, as well as for information retrieval experiments. This repository contains only the collection and relevance judgments for information retrieval task. Short answers and data for the reading comprehension task (extractive QA) can be found here. KazQAD contains just under 6,000 unique questions and… See the full description on the dataset page: https://huggingface.co/datasets/issai/kazqad-retrieval.texttext-retrieval100K<n<1M3 likes119 downloads2y agoHugging Face16friedrichor /Unite-Instruct-Retrieval-Train Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval Statistics Accessing Images and Videos 2025-06-19: We've updated the compressed archives for all image and video files to enable faster extraction.If you've already downloaded the previous files, there's no need to redownload them — the content remains exactly the same. The only difference lies in the compression method, which now allows for quicker… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/Unite-Instruct-Retrieval-Train.textfeature-extraction1M<n<10M1 likes116 downloads1y agoHugging Face17ai-forever /ria-news-retrievaltexttext-retrieval100K<n<1M1 likes114 downloads2y agoHugging Face18TUKE-KEMT /retrieval-skquad Dataset Card for retrieval-skquad Dataset Summary STS SK-QuAD Retrieval is a unique dataset designed to evaluate Slovak search performance using metrics like MRR, MAP, and NDCG, derived from the SK-QuAD dataset. It features questions and answers sourced from a search engine before annotation. The annotated data assigns categories to the best answers for each question, enhancing Slovak language search evaluation. This dataset is a significant step forward in the… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/retrieval-skquad.texttext-retrieval10K<n<100K0 likes113 downloads2y agoHugging Face19andersonbcdefg /synthetic_retrieval_tasksSynthetic data designed as prompts for generating embeddings training data for retrieval. The "iteration" column refers to how the data was generated. Iteration 1: Use the following pool of seed tasks, prompt GPT-3.5-Turbo to generate additional tasks. RETRIEVAL_EXAMPLES = [ 'Provide a scientific claim as query, retrieve documents that help verify or refute the claim.', 'Search for documents that answers a FAQ-style query on children\'s nutrition.', "Retrieve company's financial reports… See the full description on the dataset page: https://huggingface.co/datasets/andersonbcdefg/synthetic_retrieval_tasks.text100K<n<1M79 likes88 downloads3y agoHugging Face20CZLC /history_retrieval Introduction The HistoryIR dataset was annotated on top of the historical part of BUT-LCC corpus. We urged annotators to search for historical events (from their own mind, or using our inspirator, more details in the upcoming paper), using the semantic search tool we developed (translation service + English contriever model setup). Then the annotators annotated top retrieved passages as relevant or irrelevant. We've done additional filtering step that included manual verification of… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/history_retrieval.text1K<n<10K0 likes85 downloads2y agoHugging Face21MCINext /synthetic-persian-qa-retrieval Dataset Summary Synthetic Persian QA Retrieval (SynPerQARetrieval) is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on question answering. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using the GPT-4o-mini Large Language Model. It consists of question-answer pairs derived from the content of various curated Persian websites. The primary task is to retrieve the correct answer… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-qa-retrieval.text100K<n<1M0 likes84 downloads1y agoHugging Face22anonymousresearch123 /deka-retrieval Loading from datasets import load_dataset corpus = load_dataset("anonymousresearch123/deka-retrieval", "corpus", split="train") queries = load_dataset("anonymousresearch123/deka-retrieval", "queries", split="train") labels = load_dataset("anonymousresearch123/deka-retrieval", "labels", split="train") `` texttext-retrieval10K<n<100K0 likes83 downloads3d agoHugging Face23baobab-trees /wikipedia-human-retrieval-ja Japanese Wikipedia Human Retrieval dataset This is a Japanese question answereing dataset with retrieval on Wikipedia articles by trained human workers. Contributors Yusuke Oda defined the dataset specification, data structure, and the scheme of data collection. Baobab, Inc. operated data collection, data checking, and formatting. About the dataset Each entry represents a single QA session: given a question sentence, the responsible worker tried to search for… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/wikipedia-human-retrieval-ja.textquestion-answering1K<n<10K34 likes81 downloads3y agoHugging Face24MCINext /synthetic-persian-chatbot-summary-retrieval Dataset Summary Synthetic Persian Chatbot Summary Retrieval (SynPerChatbotSumSRetrieval) is a Persian (Farsi) dataset designed for the novel Summary Retrieval task, introduced as part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using the GPT-4o-mini Large Language Model and is derived from the Synthetic Persian Chatbot Dataset. The core task is to retrieve the correct, pre-generated summary that corresponds to a given user–chatbot… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-summary-retrieval.text10K<n<100K0 likes81 downloads1y agoHugging Face25MCINext /synthetic-persian-chatbot-rag-summary-retrieval Dataset Summary Synthetic Persian Chatbot RAG Summary Retrieval (SynPerChatbotRAGSumSRetrieval) is a Persian (Farsi) dataset for the Summary Retrieval task, specifically built for Retrieval-Augmented Generation (RAG) systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the Synthetic Persian Chatbot RAG Dataset. It evaluates the ability of models to match conversations—possibly… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-summary-retrieval.text1K<n<10K0 likes79 downloads1y agoHugging Face26MCINext /synthetic-persian-chatbot-rag-faq-retrieval Dataset Summary Synthetic Persian Chatbot RAG FAQ Retrieval (SynPerChatbotRAGFAQRetrieval) is a Persian (Farsi) dataset built for the Retrieval task in Retrieval-Augmented Generation (RAG)-based chatbot systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini. The dataset is designed to evaluate how well models retrieve relevant FAQ entries based on a user's message and prior conversation context. Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-faq-retrieval.text10K<n<100K0 likes79 downloads1y agoHugging Face27grimulkan /passkey-retrievalPasskey retrieval training/evaluation data in Fastchat format. You will have to split into train/evaluation manually. Articles were drawn from Long C4 in varying lengths A secret passkey was inserted somewhere in the article, randomly. The name and type of secret is randomly varied (passphrase, secret key, specific fact, favorite colors, password, etc.) and the passkey itself was randomly generated based on various proper nouns (Faker Library), words/phrases of varying lengths (WonderWords… See the full description on the dataset page: https://huggingface.co/datasets/grimulkan/passkey-retrieval.text10K<n<100K1 likes78 downloads3y agoHugging Face28MCINext /persian-web-document-retrieval Dataset Summary Persian Web Document Retrieval is a Persian (Farsi) dataset designed for the Retrieval task. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset consists of real-world queries collected from the Zarrebin search engine and web documents labeled by humans for relevance. It is curated to evaluate model performance in web search scenarios. Language(s): Persian (Farsi) Task(s): Retrieval (Web Search) Source: Collected from Zarrebin… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/persian-web-document-retrieval.text100K<n<1M0 likes77 downloads1y agoHugging Face29intfloat /personalized_passkey_retrieval Dataset Summary This dataset contains the data for personalized passkey retrieval task in the paper Improving Text Embeddings with Large Language Models. Data Fields query: a string feature. candidates: List of string feature, 100 candidates for each query. label: a int32 feature, the index of the correct candidate in the candidates list, always 0. context_length: a int32 feature, the approximate length for the candidate documents. How to use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/intfloat/personalized_passkey_retrieval.tabularn<1K10 likes72 downloads3y agoHugging Face30roshbeed /ai-residency-vector-search-retrieval-datatextn<1K0 likes71 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.