datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webfaq-retrievalWebFAQ Retrieval Dataset
Overview |
Details |
Structure |
Examples |
Considerations |
License |
Citation |
Contact |
Acknowledgement
Overview
The WebFAQ Retrieval Dataset is a carefully filtered and curated subset of the broader WebFAQ Q&A Dataset.It is purpose-built for Information Retrieval (IR) tasks, such as training and evaluating dense or sparse retrieval models in multiple languages.
Each of the… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval.rubq-retrievalcontractual-clause-retrieval
Contractual Clause Retrieval 📑
Contractual Clause Retrieval by Isaacus is a challenging legal information retrieval evaluation dataset consisting of 45 unique types of contractual clauses paired with 2 highly representative examples of each, resulting in 90 pairings.
This dataset is intended to stress test the ability of information retrieval, zero-shot classification, and NLI models to identify a broad range of common types of contractual clauses based solely on their definition… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/contractual-clause-retrieval.Unite-Base-Retrieval-Train
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Statistics
Accessing Images and Videos
2025-06-19: We've updated the compressed archives for all image and video files to enable faster extraction.If you've already downloaded the previous files, there's no need to redownload them — the content remains exactly the same. The only difference lies in the compression method, which now allows for quicker… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/Unite-Base-Retrieval-Train.australian-tax-guidance-retrieval
Australian Tax Guidance Retrieval 🏦
Australian Tax Guidance Retrieval by Isaacus is a novel, diverse, and challenging legal information retrieval evaluation dataset consisting of 112 real-life Australian tax law questions paired with expert-annotated, relevant Australian Government tax guidance and policies.
Uniquely, this dataset sources its real-life tax questions from the posts of everyday Australian taxpayers on the ATO Community forum, with relevant Australian Government… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/australian-tax-guidance-retrieval.license-tldr-retrieval
License TL;DR Retrieval 📑
License TL;DR Retrieval by Isaacus is a challenging legal information retrieval evaluation dataset consisting of 65 summary-license pairs sourced from tl;drLegal.
This dataset is intended to stress test the ability of an information retrieval model to match relevant open source licenses with summaries of their terms.
This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB), the largest, most diverse, and most comprehensive benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/license-tldr-retrieval.echr-retrieval
ECHR Retrieval 🏛️
ECHR Retrieval by Isaacus is a challenging legal information retrieval evaluation dataset consisting of 200 short summaries of findings of European Court of Human Rights decisions paired with the text of those decisions sourced from the HUDOC database.
This dataset is intended to stress test the ability of an information retrieval model to retrieve relevant court decisions given arbitrary legal holdings.
This dataset forms part of the Massive Legal Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/echr-retrieval.gdpr-holdings-retrieval
GDPR Holdings Retrieval 🔐
GDPR Holdings Retrieval by Isaacus is a novel and challenging legal information retrieval evaluation dataset consisting of 500 fact patterns paired with holdings in European regulatory and court decisions.
This dataset is intended to stress test the ability of an information retrieval model to retrieve relevant judicial and regulatory decisions given arbitrary fact patterns.
This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB), the… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/gdpr-holdings-retrieval.rag-retrieval-debug-trajectories
Rag Retrieval Debug Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/rag-retrieval-debug-trajectories.concurrentqa-retrievalConcurrentQA is a textual multi-hop QA benchmark to require concurrent retrieval over multiple data-distributions (i.e. Wikipedia and email data). This dataset was constructed by researchers at Stanford and FAIR, following the data collection process and schema of HotpotQA. This benchmark can be used to study generalization in retrieval as well as privacy when reasoning across multiple privacy scopes --- i.e. public Wikipedia documents and private emails.
This dataset is for the Retrieval… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/concurrentqa-retrieval.mteb-fr-retrieval-syntec-s2p
Syntec dataset for information retrieval
This dataset has been built from the Syntec Collective bargaining agreement. Its purpose is information retrieval.
Dataset Details
The dataset is rather small. It is intended to be used only as a test set, for fast evaluation of models.
It is split into 2 subsets :
queries : it features 100 manually created questions. Each question is mapped to the article that contains the answer.
documents : corresponds to the 90 articles from… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/mteb-fr-retrieval-syntec-s2p.ReAPR-Automatic-Program-Repair-via-Retrieval-Augmented-Large-Language-ModelsThis is the Retrieval dataset used in the paper "ReAPR: Automatic Program Repair via Retrieval-Augmented Large Language Models"
ru-law-retrieval
RuLawRetrieval: поиск статей законов РФ по вопросам
Бенчмарк поиска (retrieval) по кодексам РФ в формате MTEB (corpus / queries / qrels) со сплитами train / dev / test и подмножеством test с повторной разметкой и множественной релевантностью (golden, размечено ИИ-агентом). Сделан в проекте ru-law-retrieval, где на нём сравниваются 12 готовых эмбеддеров, BM25 и дообученная multilingual-e5-small.
English summary: a Russian legal retrieval benchmark (MTEB format). Corpus: 3,786… See the full description on the dataset page: https://huggingface.co/datasets/alekseevpavel04/ru-law-retrieval.kbmill-brick-retrieval
KBMill Brick Retrieval Demos
Portable, residual-honest knowledge bricks turned into retrieval evaluation sets for KBMill — the public mill at kbmill.com. Shelf packages live in kbmill-brick-library.
These are not unbounded wiki dumps or synthetic QA. They come from real KBMill manufacturing: bounded packages with muted residual junk filtered where applicable, craft notes, security report in the ZIP, and published, re-runnable cosine retrieval evidence.
Config
Queries… See the full description on the dataset page: https://huggingface.co/datasets/CMiller/kbmill-brick-retrieval.kazqad-retrieval
Dataset Card for KazQAD-Retrieval
Dataset Summary
KazQAD is a Kazakh open-domain Question Answering Dataset
that can be used in both reading comprehension and full ODQA settings, as well as for information retrieval experiments.
This repository contains only the collection and relevance judgments for information retrieval task.
Short answers and data for the reading comprehension task (extractive QA) can be found here.
KazQAD contains just under 6,000 unique questions and… See the full description on the dataset page: https://huggingface.co/datasets/issai/kazqad-retrieval.Unite-Instruct-Retrieval-Train
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
Statistics
Accessing Images and Videos
2025-06-19: We've updated the compressed archives for all image and video files to enable faster extraction.If you've already downloaded the previous files, there's no need to redownload them — the content remains exactly the same. The only difference lies in the compression method, which now allows for quicker… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/Unite-Instruct-Retrieval-Train.ria-news-retrievalretrieval-skquad
Dataset Card for retrieval-skquad
Dataset Summary
STS SK-QuAD Retrieval is a unique dataset designed to evaluate Slovak search performance using metrics like MRR, MAP, and NDCG, derived from the SK-QuAD dataset. It features questions and answers sourced from a search engine before annotation. The annotated data assigns categories to the best answers for each question, enhancing Slovak language search evaluation. This dataset is a significant step forward in the… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/retrieval-skquad.synthetic_retrieval_tasksSynthetic data designed as prompts for generating embeddings training data for retrieval.
The "iteration" column refers to how the data was generated.
Iteration 1: Use the following pool of seed tasks, prompt GPT-3.5-Turbo to generate additional tasks.
RETRIEVAL_EXAMPLES = [
'Provide a scientific claim as query, retrieve documents that help verify or refute the claim.',
'Search for documents that answers a FAQ-style query on children\'s nutrition.',
"Retrieve company's financial reports… See the full description on the dataset page: https://huggingface.co/datasets/andersonbcdefg/synthetic_retrieval_tasks.history_retrieval
Introduction
The HistoryIR dataset was annotated on top of the historical part of BUT-LCC corpus.
We urged annotators to search for historical events (from their own mind, or using our inspirator, more details in the upcoming paper), using the semantic search tool we developed (translation service + English contriever model setup).
Then the annotators annotated top retrieved passages as relevant or irrelevant.
We've done additional filtering step that included manual verification of… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/history_retrieval.synthetic-persian-qa-retrieval
Dataset Summary
Synthetic Persian QA Retrieval (SynPerQARetrieval) is a Persian (Farsi) dataset designed for the Retrieval task, specifically focusing on question answering. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using the GPT-4o-mini Large Language Model. It consists of question-answer pairs derived from the content of various curated Persian websites. The primary task is to retrieve the correct answer… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-qa-retrieval.deka-retrieval
Loading
from datasets import load_dataset
corpus = load_dataset("anonymousresearch123/deka-retrieval", "corpus", split="train")
queries = load_dataset("anonymousresearch123/deka-retrieval", "queries", split="train")
labels = load_dataset("anonymousresearch123/deka-retrieval", "labels", split="train")
``
wikipedia-human-retrieval-ja
Japanese Wikipedia Human Retrieval dataset
This is a Japanese question answereing dataset with retrieval on Wikipedia articles
by trained human workers.
Contributors
Yusuke Oda
defined the dataset specification, data structure, and the scheme of data collection.
Baobab, Inc.
operated data collection, data checking, and formatting.
About the dataset
Each entry represents a single QA session:
given a question sentence, the responsible worker tried to search for… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/wikipedia-human-retrieval-ja.synthetic-persian-chatbot-summary-retrieval
Dataset Summary
Synthetic Persian Chatbot Summary Retrieval (SynPerChatbotSumSRetrieval) is a Persian (Farsi) dataset designed for the novel Summary Retrieval task, introduced as part of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset was synthetically generated using the GPT-4o-mini Large Language Model and is derived from the Synthetic Persian Chatbot Dataset. The core task is to retrieve the correct, pre-generated summary that corresponds to a given user–chatbot… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-summary-retrieval.synthetic-persian-chatbot-rag-summary-retrieval
Dataset Summary
Synthetic Persian Chatbot RAG Summary Retrieval (SynPerChatbotRAGSumSRetrieval) is a Persian (Farsi) dataset for the Summary Retrieval task, specifically built for Retrieval-Augmented Generation (RAG) systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was synthetically generated using GPT-4o-mini and is derived from the Synthetic Persian Chatbot RAG Dataset. It evaluates the ability of models to match conversations—possibly… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-summary-retrieval.synthetic-persian-chatbot-rag-faq-retrieval
Dataset Summary
Synthetic Persian Chatbot RAG FAQ Retrieval (SynPerChatbotRAGFAQRetrieval) is a Persian (Farsi) dataset built for the Retrieval task in Retrieval-Augmented Generation (RAG)-based chatbot systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini. The dataset is designed to evaluate how well models retrieve relevant FAQ entries based on a user's message and prior conversation context.
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-faq-retrieval.passkey-retrievalPasskey retrieval training/evaluation data in Fastchat format. You will have to split into train/evaluation manually.
Articles were drawn from Long C4 in varying lengths
A secret passkey was inserted somewhere in the article, randomly.
The name and type of secret is randomly varied (passphrase, secret key, specific fact, favorite colors, password, etc.) and the passkey itself was randomly generated based on various proper nouns (Faker Library), words/phrases of varying lengths (WonderWords… See the full description on the dataset page: https://huggingface.co/datasets/grimulkan/passkey-retrieval.persian-web-document-retrieval
Dataset Summary
Persian Web Document Retrieval is a Persian (Farsi) dataset designed for the Retrieval task. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset consists of real-world queries collected from the Zarrebin search engine and web documents labeled by humans for relevance. It is curated to evaluate model performance in web search scenarios.
Language(s): Persian (Farsi)
Task(s): Retrieval (Web Search)
Source: Collected from Zarrebin… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/persian-web-document-retrieval.personalized_passkey_retrieval
Dataset Summary
This dataset contains the data for personalized passkey retrieval task in the paper Improving Text Embeddings with Large Language Models.
Data Fields
query: a string feature.
candidates: List of string feature, 100 candidates for each query.
label: a int32 feature, the index of the correct candidate in the candidates list, always 0.
context_length: a int32 feature, the approximate length for the candidate documents.
How to use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/intfloat/personalized_passkey_retrieval.ai-residency-vector-search-retrieval-data
