datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enwikivoyage-retrieval-202605
English Wikivoyage Retrieval 2026-05
I like travel datasets because they are about real places, real constraints,
and the small practical questions people ask before they go somewhere. I am
sharing this English Wikivoyage retrieval corpus in that spirit: as honest
work from a researcher-builder who wants to explore the world, make the
pipeline inspectable, and let other people reuse or challenge the choices.
This is not presented as a finished travel product or a private… See the full description on the dataset page: https://huggingface.co/datasets/vincentss/enwikivoyage-retrieval-202605.kbmill-brick-retrieval
KBMill Brick Retrieval Demos
Portable, residual-honest knowledge bricks turned into retrieval evaluation sets for KBMill — the public mill at kbmill.com. Shelf packages live in kbmill-brick-library.
These are not unbounded wiki dumps or synthetic QA. They come from real KBMill manufacturing: bounded packages with muted residual junk filtered where applicable, craft notes, security report in the ZIP, and published, re-runnable cosine retrieval evidence.
Config
Queries… See the full description on the dataset page: https://huggingface.co/datasets/CMiller/kbmill-brick-retrieval.financial-retrieval
Overview
This dataset contains normalized, long-form records used to benchmark multiple chatbots on financial retrieval QA.
Each row represents a single (ticker, chatbot) pair answering one question.
These records are derived from a verification step that extracts structured fields from each chatbot’s answer.
Columns (normalized dataset)
Column
Type
Description
ticker
string
Company identifier used for the question (e.g., AAPL, 7203:JP).
question
string… See the full description on the dataset page: https://huggingface.co/datasets/daloopa/financial-retrieval.medquad-retrieval-pretriage
MedQuAD Retrieval Pre-Triage Dataset
Dataset Description
This repository contains a processed, retrieval-oriented derivative of the MedQuAD medical question-answering dataset.
It was prepared for contextual medical information retrieval in SortMed, an academic medical pre-triage assistant.
The corpus is not used to train the SortMed triage classifiers. It is used by a separate semantic retrieval component that identifies medically related question-answer entries… See the full description on the dataset page: https://huggingface.co/datasets/cristian-untaru/medquad-retrieval-pretriage.retrieval_grounding
Dino Data Retrieval Grounding Preview
What This Dataset Is
This dataset is a focused retrieval-grounding preview built from four Dino Data capability slices:
search trigger detection
grounded search integration
history search trigger
history search integration
The goal is to train or inspect assistant behavior around two connected problems:
deciding when retrieval or history lookup is needed
generating answers that stay grounded to supplied evidence or prior thread… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/retrieval_grounding.nanochat-depo-retrieval-copy1-20260715
Nanochat Depo retrieval v1
Each latent 16-node graph yields eight independent, token-aligned, depth-one
query documents. This arm exposes 1 nested edge(s) per
document. Only the answer is supervised in every document; the terminal token
is supervised only for query ordinal 7. This source is separate from and does
not alter Depo-L0 v1.
MoA_Long_Retrievalreal-math-corpus-questions-with-retrievals
Real Math Corpus - Statement Dependencies and Questions
Dataset Description
This dataset contains a comprehensive collection of mathematical statements and questions extracted from the Real Math Dataset with 207 mathematical papers. The dataset is split into two parts:
Corpus: Statement dependencies and proof dependencies with complete metadata and global ID mapping
Questions: Main statements from papers treated as questions, with dependency mappings to the corpus… See the full description on the dataset page: https://huggingface.co/datasets/AK123321/real-math-corpus-questions-with-retrievals.real-math-corpus-questions-with-cross-paper-retrievals
Real Math Corpus - Statement Dependencies and Questions
Dataset Description
This dataset contains a comprehensive collection of mathematical statements and questions extracted from the Real Math Dataset with 207 mathematical papers. The dataset is split into two parts:
Corpus: Statement dependencies and proof dependencies with complete metadata and global ID mapping
Questions: Main statements from papers treated as questions, with enhanced dependency mappings to the… See the full description on the dataset page: https://huggingface.co/datasets/AK123321/real-math-corpus-questions-with-cross-paper-retrievals.nanochat-depo-retrieval-width4-20260715
Nanochat Depo retrieval v1
Each latent 16-node graph yields eight independent, token-aligned, depth-one
query documents. This arm exposes 4 nested edge(s) per
document. Only the answer is supervised in every document; the terminal token
is supervised only for query ordinal 7. This source is separate from and does
not alter Depo-L0 v1.
