datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiHopRAG
Dataset Card for Dataset Name
A Dataset for Evaluating Retrieval-Augmented Generation Across Documents
Dataset Description
MultiHop-RAG: a QA dataset to evaluate retrieval and reasoning across documents with metadata in the RAG pipelines. It contains 2556 queries, with evidence for each query distributed across 2 to 4 documents. The queries also involve document metadata, reflecting complex scenarios commonly found in real-world RAG applications.
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/yixuantt/MultiHopRAG.spoken-multihop-rag
Spoken Multi-hop QA: ASR Transcripts Across Four English Accents
ASR transcriptions of 3,000 multi-hop QA questions, each spoken in four
English accents and transcribed with Whisper-large-v3. Released as the
data companion to Better Retrieval, Worse Robustness: How Multi-hop RAG
Amplifies Upstream ASR Errors
(EMNLP 2026, Main Conference).
The dataset exists to make one thing cheap to study: what happens to a
retrieval pipeline when its query arrives through ASR rather than as… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/spoken-multihop-rag.multi-hop-qa-function-calling-format-V1.0This dataset is converted from khaimaitien/qa-expert-multi-hop-qa-V1.0 to OpenAI function calling format.
Each data point is a list of messages with role=user, assistant or function:
message that role=user, content is the question
message that role=assistant, content is not None, function_call is None: --> assistant responds with text only
message that role=assistant and function_call is not None --> assistant asks to execute a function call
function_call is of the form: {"name": "retrieve"… See the full description on the dataset page: https://huggingface.co/datasets/khaimaitien/multi-hop-qa-function-calling-format-V1.0.multihop-question-decompositionOneGen-TrainDataset-MultiHopQAMultiHopRAG
Dataset Card for Dataset Name
A Dataset for Evaluating Retrieval-Augmented Generation Across Documents
Dataset Description
MultiHop-RAG: a QA dataset to evaluate retrieval and reasoning across documents with metadata in the RAG pipelines. It contains 2556 queries, with evidence for each query distributed across 2 to 4 documents. The queries also involve document metadata, reflecting complex scenarios commonly found in real-world RAG applications.… See the full description on the dataset page: https://huggingface.co/datasets/FQAJ/MultiHopRAG.multi-hop-agentic-websearchMultihopQAwriting-in-the-margins-multihopragmultihopqacotagentsim-atc-multihop
AgentSim Agent-Trace Corpus — Multi-hop
A multi-hop sibling of the AgentSim Agent-Trace Corpus
(agentsim-atc)
with an evolved schema designed for student model distillation.
1 490 accepted SFT trajectories plus 2 980 step-level DPO preference
pairs, generated over 5 multi-hop QA datasets through a 7-action agentic
executor with an Always-Search Policy filter.
This corpus accompanies a follow-up technical report to "AgentSim: A
Platform for Verifiable Agent-Trace Simulation"… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/agentsim-atc-multihop.multi-hop-websearch-tool-callingmultihop-circuitUrdu_MultiHop_QAmultihop_rag_agentmultihoprag-indexmultihopQAhardCoTcomplex_multihop_clinical_reasoning_v5synth-multihopmultihopcot100MultiHopRAG_ja
