datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BMGQ-MultiHop-Sample
🧩 BMGQ (Sample Release) – Bottom-up Multi-hop Question Generation Dataset
A Sampled Subset of BMGQ: Complex, Retrieval-Resistant, Multi-hop Reasoning Questions
👥 Authors
Bingsen Qiu, Zijian Liu, Xiao Liu, Bingjie Wang, Feier Zhang, Yixuan Qin, Chunyan Li, Haoshen Yang, Zeren Gao
📘 Dataset Summary
BMGQ is a dataset of complex, hard-to-search, multi-hop reasoning questions automatically generated using our proposed framework:
BMGQ: A Bottom-up… See the full description on the dataset page: https://huggingface.co/datasets/Fayer/BMGQ-MultiHop-Sample.handwritten_multihop_reasoning_data
Dataset used to better understand how to:
Correct Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models
This is a handwritten dataset created to aid in better understanding the multi-hop reasoning capabilities of LLMs.
To learn how the dataset was constructed please check out the project page, paper, and demo linked below.
This is the link to the Project Page.
This repo contains the code that was used to conduct the experiments in this paper.… See the full description on the dataset page: https://huggingface.co/datasets/msakarvadia/handwritten_multihop_reasoning_data.multi_hop-NIPS2026
Dataset Card for Scientific RAG Benchmark (Scenario)
Dataset Details
Dataset Description
This dataset is part of the Scientific RAG Benchmark Collection-NIPS2026.It is designed for evaluating Retrieval-Augmented Generation (RAG) systems and large language models on domain-specific scientific question-answering tasks.
Each scenario contains expert-curated question–answer pairs grounded in peer-reviewed scientific literature, with explicit DOI references to… See the full description on the dataset page: https://huggingface.co/datasets/anonymousauthor2026nips/multi_hop-NIPS2026.MultiHopRAG-syn-data-ctx_len-4096-100MultiHopRAG-syn-data-100multihopfakepedia
