datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2WikiMultihopQA2WikiMultihopQA
2WikiMultihopQA
This repository only repackages the original 2WikiMultihopQA data so that every example follows the field layout used by HotpotQA. The content of the underlying questions, answers and contexts is unaltered.
All intellectual credit for creating 2WikiMultihopQA belongs to the authors of the paper Constructing a Multi‑hop QA Dataset for Comprehensive Evaluation of Reasoning Steps (COLING 2020) and the accompanying code/data in their GitHub project… See the full description on the dataset page: https://huggingface.co/datasets/framolfese/2WikiMultihopQA.2WikiMultihopQAMirror of https://github.com/Alab-NII/2wikimultihop2wikimultihopqa_with_q_gpt35
2WikiMultihopQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the 2WikiMultihopQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/2wikimultihopqa_with_q_gpt35.2WikiMultihopQAUpdated on https://huggingface.co/datasets/voidful/2WikiMultihopQA/blob/main/dev.json with modifications.
2wikimultihopqa\msa-2wikimultihopqa-qa-with-idsmsa-2wikimultihopqa-qa-with-idsmsa-2wikimultihopqa-docs-with-idsRAG-RL-Hotpotqa-with-2wikimsa-2wikimultihopqa-docs-with-idsadaptive_rag_2wikimultihopqa2wikimultihopqa2wikimultihopqa2WikiMultihopQA2wiki_rand1k2wikimultihopqa_mistral2wiki_rlvr_no_prompttvs-2wikimultihopqa
TVS-2WikiMultiHopQA Dataset
This repository contains the tvs-2wikimultihopqa dataset, which is a key component of the "Think, Verbalize, then Speak" (TVS) framework presented in the paper Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible Speech.
Project Page: https://yhytoto12.github.io/TVS-ReVerT
Paper: https://huggingface.co/papers/2509.16028
Code: https://github.com/yhytoto12/TVS-ReVerT
Introduction
The "Think, Verbalize, then Speak"… See the full description on the dataset page: https://huggingface.co/datasets/yhytoto12/tvs-2wikimultihopqa.2wiki_rlvrikea_2wikimultihopqa_easyikea_2wikimultihopqa_hard2wikimultihopqa_searchmsa-2wikimultihopqa-c10000-eval-queriesadaptive_rag_2wikimultihopqaIn this collection you can find 4 datasets with is_supporting=True contexts from the Adaptive RAG collection.
There are picked 4/6 datasets from Adaptive RAG datasets with is_supporting=True contexts.
Not all samples from TriviaQA and SQUAD have is_supporting=True contexts, thats why we do not include them in hf collection.
Script for data transformation from original Adaptive RAG format into our format can be found here:… See the full description on the dataset page: https://huggingface.co/datasets/aboriskin/adaptive_rag_2wikimultihopqa.RAG-RL-2Wiki-OODRAG-RL-2Wiki-Eval-Gold-Only-50k2wiki_questions2wiki_test_fulltrain_2wikimultihopqa
