datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2wikimultihopqa_with_q_gpt35
2WikiMultihopQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the 2WikiMultihopQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/2wikimultihopqa_with_q_gpt35.scholarly-metadata-corpus
Scholarly Metadata Corpus
A structured corpus of scholarly metadata collected from arXiv, designed for research on scholarly information retrieval, citation representation, bibliographic metadata, and LLM-based processing of academic literature.
Dataset Structure
The corpus is organized by source:
scholary-metadata-corpus/
├── arxiv/
│ ├── 2501.00001.json
│ ├── 2501.00002.json
│ ├── 2501.00003.json
│ ├── ...
│ └── 2501.xxxxx.json
└── ...
Each JSON file… See the full description on the dataset page: https://huggingface.co/datasets/w3nabil/scholarly-metadata-corpus.hotpotqa_with_qa_gpt35
HotpotQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the HotpotQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is formatted as… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/hotpotqa_with_qa_gpt35.fetch_playwright_with_chunk_huggingface_scholarly_3271_v3rvhazy
