scholarly
2wikimultihopqa_with_q_gpt35
2WikiMultihopQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the 2WikiMultihopQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/2wikimultihopqa_with_q_gpt35.hotpotqa_with_qa_gpt35
HotpotQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the HotpotQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is formatted as… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/hotpotqa_with_qa_gpt35.ontolearner-scholarly_knowledge
Scholarly Knowledge Domain Ontologies
Overview
The scholarly knowledge domain encompasses ontologies that systematically represent the intricate structures, processes, and governance mechanisms inherent in scholarly research, academic publications, and the supporting infrastructure. This domain is pivotal in facilitating the organization, retrieval, and dissemination of academic knowledge, thereby enhancing the efficiency and transparency of scholarly communication.… See the full description on the dataset page: https://huggingface.co/datasets/SciKnowOrg/ontolearner-scholarly_knowledge.scholarly-metadata-corpus
Scholarly Metadata Corpus
A structured corpus of scholarly metadata collected from arXiv, designed for research on scholarly information retrieval, citation representation, bibliographic metadata, and LLM-based processing of academic literature.
Dataset Structure
The corpus is organized by source:
scholary-metadata-corpus/
├── arxiv/
│ ├── 2501.00001.json
│ ├── 2501.00002.json
│ ├── 2501.00003.json
│ ├── ...
│ └── 2501.xxxxx.json
└── ...
Each JSON file… See the full description on the dataset page: https://huggingface.co/datasets/w3nabil/scholarly-metadata-corpus.Scholarly-Epistemic-Engine
Dataset Card for Scholarly-Epistemic-Engine: arXiv cs.AI Corpus and Embeddings
This dataset contains the processed text, metadata, and semantic vector embeddings of approximately 90,000 scholarly articles from the arXiv Computer Science - Artificial Intelligence (cs.AI) category, spanning from 1993 to December 2024. It is designed to support Retrieval-Augmented Generation (RAG) systems and semantic knowledge discovery.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/whyamanbhardwaj/Scholarly-Epistemic-Engine.open-scholarly-document-catalog-10k
Search interactively
·
Product and methodology
·
Full catalog
Rights-Audited Scholarly PDF Catalog: 10K Validated Sample
Evaluation sample / inspect before you buy. Metadata and source links
only. No PDFs or extracted full text are redistributed.
Build scientific-literature search, RAG, discovery, and corpus-acquisition
workflows from scholarly metadata with current PDF-link evidence and per-record
Creative Commons or NASA rights… See the full description on the dataset page: https://huggingface.co/datasets/PlethoraSolutions/open-scholarly-document-catalog-10k.
