datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2wikimultihopqa_with_q_gpt35
2WikiMultihopQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the 2WikiMultihopQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/2wikimultihopqa_with_q_gpt35.hotpotqa_with_qa_gpt35
HotpotQA Dataset with GPT-3.5 Generated Questions
Overview
This repository hosts an enhanced version of the HotpotQA dataset, where each supporting sentence in the dataset has been supplemented with questions generated using OpenAI's GPT-3.5 turbo API. The aim is to provide a richer context for each entry, potentially benefiting various NLP tasks, such as question answering and context understanding.
Dataset Format
Each entry in the dataset is formatted as… See the full description on the dataset page: https://huggingface.co/datasets/scholarly-shadows-syndicate/hotpotqa_with_qa_gpt35.ontolearner-scholarly_knowledge
Scholarly Knowledge Domain Ontologies
Overview
The scholarly knowledge domain encompasses ontologies that systematically represent the intricate structures, processes, and governance mechanisms inherent in scholarly research, academic publications, and the supporting infrastructure. This domain is pivotal in facilitating the organization, retrieval, and dissemination of academic knowledge, thereby enhancing the efficiency and transparency of scholarly communication.… See the full description on the dataset page: https://huggingface.co/datasets/SciKnowOrg/ontolearner-scholarly_knowledge.scholarly-metadata-corpus
Scholarly Metadata Corpus
A structured corpus of scholarly metadata collected from arXiv, designed for research on scholarly information retrieval, citation representation, bibliographic metadata, and LLM-based processing of academic literature.
Dataset Structure
The corpus is organized by source:
scholary-metadata-corpus/
├── arxiv/
│ ├── 2501.00001.json
│ ├── 2501.00002.json
│ ├── 2501.00003.json
│ ├── ...
│ └── 2501.xxxxx.json
└── ...
Each JSON file… See the full description on the dataset page: https://huggingface.co/datasets/w3nabil/scholarly-metadata-corpus.Scholarly-Epistemic-Engine
Dataset Card for Scholarly-Epistemic-Engine: arXiv cs.AI Corpus and Embeddings
This dataset contains the processed text, metadata, and semantic vector embeddings of approximately 90,000 scholarly articles from the arXiv Computer Science - Artificial Intelligence (cs.AI) category, spanning from 1993 to December 2024. It is designed to support Retrieval-Augmented Generation (RAG) systems and semantic knowledge discovery.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/whyamanbhardwaj/Scholarly-Epistemic-Engine.open-scholarly-document-catalog-10k
Search interactively
·
Product and methodology
·
Full catalog
Rights-Audited Scholarly PDF Catalog: 10K Validated Sample
Evaluation sample / inspect before you buy. Metadata and source links
only. No PDFs or extracted full text are redistributed.
Build scientific-literature search, RAG, discovery, and corpus-acquisition
workflows from scholarly metadata with current PDF-link evidence and per-record
Creative Commons or NASA rights… See the full description on the dataset page: https://huggingface.co/datasets/PlethoraSolutions/open-scholarly-document-catalog-10k.indian_ipo_prospectus_data_with_pageno
Dataset Card for Dataset Name
Dataset Summary
Prospectus text mining is very important for the investor community to identify major risks.
factors and evaluate the use of the amount to be raised during an IPO. For this dataset author
downloaded 100 prospectuses from the Indian Market Regulator website. The dataset contains the URL and OCR text for 100 prospectuses.
Further, the author released a Roberta LM and sentence transformer for usage.
This dataset Contains Page… See the full description on the dataset page: https://huggingface.co/datasets/scholarly360/indian_ipo_prospectus_data_with_pageno.filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_sujqk4
BlueWave Logistics Delivery-Point Registry
This dataset maintains the delivery-point registry for BlueWave Logistics (regional freight & dispatch).
Files
registry.csv — the master delivery-point registry.
review_decisions.csv — the latest Q3 2026 review report (published by the operations analyst).
registry.csv schema
Columns: id,branch,address,city,state,status,review_month
id: delivery-point identifier (e.g. DP-101).
branch: operations branch… See the full description on the dataset page: https://huggingface.co/datasets/zhuq41/filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_sujqk4.scholarly-article-citations-in-wikipediaThis dataset includes a list of citations to scholarly articles from a 2015 version of English Wikipedia.
Citations are in the form of PubMed IDs (pmid) and PubMedCentral IDs (pmcid).
Digital Object Identifiers (doi)
Format
Each row in the dataset represents a citation as a (Wikipedia article, scholarly article) pair. Metadata about when the citation was first added is included.
page_id: The identifier of the Wikipedia article (int), e.g. 1325125
page_title: The title of the… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia-community/scholarly-article-citations-in-wikipedia.contracts-extraction-instruction-llm-experiments
Dataset Card for "contracts-extraction-instruction-llm-experiments"
More Information needed
indian_ipo_prospectus_data
Dataset Card for Dataset Name
Dataset Summary
Prospectus text mining is very important for the investor community to identify major risks.
factors and evaluate the use of the amount to be raised during an IPO. For this dataset author
downloaded 100 prospectuses from the Indian Market Regulator website. The dataset contains the URL and OCR text for 100 prospectuses.
Further, the author released a Roberta LM and sentence transformer for usage.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/scholarly360/indian_ipo_prospectus_data.filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_1zytb4
BlueWave Logistics Delivery-Point Registry
This dataset maintains the delivery-point registry for BlueWave Logistics (regional freight & dispatch).
Files
registry.csv — the master delivery-point registry.
review_decisions.csv — the latest Q3 2026 review report (published by the operations analyst).
registry.csv schema
Columns: id,branch,address,city,state,status,review_month
id: delivery-point identifier (e.g. DP-101).
branch: operations branch… See the full description on the dataset page: https://huggingface.co/datasets/zhuq41/filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_1zytb4.open-scholarly-document-catalog-full
Explore the sample
·
Product and methodology
·
Founding-customer checkout
Open Scholarly Document Catalog: Full Commercial Snapshot
Commercial catalog / dated and reproducible. Metadata and source links
only. No PDFs or extracted full text are redistributed.
This gated package is the complete dated 2026-08-30 metadata snapshot from
Plethora Solutions. It contains 965,792 unique scholarly-document records in
10 Parquet shards with… See the full description on the dataset page: https://huggingface.co/datasets/PlethoraSolutions/open-scholarly-document-catalog-full.playwright_with_chunk_scholarly_huggingface_744_hptc6g-analysis
Northwind Customer Survey Responses
Survey responses collected from Northwind customers for research purposes.
Dataset ID: SRC-EPSILON
Usage: non-commercial use only
filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_ugvtc2
BlueWave Logistics Delivery-Point Registry
This dataset maintains the delivery-point registry for BlueWave Logistics (regional freight & dispatch).
Files
registry.csv — the master delivery-point registry.
review_decisions.csv — the latest Q3 2026 review report (published by the operations analyst).
registry.csv schema
Columns: id,branch,address,city,state,status,review_month
id: delivery-point identifier (e.g. DP-101).
branch: operations branch… See the full description on the dataset page: https://huggingface.co/datasets/zhuq41/filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_ugvtc2.salestech_sales_qualification_framework_bant
Dataset Card for "salestech_sales_qualification_framework_bant"
license: apache-2.0
Dataset Card for Dataset Name
Dataset Summary
The BANT technique is a sales qualifying framework that considers a prospect's budget, internal influence/ability to buy, need for the product, and timeframe for making a purchase when determining whether to pursue a sale.
Because it aids in lead qualification during the discovery call, BANT plays a vital role in the… See the full description on the dataset page: https://huggingface.co/datasets/scholarly360/salestech_sales_qualification_framework_bant.BAGELS_Limitations_data_scholarly_textterrain_generation_from_sketch_for_game_assetsscholarly_huggingface_5019_qi0anr1e_model_eval_normhuggingface_terminal_scholarly_3562_ref_testfetch_playwright_with_chunk_huggingface_scholarly_3271_v3rvhazyscholarly_huggingface_5019_qi0anr1e_model_eval_srccontracts-classification-instruction-llm-experiments
Dataset Card for "contracts-classification-instruction-llm-experiments"
More Information needed
africa-owid-annual-scholarly-publications-on-artificial-intelligence
Annual Scholarly Publications On Artificial Intelligence | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-annual-scholarly-publications-on-artificial-intelligence.asia-owid-annual-scholarly-publications-on-artificial-intelligence
Annual Scholarly Publications On Artificial Intelligence | Asia (Our World in Data)
🌏 416 observations · 48 Asia countries · 2016–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 416 observations of Annual Scholarly Publications On Artificial Intelligence data across 48 Asia countries, spanning 2016–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Annual Scholarly… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-annual-scholarly-publications-on-artificial-intelligence.europe-owid-annual-scholarly-publications-on-artificial-intelligence
Annual Scholarly Publications On Artificial Intelligence | Europe (Our World in Data)
🇪🇺 382 observations · 44 Europe countries · 2016–2024 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 382 observations of Annual Scholarly Publications On Artificial Intelligence data across 44 Europe countries, spanning 2016–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Annual… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-annual-scholarly-publications-on-artificial-intelligence.
