datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Intiial-Knowledge-And-Detailed-Assessment-JSON-Format-Datapeka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
Preprint: https://arxiv.org/pdf/2411.04920
Web interface for browsing GPTKB: https://gptkb.org
Current_Trivia_Knowledge-benchmark
Current Trivia Knowledge RAG Benchmark
Short Summary:
A 140-QA pair (70 train, 70 test) dataset for real-world RAG evaluation. It features current knowledge questions unavailable to LLMs trained before 2024 (e.g., GPT-4o) across diverse domains, and includes human feedback for the training set, enabling robust assessment of contextual information's critical impact on LLM accuracy.
Introduction & Motivation:
This dataset addresses the critical need for a dynamic… See the full description on the dataset page: https://huggingface.co/datasets/TPelc/Current_Trivia_Knowledge-benchmark.cooking-knowledge-basics
Comprehensive Cooking Knowledge Q&A Dataset
This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.qulture-general-knowledge-dataset
Qulture General Knowledge Question Dataset
An open dataset containing 16 families of general knowledge questions, for a total of 64 records.
GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information.
Papers:
GPTKB methodology: https://arxiv.org/pdf/2411.04920
GPTKB v1.5: https://arxiv.org/pdf/2507.05740
Citations:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
@article{GPTKB15,
title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.NLP-KnowledgeGraph
Dataset Card for Dataset Name
Dataset Summary
KG dataset created by using spaCy PoS and Dependency parser.
Supported Tasks and Leaderboards
Can be leveraged for token classification for detection of knowledge graph entities and relations.
Languages
English
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
Important fields for the token classification task are
tokens - tokenized text
tags - Tags… See the full description on the dataset page: https://huggingface.co/datasets/vishnun/NLP-KnowledgeGraph.KnowledgeBerg
KnowledgeBerg
KnowledgeBerg is a multilingual benchmark for knowledge-grounded reasoning over enumerated answer sets.
The default configuration is English. Use the dataset viewer's configuration selector to open each language separately.
Data
Each language is stored as one CSV file under data/<Language>/.
Columns:
id
original_id
domain
original_question
original_answer
question_type
implicit_question
options
correct_option
reasoning_depth
knowledge_width
validation_cot
islamic-knowledge-ai-survey
Advances in AI Systems on Islamic Knowledge Capabilities: A Critical Survey
A comprehensive systematic survey of 160+ papers (2016–2026) examining how AI systems operationalize Islamic knowledge, spanning NLP, information retrieval, speech processing, multimodal learning, educational technology, and LLM alignment.
Abstract
AI systems are increasingly mediating how Islamic communities access, study, and apply Islamic sources; still, research on Islamic-knowledge… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/islamic-knowledge-ai-survey.medical_knowledge_from_extractsThis dataset is used to train LLMs for medical knowledge extraction tasks
healthcare-disease-knowledge
Disease Symptoms & Treatment Dataset
This dataset contains structured information about 687 diseases and their associated details.It is intended for research, educational, and prototyping purposes in healthcare-related ML/NLP tasks.
Contents
Each row corresponds to one disease, with 17 columns:
disease — Name of the disease
main_link — Reference link
Diagnosis_treatment_link — Link to diagnosis/treatment page
Doctors_departments_link — Relevant medical… See the full description on the dataset page: https://huggingface.co/datasets/harmesh95/healthcare-disease-knowledge.rivers-knowledge-base
Rivers Knowledge Base Dataset
Full structured knowledge base combining DBpedia extractions with LLM-augmented data for U.S. rivers. This repository contains raw DBpedia SPARQL query results, augmented hydrological measurements, alternative river names, and geographic and administrative metadata. The knowledge base serves as the source data for knowledge graph construction used in the Licensing Oracle experiments.
Citation
@article{ackermann2025stemming,
title={Stemming… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/rivers-knowledge-base.NeuroExplain-Neurological-Knowledge
NeuroExplain: Neurological Knowledge Dataset
About
This dataset is created for the NeuroExplain: LLM-Based Neurological Report Interpretation and Patient Education System project.
The purpose of this dataset is to provide structured neurological information that can be used to develop an AI system capable of explaining neurological conditions in simple and understandable language.
Dataset Contents
The dataset currently contains information about… See the full description on the dataset page: https://huggingface.co/datasets/aashrithareddyy/NeuroExplain-Neurological-Knowledge.Diverse-Knowledge
Everything Data
This data is synthetically generated by a ton of open and closed source models. This is basically a parsed version of yearly log form a small dialouge based testing to anylyze model's response on it then perform human evals on it.
The data contains information about everything from every domain, most of the pairs included in this data are preferred by humans as the model's response.
It can be used for topic modeling, or human preference evals etc.
Rest anyone can do… See the full description on the dataset page: https://huggingface.co/datasets/kunu5402/Diverse-Knowledge.keen_estimating_knowledge_in_llmsrailway-knowledge-QAknowledge-graph
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset can be used for training the SVO
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/Aditya210/knowledge-graph.legal-file-handover-knowledge-action-deadline-coherence-risk-v0.1What this dataset does
You receive
handover summary
key facts
strategy
tasks
deadlines
new actions
You decide
coherent
or
incoherent
Daily use
transfer QC
missed deadline prevention
strategy drift detection
Knowledge-Driven-Dialoguesfire-knowledge-embedding-wikipedia
Fire Domain Knowledge Safety Wikipedia Embeddings
Pre-computed vector embeddings from Wikipedia articles covering fire hazard safety, environmental effects, and related topics — ready to drop into your RAG pipeline without any embedding overhead.
Dataset Details
Property
Details
Embedding Model
nomic-ai/nomic-embed-text-v1.5 (135M)
Vector Dimensions
768
Source
Wikipedia
Topics
Fire safety, hazard prevention, environmental impact, and more
Format
CSV… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/fire-knowledge-embedding-wikipedia.herbal-knowledge-embedding-wikipedia
Herbs Domain Knowledge Safety Wikipedia Embeddings
Pre-computed vector embeddings from Wikipedia articles covering Assorted Herbs, spices, and other botanical items used for alternative medicine — ready to drop into your RAG pipeline without any embedding overhead.
Dataset Details
Property
Details
Embedding Model
nomic-ai/nomic-embed-text-v1.5 (135M)
Vector Dimensions
768
Source
Wikipedia
Topics
herbs, plants, spices, and other botanical items for… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/herbal-knowledge-embedding-wikipedia.askhistorians-knowledge-filling
Knowledge Filling Dataset
The dataset of our paper Knowledge Acquisition through Continued Pretraining is Difficult: A Case Study on r/AskHistorians from the "ACL 2024 Workshop Towards Knowledgeable Language Models".
Universal-Knowledge-GPTsr_spill_knowledgeBeginner-Golf-Knowledgeknowledge_transfer_1500_base
Dataset Summary
knowledge_transfer_1500_base is a dataset consisting of 1500 prompts distributed across 10 different topics. The prompts are designed to facilitate sentence completion by models, with half of the sentences containing fewer than 5 words and the other half containing between 5 and 10 words. The dataset is intended to serve as a foundation for knowledge transfer tasks, specifically for creating response datasets that can be useful for fine-tuning models to mimic the… See the full description on the dataset page: https://huggingface.co/datasets/oopere/knowledge_transfer_1500_base.durecdial-knowledgedropthe-knowledge-graph
DropThe Entity Relationship Graph
A large-scale entity relationship dataset containing 2.9 million typed, directional connections between 1.8 million entities spanning entertainment, media, finance, and technology. Extracted from the DropThe knowledge graph.
Dataset Description
While most open knowledge graphs focus on encyclopedic facts (Wikidata) or narrow domains (MovieLens for ratings), this dataset captures operational relationships -- the connections that actually… See the full description on the dataset page: https://huggingface.co/datasets/DropTheHQ/dropthe-knowledge-graph.Knowledge_blanks_finetuning
