datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.knowledge-graph-risk-engine-20260828-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260828-dataset.aicoolies-developer-tools-knowledge-graph
aicoolies-developer-tools-knowledge-graph
Public catalog dump from aicoolies.com: tools, comparisons, and scored reviews as JSON.
This dataset is not a coding agent. It does not edit repositories, run tools, or execute code. It is a machine-readable snapshot of the public aicoolies Developer Tools Knowledge Graph so humans and research agents can reuse the catalog without scraping HTML.
Canonical open-data page: https://aicoolies.com/data
Homepage: https://aicoolies.com… See the full description on the dataset page: https://huggingface.co/datasets/rasitakyol/aicoolies-developer-tools-knowledge-graph.knowledge-graph-risk-engine-20260907-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260907-dataset.knowledge-graph-rag-retrieval-artifacts
Knowledge Graph RAG Retrieval Artifacts
Pinned vector-retrieval artifacts for the Knowledge Graph RAG Assistant, a Washington State University capstone project combining knowledge-graph and dense-vector retrieval.
This repository is a project-owned, documented mirror of the two binary artifacts used by the maintained application. The files are byte-identical to the current artifacts originally hosted in miverson9/acme10-he-ragapp-embeddings at revision… See the full description on the dataset page: https://huggingface.co/datasets/ethanvillalovoz/knowledge-graph-rag-retrieval-artifacts.knowledge-graph-risk-engine-20260917-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260917-dataset.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/wuwu616/huatuo_knowledge_graph_qa.NEOMANITAI-Knowledge-Graph
⚠️ LIVING WORK DOCUMENT — DRAFT STATE ⚠️
This dataset is part of the AUGMANITAI Compendium, a living research work document, continuously updated. Each entry is a priority anchor for terminological provenance — not a final reference. Errors, omissions and improvements are expected and explicitly part of the evolving methodology.
LEBENDES ARBEITSDOKUMENT — ENTWURFSSTADIUM. Laufend aktualisiert. Prioritäts-Anker, nicht finale Referenz.
Author: Andreas Ehstand · ORCID: 0009-0006-3773-7796 ·… See the full description on the dataset page: https://huggingface.co/datasets/AndreasEhstand/NEOMANITAI-Knowledge-Graph.knowledge-graph-risk-engine-20260719-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260719-dataset.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/Pete1994/huatuo_knowledge_graph_qa.knowledge-graph-risk-engine-20260818-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260818-dataset.Turkish-Knowledge-Graph-Corpus
Turkish Knowledge Graph Corpus
Triple'lar, Knowledge Graph Production platformunun Triple Benchmark
sayfasında agentic child/parent RAG doğrulamasından "verified" (doğrulanmış)
statüsüyle geçen çıktılardan oluşuyor. Her satır tekil bir (baş, ilişki, uç)
üçlüsü -- aynı üçlü normalize edilmiş haliyle (küçük/büyük harf ve boşluk
farkları yok sayılarak) yalnızca bir kez yer alır.
Alanlar
input_text: Triple'ın çıkarıldığı kaynak metin/cümle (bilinmiyorsa boş).
subject… See the full description on the dataset page: https://huggingface.co/datasets/SalihHub/Turkish-Knowledge-Graph-Corpus.graphiti_knowledge_graphhan-distributed-knowledge-graph-interaction-dataset-v1
Humanoid Distributed Knowledge Graph Interaction Dataset
This dataset captures semantic reasoning
and knowledge graph traversal
performed by humanoid agents
within decentralized knowledge systems.
It models entity linking,
context expansion,
and inference path efficiency.
Objective
To enhance structured reasoning
and distributed knowledge utilization
across humanoid agents.
Data Fields
agent_id
query_id
entity_nodes_traversed
relation_paths
inference_depth… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-distributed-knowledge-graph-interaction-dataset-v1.wikitolica-knowledge-graph
🕸️ Wikitolica Knowledge Graph
Este repositorio gestiona los metadatos y la preservación digital del Knowledge Graph de la Enciclopedia Católica Wikitólica. Para garantizar la frescura de los datos, el dataset principal se sirve de forma dinámica desde nuestra infraestructura oficial.
📊 Acceso a los Datos
Los datos están disponibles en los siguientes puntos de enlace:
Dataset Principal (JSON-LD): https://www.wikitolica.com/knowledge-graph.jsonld
Portal de Datos… See the full description on the dataset page: https://huggingface.co/datasets/CursoCatolico/wikitolica-knowledge-graph.knowledge-graph-risk-engine-20260808-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260808-dataset.FactoryBench-KnowledgeGraphknowledge-graph-risk-engine-20260729-dataset
Knowledge Graph Risk Engine Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Risk teams need relationship-level explanations instead of opaque entity scores.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic:… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/knowledge-graph-risk-engine-20260729-dataset.knowledge-graph-triplets-sharegptknowledgegraph-creative-writing-subsethuatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/chatarchive/huatuo_knowledge_graph_qa.wikipedia_knowledge_graph_en
Dataset Card for Wikipedia Knowledge Graph
The dataset contains 16_958_654 extracted ontologies from a subset of selected wikipedia articles.
Dataset Creation
The dataset was created via LLM processing a subset of the English Wikipedia 20231101.en dataset.
The initial knowledge base dataset was used as a basis to extract the ontologies from.
Pipeline: Wikipedia article → Chunking → Fact extraction (Knowledge base dataset) → Ontology extraction from facts →… See the full description on the dataset page: https://huggingface.co/datasets/Jayesh2160/wikipedia_knowledge_graph_en.FactoryBench-KnowledgeGraphvietnam-durian-knowledge-graph
DurianExpert-VN: The First Structured Knowledge Graph for Vietnam's Durian Industry
1. Project Overview
This dataset fills a critical gap in Agricultural AI. While mainstream models understand general botany, they lack the "local expertise" required for **Durian (Durio zibethinus)**—the highest-value fruit in Southeast Asia.
2. The Challenge: Why it's "Uncharted"
Linguistic Barrier: Farmers use vernacular terms like "mắt cua" (crab eyes) or "mũi giáo" (spear… See the full description on the dataset page: https://huggingface.co/datasets/batinho/vietnam-durian-knowledge-graph.
