datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikidata-truthy
Wikidata Truthy
Dataset Description
Core facts from Wikidata (preferred statements only)
Original Source: https://dumps.wikimedia.org/wikidatawiki/entities/latest-truthy.nt.bz2
Dataset Summary
This dataset contains RDF triples from Wikidata Truthy converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally ntriples, converted to HuggingFace Dataset
Size: 100.0 GB (extracted)
Entities: ~100M
Triples: ~2B
Original… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/wikidata-truthy.BRINK-Wikidata5m
BRINK-Wikidata5m
BRINK (Benchmark for Reasoning under Incomplete Knowledge) is a benchmark for evaluating Knowledge Graph–based Retrieval-Augmented Generation (KG-RAG) under incomplete knowledge. Unlike standard KGQA benchmarks, BRINK is designed so that each question cannot be answered by directly retrieving a single explicit supporting triple. Instead, the answer must be inferred from alternative reasoning paths that remain in the graph after the directly supporting fact is… See the full description on the dataset page: https://huggingface.co/datasets/ZDZR/BRINK-Wikidata5m.shinto-wikidata-qa
Shinto Wikidata QA
Instruction/QA pairs about the Shinto domain — Shinto shrines, kami (deities, with
genealogy), and key texts (Engishiki, Kojiki, Nihon Shoki) — generated from Wikidata
structured facts.
Built for the Adaption Labs AutoScientist Challenge (All Other Domains track).
Credit: Adaptive Data by Adaption.
Source & license
Source: Wikidata Query Service (https://query.wikidata.org). All statement data is
CC0 / public domain, so this derived dataset is… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/shinto-wikidata-qa.normalized-wikidata
Normalized Wikidata
A preprocessed text-form view of Wikidata, optimised for training language
models or knowledge-graph world models. The goal is a corpus where the
semantic content of Wikidata triples comes through cleanly, with the
catalog-and-identifier clutter that dominates raw Wikidata by volume stripped
out.
License inherits from Wikidata: CC-BY-SA 4.0.
This dataset is the input to a corresponding series of Loka world-model
checkpoints at EmmaLeonhart/loka.
Each snapshot… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/normalized-wikidata.impuls-wikidata-kb
IMPULS R&D Knowledge Base
A multilingual knowledge base of 4,265 R&D concepts derived from Wikidata, designed for query expansion in scientific and research project search systems.
Dataset Description
This knowledge base was created as part of the IMPULS project (AINA Challenge 2024), a collaboration between SIRIS Academic and Generalitat de Catalunya to build a multilingual semantic search system for R&D ecosystems.
The KB contains scientific and technological concepts… See the full description on the dataset page: https://huggingface.co/datasets/SIRIS-Lab/impuls-wikidata-kb.
