datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikidata-truthy
Wikidata Truthy
Dataset Description
Core facts from Wikidata (preferred statements only)
Original Source: https://dumps.wikimedia.org/wikidatawiki/entities/latest-truthy.nt.bz2
Dataset Summary
This dataset contains RDF triples from Wikidata Truthy converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally ntriples, converted to HuggingFace Dataset
Size: 100.0 GB (extracted)
Entities: ~100M
Triples: ~2B
Original… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/wikidata-truthy.BRINK-Wikidata5m
BRINK-Wikidata5m
BRINK (Benchmark for Reasoning under Incomplete Knowledge) is a benchmark for evaluating Knowledge Graph–based Retrieval-Augmented Generation (KG-RAG) under incomplete knowledge. Unlike standard KGQA benchmarks, BRINK is designed so that each question cannot be answered by directly retrieving a single explicit supporting triple. Instead, the answer must be inferred from alternative reasoning paths that remain in the graph after the directly supporting fact is… See the full description on the dataset page: https://huggingface.co/datasets/ZDZR/BRINK-Wikidata5m.shinto-wikidata-qa
Shinto Wikidata QA
Instruction/QA pairs about the Shinto domain — Shinto shrines, kami (deities, with
genealogy), and key texts (Engishiki, Kojiki, Nihon Shoki) — generated from Wikidata
structured facts.
Built for the Adaption Labs AutoScientist Challenge (All Other Domains track).
Credit: Adaptive Data by Adaption.
Source & license
Source: Wikidata Query Service (https://query.wikidata.org). All statement data is
CC0 / public domain, so this derived dataset is… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/shinto-wikidata-qa.impuls-wikidata-kb
IMPULS R&D Knowledge Base
A multilingual knowledge base of 4,265 R&D concepts derived from Wikidata, designed for query expansion in scientific and research project search systems.
Dataset Description
This knowledge base was created as part of the IMPULS project (AINA Challenge 2024), a collaboration between SIRIS Academic and Generalitat de Catalunya to build a multilingual semantic search system for R&D ecosystems.
The KB contains scientific and technological concepts… See the full description on the dataset page: https://huggingface.co/datasets/SIRIS-Lab/impuls-wikidata-kb.
