datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BRINK-Wikidata5m
BRINK-Wikidata5m
BRINK (Benchmark for Reasoning under Incomplete Knowledge) is a benchmark for evaluating Knowledge Graph–based Retrieval-Augmented Generation (KG-RAG) under incomplete knowledge. Unlike standard KGQA benchmarks, BRINK is designed so that each question cannot be answered by directly retrieving a single explicit supporting triple. Instead, the answer must be inferred from alternative reasoning paths that remain in the graph after the directly supporting fact is… See the full description on the dataset page: https://huggingface.co/datasets/ZDZR/BRINK-Wikidata5m.wikidataSR-KI Dataset
The dataset accompanying SR-KI: Scalable and Real-Time Knowledge Integration into LLMs via Supervised Attention (AAAI 2026).
Overview
The SR-KI Dataset provides Chinese question-answering data for training and evaluating the supervised-attention knowledge-integration method introduced in the SR-KI paper. Each example pairs a question and answer with the supporting knowledge and its corresponding material identifier, enabling models to produce… See the full description on the dataset page: https://huggingface.co/datasets/SharkSpicy/wikidata.shinto-wikidata-qa
Shinto Wikidata QA
Instruction/QA pairs about the Shinto domain — Shinto shrines, kami (deities, with
genealogy), and key texts (Engishiki, Kojiki, Nihon Shoki) — generated from Wikidata
structured facts.
Built for the Adaption Labs AutoScientist Challenge (All Other Domains track).
Credit: Adaptive Data by Adaption.
Source & license
Source: Wikidata Query Service (https://query.wikidata.org). All statement data is
CC0 / public domain, so this derived dataset is… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/shinto-wikidata-qa.WikidataThis dataset accompanies the paper:
When Do LLMs Admit Their Mistakes? Understanding the Role of Model Belief in Retraction
It includes the original Wikidata questions used in our experiments, with train/test split. For a detailed explanation of the dataset construction and usage, please refer to the paper.
Code: https://github.com/ayyyq/llm-retraction
Citation
@misc{yang2025llmsadmitmistakesunderstanding,
title={When Do LLMs Admit Their Mistakes? Understanding the Role of… See the full description on the dataset page: https://huggingface.co/datasets/ayyyq/Wikidata.
