datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikidata_triple_enwikidataSR-KI Dataset
The dataset accompanying SR-KI: Scalable and Real-Time Knowledge Integration into LLMs via Supervised Attention (AAAI 2026).
Overview
The SR-KI Dataset provides Chinese question-answering data for training and evaluating the supervised-attention knowledge-integration method introduced in the SR-KI paper. Each example pairs a question and answer with the supporting knowledge and its corresponding material identifier, enabling models to produce… See the full description on the dataset page: https://huggingface.co/datasets/SharkSpicy/wikidata.wikidata-entity-translationsshinto-wikidata-qa
Shinto Wikidata QA
Instruction/QA pairs about the Shinto domain — Shinto shrines, kami (deities, with
genealogy), and key texts (Engishiki, Kojiki, Nihon Shoki) — generated from Wikidata
structured facts.
Built for the Adaption Labs AutoScientist Challenge (All Other Domains track).
Credit: Adaptive Data by Adaption.
Source & license
Source: Wikidata Query Service (https://query.wikidata.org). All statement data is
CC0 / public domain, so this derived dataset is… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/shinto-wikidata-qa.wikidata-parallel-descriptions-en-ja
Wikidata parallel descriptions en-ja
Parallel corpus for machine translation generated from wikidata dump (2024-05-06).
Currently we processed only English/Japanese pair.
The jsonl file is ready-to-train by Hugging Face transformers trainer for translation tasks.
Dataset Details
https://www.wikidata.org/wiki/Wikidata:Database_download
Dataset Creation
As Wikidata description field does not represent exact direct translation, filtering is required for… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/wikidata-parallel-descriptions-en-ja.en_wikidata_5M_entities
en_wikidata_5M_entities
Hugging Face dataset card for a large, English-only Wikidata slice with optional Wikipedia links and Wikimedia Commons image URLs.
One file, five million entities.
Filename: en_wikidata_5M_entities.jsonl.gz
TL;DR
Format: JSON Lines, gzip-compressed (.jsonl.gz)
Rows: 5,000,000 entities (one JSON object per line)
Language: English labels/descriptions
Fields: qid, label, description, enwiki_title, wikipedia_url, images (list of URLs), has_image… See the full description on the dataset page: https://huggingface.co/datasets/Vijaysr4/en_wikidata_5M_entities.WikidataThis dataset accompanies the paper:
When Do LLMs Admit Their Mistakes? Understanding the Role of Model Belief in Retraction
It includes the original Wikidata questions used in our experiments, with train/test split. For a detailed explanation of the dataset construction and usage, please refer to the paper.
Code: https://github.com/ayyyq/llm-retraction
Citation
@misc{yang2025llmsadmitmistakesunderstanding,
title={When Do LLMs Admit Their Mistakes? Understanding the Role of… See the full description on the dataset page: https://huggingface.co/datasets/ayyyq/Wikidata.wikidata_rdf_massive_objects_ENWikidata_Query_Logs_Dataseten_wikidata_5M_entities
en_wikidata_5M_entities
Hugging Face dataset card for a large, English-only Wikidata slice with optional Wikipedia links and Wikimedia Commons image URLs.
One file, five million entities.
Filename: en_wikidata_5M_entities.jsonl.gz
TL;DR
Format: JSON Lines, gzip-compressed (.jsonl.gz)
Rows: 5,000,000 entities (one JSON object per line)
Language: English labels/descriptions
Fields: qid, label, description, enwiki_title, wikipedia_url, images (list of URLs)… See the full description on the dataset page: https://huggingface.co/datasets/dhruv-anand-aintech/en_wikidata_5M_entities.wikidata_triple_jawikidata_oven_subgraphwikiDatasetv3wikidata_b21_datasetДатасет составлен на основе KG WikiData. Описание всех файлом можно почитать на официальном сайте.
Там же можно скачать все файлы
Для начала были найдены тройки связанных вершин, после сгенерированы вопросы для них.
После для связанных троек также были сгенерированы усложненные вопросы вида bridge_2_1 (как описано в статье).
wikidata
