CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RJZ /wikidata_triple_entext10M<n<100M1 likes1.7k downloads2y agoHugging Face02SharkSpicy /wikidataSR-KI Dataset The dataset accompanying SR-KI: Scalable and Real-Time Knowledge Integration into LLMs via Supervised Attention (AAAI 2026). Overview The SR-KI Dataset provides Chinese question-answering data for training and evaluating the supervised-attention knowledge-integration method introduced in the SR-KI paper. Each example pairs a question and answer with the supporting knowledge and its corresponding material identifier, enabling models to produce… See the full description on the dataset page: https://huggingface.co/datasets/SharkSpicy/wikidata.textquestion-answering100K<n<1M1 likes77 downloads19d agoHugging Face03agentlans /wikidata-entity-translationstexttranslation10M<n<100M1 likes76 downloads3mo agoHugging Face04EmmaLeonhart /shinto-wikidata-qa Shinto Wikidata QA Instruction/QA pairs about the Shinto domain — Shinto shrines, kami (deities, with genealogy), and key texts (Engishiki, Kojiki, Nihon Shoki) — generated from Wikidata structured facts. Built for the Adaption Labs AutoScientist Challenge (All Other Domains track). Credit: Adaptive Data by Adaption. Source & license Source: Wikidata Query Service (https://query.wikidata.org). All statement data is CC0 / public domain, so this derived dataset is… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/shinto-wikidata-qa.textquestion-answering100K<n<1M0 likes54 downloads3mo agoHugging Face05Mitsua /wikidata-parallel-descriptions-en-ja Wikidata parallel descriptions en-ja Parallel corpus for machine translation generated from wikidata dump (2024-05-06). Currently we processed only English/Japanese pair. The jsonl file is ready-to-train by Hugging Face transformers trainer for translation tasks. Dataset Details https://www.wikidata.org/wiki/Wikidata:Database_download Dataset Creation As Wikidata description field does not represent exact direct translation, filtering is required for… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/wikidata-parallel-descriptions-en-ja.texttranslation1M<n<10M9 likes53 downloads2y agoHugging Face06Vijaysr4 /en_wikidata_5M_entities en_wikidata_5M_entities Hugging Face dataset card for a large, English-only Wikidata slice with optional Wikipedia links and Wikimedia Commons image URLs. One file, five million entities. Filename: en_wikidata_5M_entities.jsonl.gz TL;DR Format: JSON Lines, gzip-compressed (.jsonl.gz) Rows: 5,000,000 entities (one JSON object per line) Language: English labels/descriptions Fields: qid, label, description, enwiki_title, wikipedia_url, images (list of URLs), has_image… See the full description on the dataset page: https://huggingface.co/datasets/Vijaysr4/en_wikidata_5M_entities.textother1M<n<10M2 likes52 downloads1y agoHugging Face07ayyyq /WikidataThis dataset accompanies the paper: When Do LLMs Admit Their Mistakes? Understanding the Role of Model Belief in Retraction It includes the original Wikidata questions used in our experiments, with train/test split. For a detailed explanation of the dataset construction and usage, please refer to the paper. Code: https://github.com/ayyyq/llm-retraction Citation @misc{yang2025llmsadmitmistakesunderstanding, title={When Do LLMs Admit Their Mistakes? Understanding the Role of… See the full description on the dataset page: https://huggingface.co/datasets/ayyyq/Wikidata.textquestion-answering1K<n<10K0 likes39 downloads1y agoHugging Face08Yomm1927 /wikidata_rdf_massive_objects_ENtext1M<n<10M0 likes34 downloads2y agoHugging Face09momo4382 /Wikidata_Query_Logs_Datasettext100K<n<1M0 likes26 downloads6mo agoHugging Face10dhruv-anand-aintech /en_wikidata_5M_entities en_wikidata_5M_entities Hugging Face dataset card for a large, English-only Wikidata slice with optional Wikipedia links and Wikimedia Commons image URLs. One file, five million entities. Filename: en_wikidata_5M_entities.jsonl.gz TL;DR Format: JSON Lines, gzip-compressed (.jsonl.gz) Rows: 5,000,000 entities (one JSON object per line) Language: English labels/descriptions Fields: qid, label, description, enwiki_title, wikipedia_url, images (list of URLs)… See the full description on the dataset page: https://huggingface.co/datasets/dhruv-anand-aintech/en_wikidata_5M_entities.textother1M<n<10M0 likes24 downloads4mo agoHugging Face11RJZ /wikidata_triple_jatext10M<n<100M1 likes21 downloads2y agoHugging Face12zhKingg /wikidata_oven_subgraphtext10K<n<100K0 likes13 downloads10mo agoHugging Face13bysdynamo /wikiDatasetv3textn<1K0 likes10 downloads2y agoHugging Face14VavGreg /wikidata_b21_datasetgatedДатасет составлен на основе KG WikiData. Описание всех файлом можно почитать на официальном сайте. Там же можно скачать все файлы Для начала были найдены тройки связанных вершин, после сгенерированы вопросы для них. После для связанных троек также были сгенерированы усложненные вопросы вида bridge_2_1 (как описано в статье). text100K<n<1M1 likes5 downloads5mo agoHugging Face15sneka2001 /wikidatatext10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.