datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikidata1
Wikipedia N-Link Basins
A novel graph-theoretic analysis of Wikipedia's internal link structure, revealing deterministic "basins of attraction" under N-link traversal rules.
Dataset Description
This dataset demonstrates that Wikipedia's 17.9 million pages partition into coherent regions when following a simple rule: from any page, always follow the Nth link. Every page eventually reaches a cycle, and pages sharing the same terminal cycle form a basin of attraction.… See the full description on the dataset page: https://huggingface.co/datasets/mgmacleod/wikidata1.ko_wikidata_QA
업데이트 로그
2023-11-03 : MarkrAI의 Dedup 적용.
한국어 위키 데이터 QA셋
본 데이터는 Synatra-7B-Instruct 모델과 ChatGPT를 사용하여, 제작된 QA셋입니다.
해당 데이터를 직접적으로 상업적으로 사용하는 것은 허용되지 않으며, 데이터를 이용하여 훈련된 모델에 대한 상업적 사용은 허용됩니다.
아직 완벽히 정제되지는 않았으며, 오류나 수정사항에 대해서는 PR 부탁드립니다.
BRINK-Wikidata5m
BRINK-Wikidata5m
BRINK (Benchmark for Reasoning under Incomplete Knowledge) is a benchmark for evaluating Knowledge Graph–based Retrieval-Augmented Generation (KG-RAG) under incomplete knowledge. Unlike standard KGQA benchmarks, BRINK is designed so that each question cannot be answered by directly retrieving a single explicit supporting triple. Instead, the answer must be inferred from alternative reasoning paths that remain in the graph after the directly supporting fact is… See the full description on the dataset page: https://huggingface.co/datasets/ZDZR/BRINK-Wikidata5m.freebase-wikidata-mapping
mapping between freebase and wikidata entities
This dataset maps freebase ids to wikidata ids and labels. It is useful for visualising and better understanding when working with datasets like fb15k-237
How it was created:
Download freebase-wikidata mapping from here. [compressed size: 21.2 MB]
Download wikidata entities data from here. [compressed size: 81GB]
Align labels with the freebase,wikidata id
Wikidata5mWikidata-celebrity-parentWikidataMediaEntitiescollection of media related keywords from Wikidata collected via SPARQL queries
wikiDatawikidatawikidata-geoloc-properties-20220907wikidata_reference
Dataset Card for Triple-to-Text Alignment Dataset
Dataset Summary
The Triple-to-Text Alignment dataset aligns Knowledge Graph (KG) triples from Wikidata with diverse, real-world textual sources extracted from the web. Unlike previous datasets that rely primarily on Wikipedia text, this dataset provides a broader range of writing styles, tones, and structures by leveraging Wikidata references from various sources such as news articles, government reports, and scientific… See the full description on the dataset page: https://huggingface.co/datasets/sven-h/wikidata_reference.wiki-datawikidata
