datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KGQAGen-10k
KGQAGen-10k
KGQAGen-10k is a high-quality example dataset generated using our KGQAGen(github.com/liangliang6v6/KGQAGen) framework for multi-hop Knowledge Graph Question Answering (KGQA). It showcases how large-scale, verifiable QA benchmarks can be automatically constructed from Wikidata using a combination of subgraph expansion, SPARQL validation, and LLM-guided generation.
This 10k release serves as a representative sample demonstrating the scalability, reasoning depth, and… See the full description on the dataset page: https://huggingface.co/datasets/lianglz/KGQAGen-10k.KGQA_model_selection_dataKGQA-data-triplesKGQA_prompt_contextKGQA_promptsKORA-Benchmark
KORA Benchmark
Resources for reproducing KORA: Adaptive Multi-Agent Orchestrated Retrieval over Knowledge Graphs — including the BioCQ benchmark dataset, entity resolution indexes, and the combined biomedical knowledge graph.
Repository Contents
Path
Description
benchmark/
BioCQ question splits (train / val / test / full)
indexes/scispacy*/
Pre-built SciSpaCy entity resolution indexes (~1 GB)
indexes/ark_bm25/
Pre-built ARK BM25 retrieval indexes… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-kgqa/KORA-Benchmark.KGQA_triples
