datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GraphRAG-Bench
GraphRAG-Bench : A Comprehensive Benchmark for Evaluating Graph Retrieval-Augmented Generation Models
🎉News •
📖About •
🏆Leaderboards •
🧩Task Examples
🔧Getting Started •
📬Contact •
📝Citation
This repository is for the GraphRAG-Bench project, a comprehensive benchmark for evaluating Graph Retrieval-Augmented Generation models.
🎉 News
[2025-05-25] We release GraphRAG-Bench, the benchmark for evaluating GraphRAG… See the full description on the dataset page: https://huggingface.co/datasets/GraphRAG-Bench/GraphRAG-Bench.cvefixes-security-ir-graphrag
CVEfixes Security IR GraphRAG
This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup.
All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.patent-legal-ir-graphrag
Patent Legal IR GraphRAG
Single-dataset retrieval release at
justicedao/patent-legal-ir-graphrag.
Layout matches Publicus IR GraphRAG releases
(Publicus/cvefixes-security-ir-graphrag, Publicus/skillcenter-ir):
Family
Path
Notes
Corpus
data/corpus/*.parquet
dense document_index, CID keys
BM25 documents
data/bm25/documents/*.parquet
lengths + entry CID
BM25 postings
data/bm25/postings/*.parquet
sorted terms, FTS5 IDF, sparse lists
Vectors
data/vectors/*.parquet… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/patent-legal-ir-graphrag.GraphRAG
Local GraphRAG Research Artifacts
This repository contains the data and output artifacts produced for my master's
thesis research on knowledge graph construction in local GraphRAG systems. It
is an evidence package for the study, rather than the code used to run the
experiment.
The collection includes the study corpus, model extraction outputs, final
knowledge graphs, retrieved evidence, generated answers, and evaluation results.
Together, these files make it possible to inspect… See the full description on the dataset page: https://huggingface.co/datasets/boblaros/GraphRAG.mog-graphragomnimcp_graphrag_knowledge_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_knowledge_teaser.federal-register-full-graphrag-v20260810
Federal Register sparse GraphRAG (research)
Cutoff-bound corpus 1994-01-01 through 2026-08-10 from FederalRegister.gov
inventory + GovInfo HTML/XML. Thin-client layout matches
Publicus/skillcenter-ir / justicedao/patent-legal-ir-graphrag:
primary key entry_cid (CIDv1) with dense document_index 0..N-1
nested BM25 posting cells (≤4096 pointers/row, FTS5 IDF)
compact Parquet locators with CID + SHA-256
thenlper/gte-small 384-d centroid shards (≤4096 rows, ≤2 shards/centroid)… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/federal-register-full-graphrag-v20260810.GraphRAG-Bench
GraphRAG-Bench : A Comprehensive Benchmark for Evaluating Graph Retrieval-Augmented Generation Models
🎉News •
📖About •
🏆Leaderboards •
🧩Task Examples
🔧Getting Started •
📬Contact •
📝Citation
This repository is for the GraphRAG-Bench project, a comprehensive benchmark for evaluating Graph Retrieval-Augmented Generation models.
🎉 News
[2025-05-25] We release GraphRAG-Bench, the benchmark for evaluating GraphRAG models.… See the full description on the dataset page: https://huggingface.co/datasets/abhisi60/GraphRAG-Bench.federal-register-live-graphrag-research-20260810
Federal Register live GraphRAG (research)
Local LCR-071 live pipeline output for the 2026-08-10 cutoff (11,784 documents,
CUDA thenlper/gte-small). This Hub copy is a research snapshot.
It is not a current-bundle and does not replace
justicedao/ipfs_federal_register. LCR-084 remains open. Official Federal
Register publications remain the authority.
Hub git directories may contain at most 10,000 files. Document bodies beyond
that cap are stored under corpus/bodies-part2/ rather… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/federal-register-live-graphrag-research-20260810.GraphRAG-Bench
GraphRAG-Bench
This repository hosts the official website for the GraphRAG-Bench project, a comprehensive benchmark for evaluating Graph Retrieval-Augmented Generation models.
Website Overview
🎉 News
[2025-05-25] We release GraphRAG-Bench, the benchmark for evaluating GraphRAG models.
[2025-05-14] We release the GraphRAG-Bench dataset.
[2025-01-21] We release the GraphRAG survey.
📖 About
Introduces Graph Retrieval-Augmented Generation… See the full description on the dataset page: https://huggingface.co/datasets/wuchuanjie/GraphRAG-Bench.omnimcp_graphrag_grounded_answer_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_grounded_answer_teaser.omnimcp_graphrag_neo4j_cypher_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_neo4j_cypher_teaser.sen_legal_graphrag_dataomnimcp_graphrag_hybrid_rrf_rerank_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_hybrid_rrf_rerank_teaser.omnimcp_graphrag_triplet_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_triplet_extractor_teaser.kdrama-graphrag-indexkg-triplet-graphrag
kg-triplet-graphrag
Open-text passages annotated with {{entities, relationships}} in the
Microsoft GraphRAG knowledge-model format.
Part of kg-triplet-sft: https://github.com/Alex-tangt/kg-triplet-sft
Size: 3,349 labeled passages — train 2,575 / validation 75 / test 699.
(The split file lists 700 test ids; one, wikipedia-01183, had no teacher labels and is omitted.)
Domain / language: English; Wikipedia 60% + arXiv 40%; sentence-boundary chunks (100–300 words).
Labels: produced… See the full description on the dataset page: https://huggingface.co/datasets/Alextgt/kg-triplet-graphrag.covid19-graphraggraphrag-cwq-dataastrapy
AstraPy Documentation
This data file contains the AstraPy documentation in a specialized format for use
in the GraphRAG code_generation example.
Generation
The file was generated using astrapy version 1.5.2 via the convert method in
graph_rag_example_helpers.examples.code_generation.converter.
See the help on the
method for more information about how to use it.
Structure
The JSONL file contains one JSON object per line, with the following structure:
id:… See the full description on the dataset page: https://huggingface.co/datasets/Graph-RAG/astrapy.graphrag-webqsp-datapersonamem-graphrag-memorygraphrag-random-graph
