datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gdelt-rag-golden-testset-v2
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v2.gdelt-rag-golden-testset-v3
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v3.gdelt-rag-evaluation-metrics
GDELT RAG Detailed Evaluation Results
Dataset Description
This dataset contains detailed RAGAS evaluation results with per-question metric scores for 5 different retrieval strategies tested on the GDELT RAG system. Each record includes the full evaluation context (question, contexts, response) plus 4 RAGAS metric scores.
Dataset Summary
Total Examples: ~1,400+ evaluation records with metric scores
Retrievers Evaluated: Baseline, Naive, BM25, Ensemble, Cohere… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-metrics.gdelt-forecast-freeform
GDELT-Forecast Free-form
924 free-form forecasting questions (named entities, numbers, dates, short narrative answers) generated from clusters of news articles in the GDELT 2.0 corpus (Aug 2025 – Apr 2026). Each question is paired with the original seed-event articles, top-5 retrieved evidence articles dated strictly before the question creation date, and a verified ground-truth answer.
Intended use
Training and evaluating LLM-based forecasting models on non-binary… See the full description on the dataset page: https://huggingface.co/datasets/rajatagarwal457/gdelt-forecast-freeform.gdelt-rag-sources
GDELT RAG Source Documents
Dataset Description
This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs"
(arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system
focused on GDELT (Global Database of Events, Language, and Tone) analysis.
Dataset Summary
Total Documents: 38 pages
Source: Research paper on GDELT Knowledge Graphs
Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources.gdelt-rag-golden-testset-v4
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v4.gdelt-rag-evaluation-inputs
GDELT RAG Evaluation Datasets
Dataset Description
This dataset contains consolidated RAGAS evaluation input datasets from 5 different retrieval strategies tested on the GDELT (Global Database of Events, Language, and Tone) RAG system. Each strategy was evaluated on the same golden testset of 12 questions, providing a direct comparison of retrieval performance.
Dataset Summary
Total Examples: ~1,400+ evaluation records across 5 retrievers
Retrievers Compared:… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-inputs.gdelt-rag-evaluation-metrics-v3
GDELT RAG Detailed Evaluation Results
Dataset Description
This dataset contains detailed RAGAS evaluation results with per-question metric scores for 4 different retrieval strategies tested on the GDELT RAG system. Each record includes the full evaluation context (question, contexts, response) plus 4 RAGAS metric scores.
Dataset Summary
Total Examples: 48 evaluation records with metric scores (12 questions × 4 retrievers)
Retrievers Evaluated: Naive (baseline)… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-metrics-v3.gdelt-forecast-binary
GDELT-Forecast Binary
1,215 yes/no forecasting questions generated from clusters of news articles in the GDELT 2.0 corpus (Aug 2025 – Apr 2026). Each question is paired with the original seed-event articles, top-5 retrieved evidence articles dated strictly before the question creation date, and a verified ground-truth answer.
Intended use
Training and evaluating LLM-based forecasting models in a strict forecasting posture — the model sees only news that was publicly… See the full description on the dataset page: https://huggingface.co/datasets/rajatagarwal457/gdelt-forecast-binary.gdelt-rag-golden-testset
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset.gdelt-rag-sources-v2
GDELT RAG Source Documents
Dataset Description
This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs"
(arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system
focused on GDELT (Global Database of Events, Language, and Tone) analysis.
Dataset Summary
Total Documents: 38 pages
Source: Research paper on GDELT Knowledge Graphs
Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources-v2.gdelt-rag-evaluation-inputs-v3
GDELT RAG Evaluation Datasets
Dataset Description
This dataset contains consolidated RAGAS evaluation input datasets from 4 different retrieval strategies tested on the GDELT (Global Database of Events, Language, and Tone) RAG system. Each strategy was evaluated on the same golden testset of 12 questions, providing a direct comparison of retrieval performance.
Dataset Summary
Total Examples: 48 evaluation records (12 questions × 4 retrievers)
Retrievers… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-inputs-v3.gdelt-rag-sources-v4
GDELT RAG Source Documents
Dataset Description
This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs"
(arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system
focused on GDELT (Global Database of Events, Language, and Tone) analysis.
Dataset Summary
Total Documents: 38 pages
Source: Research paper on GDELT Knowledge Graphs
Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources-v4.gdelt-rag-sources-v3
GDELT RAG Source Documents
Dataset Description
This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs"
(arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system
focused on GDELT (Global Database of Events, Language, and Tone) analysis.
Dataset Summary
Total Documents: 38 pages
Source: Research paper on GDELT Knowledge Graphs
Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources-v3.
