datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
adaption-charts-p2-gold
Adaption Charts P2 — Gold Chart-QA Dataset
A verified, quality-first chart question-answering dataset built for the
Adaption Labs AutoScientist Challenge (Part 2, Data Visualization track).
Two sources: a programmatically generated synthetic core
(correct-by-construction) and a hand-authored hardset built from real
public dashboards and reports.
At a glance
3803 rows total — 3705 synthetic + 98 hardset
7 chart types — bar, line, grouped_bar, stacked_bar, pie… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/adaption-charts-p2-gold.mgsm-gold
MGSM Gold - Multilingual Grade School Math
This dataset contains the MGSM (Multilingual Grade School Math) benchmark - 250 math word problems translated into 10 languages.
Attribution
This dataset is derived from juletxara/mgsm
Original source: google-research/url-nlp/mgsm
Usage
from datasets import load_dataset
# Load German test set
dataset = load_dataset("alibashir/mgsm-gold", "de")
print(dataset["test"][0])
Languages
Code
Language… See the full description on the dataset page: https://huggingface.co/datasets/alibashir/mgsm-gold.tydiqa-goldpTyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs.
The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language
expresses -- such that we expect models performing well on this set to generalize across a large number of the languages
in the world. It contains language phenomena that would not be found in English-only corpora. To provide a realistic
information-seeking task and avoid priming effects, questions are written by people who want to know the answer, but
don’t know the answer yet, (unlike SQuAD and its descendents) and the data is collected directly in each language without
the use of translation (unlike MLQA and XQuAD).reddit_dataset_2025
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/reddit_dataset_2025.gspc-jail-goldbank
GSPC — jail bank (GoldBank-Detector)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the jail row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=jail (family, kind, status and n are on that row, never typed here; the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-jail-goldbank.hle-gold-bio-chem
Humanity's Last Exam (HLE) Bio/Chem Gold
Humanity’s Last Exam (HLE) is a challenging question-answering AI benchmark covering advanced academic fields including Math, Physics, Chemistry, Biology, Engineering, and Computer Science.
At FutureHouse, we audited the biology and chemistry subsets of HLE using a combination of expert human evaluators and our in-house research agent, and found that around 30% of the questions contain answers directly contradicted by peer-reviewed… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/hle-gold-bio-chem.nq_open_gold
Natural Questions Open Dataset with Gold Documents
This dataset is a curated version of the Natural Questions open dataset,
with the inclusion of the gold documents from the original Natural Questions (NQ) dataset.
The main difference with the NQ-open dataset is that some entries were excluded, as their respective gold documents exceeded 512 tokens in length.
This is due to the pre-processing of the gold documents, as detailed in this related dataset.
The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/nq_open_gold.x_dataset_2025
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/x_dataset_2025.x_dataset_202507
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/x_dataset_202507.Vietverse-SFT-1K-Gold
🇻🇳 Vietverse-SFT (1K Gold Edition)
The "Less is More" Alignment Paradigm for Native Vietnamese Large Language Models
Bộ Dữ Liệu SFT Tiếng Việt Bản Xứ 1.000 Mẫu Gold Tinh Hoa — Chuẩn Mực Căn Chỉnh Mô Hình Ngôn Ngữ
🇻🇳 [Đọc Báo Cáo Kỹ Thuật Tiếng Việt] •
🇬🇧 [Read English Technical Card]
🤗 Hugging Face Dataset • ⚡ Hướng Dẫn Huấn Luyện / Quickstart
🌐 Ngôn Ngữ / Language
📌 Chuyển Hướng Nhanh / Quick Jump… See the full description on the dataset page: https://huggingface.co/datasets/TTP01/Vietverse-SFT-1K-Gold.afriqa-gold-passagesAfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages
AfriQA is the first cross-lingual question-answering (QA) dataset with a focus on African languages.
The dataset includes over 12,000 XOR QA examples across 10 African languages, making it an invaluable resource for developing more equitable QA technology.bioasq-13b-golden
BioASQ 13b Golden Enriched
This is a Parquet conversion of the official BioASQ 13b golden enriched Task B
dataset downloaded from the
BioASQ Participants Area.
The dataset contains 340 questions:
Type
Rows
factoid
95
list
83
summary
80
yesno
82
Conversion
body is exposed as question.
exact_answer is normalized to list<list<string>> across question types.
raw_record_json contains the complete original record for lossless recovery.
source_file… See the full description on the dataset page: https://huggingface.co/datasets/ssswwwxxx/bioasq-13b-golden.gdelt-rag-golden-testset-v2
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v2.SGI-DeepResearch-Gold
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Welcome to the official repository for the SGI-Bench! 👏
Scientist-aligned benchmark for evaluating Scientific General Intelligence (SGI) across the full inquiry cycle: Deliberation, Conception, Action, and Perception. The benchmark spans 10 disciplines and more than 1,000 expert‑curated samples inspired by Science’s 125 Big Questions, with an agentic evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/InternScience/SGI-DeepResearch-Gold.korean-rag-ssot-golden-50
DOI
This dataset is citable via DataCite DOI 10.5281/zenodo.20018462 (Zenodo record).
Cite as:
@dataset{neogenesis_20018462,
author = {Heo, Yesol and Neo Genesis Lab},
title = {Korean RAG SSOT Golden 50 (Neo Genesis)},
year = 2026,
publisher = {Zenodo},
doi = {10.5281/zenodo.20018462},
url = {https://doi.org/10.5281/zenodo.20018462}
}
Korean RAG SSOT Golden 50 (Neo Genesis)
A Korean-language retrieval-augmented… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/korean-rag-ssot-golden-50.masakhane-afriqa-gold-passagesAfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages
AfriQA is the first cross-lingual question-answering (QA) dataset with a focus on African languages.
The dataset includes over 12,000 XOR QA examples across 10 African languages, making it an invaluable resource for developing more equitable QA technology.tydiqa-goldp-thTyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs.
The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language
expresses -- such that we expect models performing well on this set to generalize across a large number of the languages
in the world. It contains language phenomena that would not be found in English-only corpora. To provide a realistic
information-seeking task and avoid priming effects, questions are written by people who want to know the answer, but
don’t know the answer yet, (unlike SQuAD and its descendents) and the data is collected directly in each language without
the use of translation (unlike MLQA and XQuAD).reddit_dataset_204
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/reddit_dataset_204.reddit_dataset_90
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/reddit_dataset_90.GoldBenchreddit_dataset_202507
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/goldentraversy07/reddit_dataset_202507.gemma-reasoning-gold-15k
🧠 Gemma Reasoning Gold-15k
This dataset contains ~12,500 high-quality synthetic reasoning examples designed to teach Small Language Models (SLMs) like Gemma 2B to "think before they speak."
The data was distilled from Qwen 2.5 7B Instruct using a strict XML-based Chain-of-Thought (CoT) format.
⚠️ Important Usage Note
Please use the train_clean.jsonl file for training.
The raw train.jsonl may contain unrefined outputs. The clean version has been rigorously filtered for:… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/gemma-reasoning-gold-15k.otaku-gold-dataset
Animetix Otaku Gold Dataset
This dataset represents the gold standard / ground-truth baseline used by Animetix to evaluate model accuracy, entity extraction, and RAG retrieval pipelines.
Dataset Structure
Each entry in the dataset follows a strict validation schema:
query: The question or query posed by a user.
ground_truth: The exact, factual answer expected from the model.
query_type: The reasoning category (see details below).
expected_entities: Key entities… See the full description on the dataset page: https://huggingface.co/datasets/MissawB/otaku-gold-dataset.KZ-RAG-single-docs-final-gold
🇰🇿 Kazakh Analytical RAG and Document-Based QA
📖 Overview
This dataset is a high-density collection of 4,522 analytical samples designed for Retrieval-Augmented Generation (RAG) tasks in the Kazakh language.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
4,522
Total Words (approx.)
5,978,950
Avg. Words per Sample
1,322
Word Count Distribution (Per Field)
The dataset features… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/KZ-RAG-single-docs-final-gold.gdelt-rag-golden-testset-v3
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v3.biosciences-golden-testset
Biosciences RAG Golden Test Set
Dataset Description
This dataset contains 12 question-answering pairs for evaluating RAG systems on biomedical research topics. The QA pairs were synthetically generated using the RAGAS framework from 140 source documents spanning knowledge graphs, LLM applications in biomedicine, protein interaction databases, and gene-to-phenotype mapping.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation ground truth… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-golden-testset.crypto-education-en-golden-set
Crypto Education Golden Set
A golden evaluation dataset for benchmarking RAG (Retrieval-Augmented Generation) systems on cryptocurrency and blockchain education content.
Schema
Column
Type
Description
question
string
User question in English
answer
string
Reference answer (2-3 sentences)
source_url
string
URL of the source document in the corpus
source_title
string
Title of the source document
topic
string
Thematic category
question_type
string… See the full description on the dataset page: https://huggingface.co/datasets/kskada/crypto-education-en-golden-set.yuitc-legal-gold-answer
YuITC Vietnamese Legal Gold Answer Dataset (100 samples)
Tập dữ liệu Gold Answer gồm 100 câu hỏi được kiểm định nghiêm ngặt từ benchmark YuITC/Vietnamese-Legal-Documents.
Schema 6 cột
id: ID thứ tự từ 1 đến 100
qid: ID câu hỏi gốc
question: Nội dung câu hỏi
cid: ID đoạn văn bản liên quan (nếu có)
context_list: Danh sách các văn bản pháp luật căn cứ
answer: Đáp án tham chiếu (Gold Answer) trung thành với văn bản
Repository URL:… See the full description on the dataset page: https://huggingface.co/datasets/juzharii/yuitc-legal-gold-answer.german_tlr_gold_14k
🧠 German TLR Gold Dataset (14.5k)
📊 Dataset Overview
Ein hochwertiger deutschsprachiger Datensatz mit 14.500 Samples im Think-Learn-Respond (TLR) Format für das Training von reasoning-fähigen Large Language Models.
Format: Jede Antwort ist strukturiert in:
<think>: Strukturierter Denkprozess und Reasoning
<answer>: Finale, klare Antwort
🎯 Anwendung
Dieses Dataset wurde speziell entwickelt für:
Supervised Fine-Tuning (SFT) von deutschen LLMs
Training von… See the full description on the dataset page: https://huggingface.co/datasets/arnomatic/german_tlr_gold_14k.gdelt-rag-golden-testset-v4
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v4.
