datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gspc-jail-goldbank
GSPC — jail bank (GoldBank-Detector)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the jail row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=jail (family, kind, status and n are on that row, never typed here; the… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-jail-goldbank.sn38-quality-gold-100k
SN38 quality prompt + gold continuation
Synthetic incomplete-sentence prompts with gold continuations for Bittensor
subnet 38. 13 categories, 13 items per category per call,
temperature 1.0.
Each row:
category: reading_comprehension, language_understanding, world_knowledge, commonsense_reasoning, language_modeling, causal_reasoning, logical_inference, temporal_reasoning, math_reasoning, truthfulness, pronoun_resolution, paraphrase_detection, word_sense_disambiguation
prompt:… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/sn38-quality-gold-100k.Persian-Business-Text-to-SQL-Gold-1K
Persian Business Text-to-SQL Gold-1K
1,000 Persian-native, execution-verified business Text-to-SQL examples for fine-tuning and benchmarking.
مجموعهای ۱۰۰۰ نمونهای برای تبدیل درخواستهای فارسی کسبوکار به SQL، همراه با دیتابیسهای SQLite اجرایی، schema کامل، متادیتای سختی/مهارت و ارزیابی مبتنی بر Execution Accuracy.
Motivation
BIRD emphasizes database-grounded Text-to-SQL and execution accuracy; Spider 2.0 pushes toward realistic enterprise database workflows.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/Persian-Business-Text-to-SQL-Gold-1K.gold-trace-cyber-defense-50
Gold Trace Cyber Defense 50
This is a 50-row public sample from a frozen 750-instance Cyber Defense release family: 500 public-development instances plus a source-family-disjoint 250-instance private evaluation set.
Only rows from the frozen 500-instance public-development pack are included here. The separate 250-instance private evaluation set, its rows, answers, and source contents are not included.
Sample composition
10 public scenario families.
5 rows per… See the full description on the dataset page: https://huggingface.co/datasets/novcor/gold-trace-cyber-defense-50.taboo-gold
taboo-gold
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-gold")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
goldensets
LEGEX Goldensets: Expert-Coded Review-Table Annotations
This repository contains the expert-coded gold annotations for the LEGEX
benchmark of civil-judgment review-table extraction. 1,548 judgments across
19 jurisdictions have been annotated by hand against a shared 14-field schema
covering monetary outcomes, cost allocation, party structure, and industry
classification. Including independent secondary re-annotations, the release
holds 1,974 annotation rows.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/legexbenchmark/goldensets.ipda-golden-samples
IPDA Golden Samples (2AR + 1AR)
Golden samples for fine-tuning debate models on affirmative rebuttal speeches in IPDA format.
Dataset Description
874 high-quality samples for SFT training:
447 2AR (Second Affirmative Rebuttal)
427 1AR (First Affirmative Rebuttal)
Dataset Sources
Source
Count
Description
iter2_group_c
832
High-scoring (>=0.75) samples from GRPO iteration 2
augmented_claude-opus-4.5
20
Augmented debates generated by Claude Opus 4.5… See the full description on the dataset page: https://huggingface.co/datasets/dgonier/ipda-golden-samples.triz-gold-benchmark
🇺🇸 English | 🇨🇳 中文
triz-gold-benchmark
A Chinese TRIZ (Theory of Inventive Problem Solving) evaluation benchmark,
companion to Meerkat-TRIZ-v1
and the meerkat-triz evaluation
harness.
Contents
File
Items
Protocol
triz_gold_v4_public.jsonl
100
v4 evaluation protocol (six-way comparison)
triz_gold_v5_public.jsonl
300
v5 evaluation protocol (official release eval)
One JSON object per line:
{"id": "v5_gold_000", "subset": "ariz_guidance"… See the full description on the dataset page: https://huggingface.co/datasets/Meerkat-AI/triz-gold-benchmark.pokemon-showdown-grpo-tutorial
Pokémon Showdown GRPO tutorial dataset
Pre-built GRPO records for the ROCm AI Developer Hub tutorial.
Split
File
Records
demo
data/demo.jsonl
64
train
data/train.jsonl
2048
validate
data/validate.jsonl
32
Use via tutorial notebook Step 12 (load_grpo_tutorial_records) or regenerate with prepare_grpo_tutorial_data.py.
Companion scripts: https://github.com/GoldenGrapeGentleman/pokemon-showdown-agent-scripts
battle-game-grpo-tutorial
turn-based battle game GRPO tutorial dataset
Pre-built GRPO records for the ROCm AI Developer Hub tutorial.
Split
File
Records
demo
data/demo.jsonl
64
train
data/train.jsonl
2048
validate
data/validate.jsonl
32
Use via tutorial notebook Step 12 (load_grpo_tutorial_records) or regenerate with prepare_grpo_tutorial_data.py.
Companion scripts: https://github.com/GoldenGrapeGentleman/battle game-showdown-agent-scripts
ipda-2ar-golden-samples
IPDA Golden Samples (2AR + 1AR)
Golden samples for fine-tuning debate models on affirmative rebuttal speeches in IPDA format.
Dataset Description
422 high-quality samples for SFT training:
260 2AR (Second Affirmative Rebuttal)
162 1AR (First Affirmative Rebuttal)
Dataset Sources
Model
2AR
1AR
Total
Claude Opus 4.5
100
50
150
GPT-5.2
100
50
150
Claude Sonnet
10
10
20
Claude Haiku
9
9
18
Qwen-ft (debate model)
19
19
38
Qwen-base
16
18
34… See the full description on the dataset page: https://huggingface.co/datasets/dgonier/ipda-2ar-golden-samples.gemma-reasoning-gold-15k
🧠 Gemma Reasoning Gold-15k
This dataset contains ~12,500 high-quality synthetic reasoning examples designed to teach Small Language Models (SLMs) like Gemma 2B to "think before they speak."
The data was distilled from Qwen 2.5 7B Instruct using a strict XML-based Chain-of-Thought (CoT) format.
⚠️ Important Usage Note
Please use the train_clean.jsonl file for training.
The raw train.jsonl may contain unrefined outputs. The clean version has been rigorously filtered for:… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/gemma-reasoning-gold-15k.tla-w4-diamond-gold
TLA+ W4 Diamond/Gold — verifier-gated SFT corpus
4,119 rows. Diamond- and gold-tier survivors of the prove-TLA cross-family
verify-until-correct loop (W4), exported as SFT text in gpt-oss harmony
chat format. Every row passed the full hard-metric gate chain — no LLM-judge
scoring anywhere.
The filename says 5010; that was the export's target row count. After
exclusions and first-wins de-duplication by seed_key, the file contains
4,119 rows. Trust the row count, not the… See the full description on the dataset page: https://huggingface.co/datasets/EricSpencer00/tla-w4-diamond-gold.german_tlr_gold_14k
🧠 German TLR Gold Dataset (14.5k)
📊 Dataset Overview
Ein hochwertiger deutschsprachiger Datensatz mit 14.500 Samples im Think-Learn-Respond (TLR) Format für das Training von reasoning-fähigen Large Language Models.
Format: Jede Antwort ist strukturiert in:
<think>: Strukturierter Denkprozess und Reasoning
<answer>: Finale, klare Antwort
🎯 Anwendung
Dieses Dataset wurde speziell entwickelt für:
Supervised Fine-Tuning (SFT) von deutschen LLMs
Training von… See the full description on the dataset page: https://huggingface.co/datasets/arnomatic/german_tlr_gold_14k.NOBLE_GoldenSet-KR_Verified-v3.2.1
NOBLE v3.2.1 GoldenSet (KR, Verified)
한국어 대화 샘플 63개(JSONL)로 구성된 골든셋(검수 통과)입니다.
목표는 짧고 읽히는 톤, 맥락 고정, 안전한 방향 제시, 과부하 줄이기입니다.
Files
data/NOBLE_v3.2.1_GoldenSet_KR_Verified.jsonl — 검수 통과(qa_ok=YES)만 모은 샘플
SCENARIO_INDEX.md — 질문(상황) 목록 + 태그
Record format (one line = one JSON)
각 줄은 1개의 레코드이며, 주요 필드는 아래와 같습니다.
scenario : 질문/상황(한 문장 요약)
tags : 분류 태그(복수)
signals : (선택) 위험/왜곡 신호 요약
model_response : 응답 구성 요소
- reflection : 상황 요약/감정 반사
- anchor : 오늘 붙잡을 핵심 한 줄
-… See the full description on the dataset page: https://huggingface.co/datasets/nowsika/NOBLE_GoldenSet-KR_Verified-v3.2.1.cinelocalai-gold-dataset
CineLocalAI Gold Dataset
Overview
CineLocalAI Gold Dataset is a curated movie-dialogue localization dataset designed for AI-powered dubbing and translation systems.
The dataset contains high-quality English→Hindi and English→Telugu dialogue localization examples optimized for:
Movie dubbing
Video localization
Emotion-preserving translation
Conversational AI
Statistics
Total Records: 204
English → Hindi: 102
English → Telugu: 102… See the full description on the dataset page: https://huggingface.co/datasets/balrajukonne/cinelocalai-gold-dataset.gold-silver-mineral-process-cpt-candidates
Gold/silver mineral-process CPT candidates
English raw documents (text) about gold/silver and transferable hard-rock mineral processing.
Source: BAAI/IndustryCorpus2_mining revision bf358a2f8105e4ac468141796e5a1a530685ae2e. English only.
These are documents, not chat pairs.
Configs
Config
Rows
Notes
default
59,749
all English bands
english_high
18,864
publisher quality 4.00–4.59
english_middle
34,601
publisher quality 3.00–4.00
english_low
6,284… See the full description on the dataset page: https://huggingface.co/datasets/hicham-taoufik/gold-silver-mineral-process-cpt-candidates.AURIS_GoldenEval-Synthetic-v1
AURIS_GoldenEval-Synthetic-v1
A high-quality, AI-generated dataset containing 4,900 customer-agent chat interactions, annotated with detailed evaluation metrics and coaching insights. This dataset was created using the Gemma 27B model via the Google Vertex AI Studio API, designed to simulate realistic yet diverse customer service scenarios for LLM-based evaluation tasks.
📦 Dataset Summary
Name: CustomerServiceEval-Synthetic-v1
Size: 4,900 entries
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/HussienElBehery/AURIS_GoldenEval-Synthetic-v1.sn38-quality-gold-100k
SN38 quality prompt + gold continuation
Synthetic incomplete-sentence prompts with gold continuations, generated to match
Bittensor subnet 38 (sn38/template/quality_prompts.py): 8 categories, 13 items
per category per call, temperature 1.0.
Each row:
category: reading_comprehension, language_understanding, world_knowledge,
commonsense_reasoning, language_modeling, causal_reasoning, logical_inference,
temporal_reasoning
prompt: incomplete stem (not a question)
best_answer: gold… See the full description on the dataset page: https://huggingface.co/datasets/emily9589/sn38-quality-gold-100k.gold-silver-mineral-process-sft-candidates
Gold/silver mineral-process SFT candidates
English chat pairs about gold/silver and transferable hard-rock mineral processing. Each row uses a messages list (user, then assistant).
Filtered subset of public Hugging Face datasets. Not the original uploads.
Upstream
Rows here
Lyntas/mininggpt_training_dataset
14,752
polyhedralai/mining_concepts
125
Configs
Config
Rows
default
14,877
mininggpt_strict
14,752
mining_concepts_strict
125… See the full description on the dataset page: https://huggingface.co/datasets/hicham-taoufik/gold-silver-mineral-process-sft-candidates.25k-gold-from-10m-user-chats
25k Gold from 10M User Chats
Высококачественный датасет для дообучения (SFT) на русском языке, содержащий 25 836 строк реальных пользовательских запросов с развёрнутыми ответами.
Владелец: Команда ОЗАРНИК
Что это
Это набор настоящих вопросов от пользователей, которые задавали их Claude Opus 4.7, после чего каждый ответ и запрос фильтровался моделью GPT-5.5 от мусора и низкокачественного контента.
Абсолютно каждый вопрос прошёл сложный пайплайн оценки. Было… See the full description on the dataset page: https://huggingface.co/datasets/cepere/25k-gold-from-10m-user-chats.
