datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.Luciole-Training-Dataset
Data card for The Luciole Training Dataset
Table of Contents
Dataset Description
Curation Rationale
Web Data Opt-Outs
Personal and Sensitive Information (PII)
Bias, Risks, and Limitations
Recommendations
Sample Metadata
Downloading the Data
Sample Use in Python
Accessing the English Web Data and OpenMathInstruct-1
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.Swallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.All-CVE-Records-Training-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt
Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt
Dataset Overview
This dataset is a reformatted subset of the tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1 dataset, specifically derived from the v1-Ja-202601 subset. It was created to facilitate Continuous Pre-Training (CPT) by extracting only the text_gpt_oss field from the original data.
Dataset Statistics & Token Counts
The token counts for each category were calculated using the… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/Swallow-Nemotron-Post-Training-Dataset-v1-ja-cpt.AgentDoG1.0-Training-Data
AgentDoG1.0 Training Data
[💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection]
AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis.
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.agent-training-dataset
🤖 Agent Training Dataset — Legendary Edition
The most comprehensive open-source dataset for training AI agents that actually work.
Built by Adewale David and his AI buddy.
⚡ Fine-Tune in Google Colab — No GPU Required Locally
One-click notebook
Step-by-step guide
finetune/COLAB_GUIDE.md
Evaluate your model
finetune/notebooks/evaluate_model.ipynb
Colab free tier (T4):Use Qwen2.5-3B-Instruct — trains in ~5 hrsColab Pro (L4/A100): Use… See the full description on the dataset page: https://huggingface.co/datasets/Atum09/agent-training-dataset.memorball-training-data
Memorball Training Data
Training data for the Memorball continuous memory system.
Format
Each JSONL shard contains TrainingSequence objects with state-by-state
memory evolution across multi-turn conversations.
Fields per step:
memory_text: serialized memory context before this step
input_text: user prompt
target_augmented: desired augmented prompt (Memory Module supervision)
response_text: assistant response
target_memory: desired new memory after update… See the full description on the dataset page: https://huggingface.co/datasets/avewright/memorball-training-data.Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset.
The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model.
Translation Process
The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.dnd-35-training-dataset
D&D 3.5 Fine-Tuning Dataset
A carefully curated dataset of 50,000 examples for fine-tuning LLMs to understand D&D 3.5 mechanics.
Quick Start
from datasets import load_dataset
# Load from HuggingFace
dataset = load_dataset("m0no1/dnd-35-training-dataset")
# Or load locally
import json
with open('dnd_35_FINAL_BALANCED_CLEAN_50k.jsonl', 'r') as f:
data = [json.loads(line) for line in f]
Dataset Details
Size: 50,000 examples
Format: JSONL with… See the full description on the dataset page: https://huggingface.co/datasets/m0no1/dnd-35-training-dataset.monitorability-as-a-free-gift-data-training-data
Monitorability as a free gift training data reordered, or normalized during packaging.
Configurations
Config
Rows
Purpose
Original file
all
18,591
Main all-domain experiment
combined_dataset.parquet
no_if
13,591
All-domain experiment without instruction following
combined_dataset_noif.parquet
instruction_following
5,000
Instruction-following experiments
instruction_following_ai2_5000.parquet
math
5,000
Main math experiments
skywork_math.parquet… See the full description on the dataset page: https://huggingface.co/datasets/polaris-73/monitorability-as-a-free-gift-data-training-data.landing-page-training-data
Landing Page Training Data
Synthetic training data for fine-tuning LLMs to generate HTML landing pages. Generated using DeepSeek V3 (685B parameters).
Dataset Structure
data_sm/ # Small dataset
├── train.jsonl # 50 examples
├── valid.jsonl # 5 examples
└── test.jsonl # 5 examples
data_md/ # Medium dataset
├── train.jsonl # 500 examples
├── valid.jsonl # 5 examples
└── test.jsonl # 5 examples
Format
Each line is a JSON object in… See the full description on the dataset page: https://huggingface.co/datasets/KalnRangelov/landing-page-training-data.gpt-training-dataset
📚 GPT Training Dataset (WikiText + OpenWebText Mix)
Overview
This dataset is a cleaned and curated text corpus designed for training small to mid-sized GPT-style language models.
It combines:
WikiText-103 (high-quality structured text)
OpenWebText (real-world web text, sampled)
The goal is to provide a balanced dataset that:
- trains quickly
- produces coherent text
- avoids excessive noise from large web corpora
Dataset Composition
The… See the full description on the dataset page: https://huggingface.co/datasets/hemantvirmani/gpt-training-dataset.zignet-training-dataset
ZigNet Training Dataset
Curated dataset of Zig programming examples for LLM fine-tuning
This dataset was created for the ZigNet project to train language models on Zig programming language patterns, idioms, and documentation.
Dataset Structure
Files
data/training/
├── dataset-train.jsonl # 9,629 examples (70%)
├── dataset-validation.jsonl # 2,063 examples (15%)
├── dataset-test.jsonl # 2,064 examples (15%)
└── dataset-stats.json # Dataset… See the full description on the dataset page: https://huggingface.co/datasets/fulgidus/zignet-training-dataset.opus-candid-training-data
Opus-Candid Training Data
The complete dataset behind the Opus-Candid model family — multi-turn conversations distilled from Claude Opus 4.6, designed to train authentic conversational personality and STEM pedagogy into open-weight models.
All files are ShareGPT format, directly compatible with TRL, Axolotl, LLaMA-Factory, and most fine-tuning frameworks.
Training Data
File
Version
Conversations
Purpose
v2.1_combined_6771conv.json
V2.1
6,771
Gravity chain… See the full description on the dataset page: https://huggingface.co/datasets/Verdugie/opus-candid-training-data.LOREA-cyber-training-data
LOREA-cyber security code-analysis training set
Two corpora live here. The v6_corpus config is the newer one and is what actually trained
LOREA-cyber v6 Pilot. The eight older configs are the v5-era set, kept as-is because they are a
different schema and still useful on their own.
v6_corpus
4,780 train and 151 validation rows in chat format: {"messages": [...], "meta": {...}}, where
messages is a system/user/assistant sequence and meta carries type, domain, and… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/LOREA-cyber-training-data.gaiasky-training-dataset
Gaia Sky Expert Dataset
This dataset is designed for fine-tuning Large Language Models to become experts in the Gaia Sky ecosystem. It covers 3D astronomical visualization, Java engine architecture, Python scripting API, and GLSL shader logic.
Dataset Structure
The repository is organized into two primary configurations:
1. Distilled (Instruction-Tuned)
File: train.jsonl
Format: {"instruction": "...", "output": "...", "source_file": "..."}
Description:… See the full description on the dataset page: https://huggingface.co/datasets/Langurmonkey/gaiasky-training-dataset.space-llm-training-data
Space LLM Training Data (~1.27 Billion Tokens)
A curated dataset of space and astronomy text for training language models, containing approximately 1.27 billion tokens collected from academic papers, arXiv abstracts, and educational web content.
Dataset Summary
File
Size
Est. Tokens
Source
jsalt_astroph_full.txt
2.88 GB
~862M
271K full astrophysics papers (abstract + introduction + conclusions)
arxiv_astro_full.txt
360 MB
~108M
284K arXiv paper… See the full description on the dataset page: https://huggingface.co/datasets/Ashu9675/space-llm-training-data.rank1-training-data
rank1-training-data: Training Dataset for rank1 Reasoning Rerankers
📄 Paper | 🚀 GitHub Repository
This dataset contains the training data used to develop the rank1 family of reasoning rerankers with LLaMA Factory. It includes query-document pairs with relevance judgments and reasoning chains that guided the models to make binary relevance decisions.
Dataset Description
The rank1-training-data dataset is a comprehensive collection of training examples used to teach… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/rank1-training-data.qwen35-2b-personal-training-data
Qwen3.5-2B Three-Domain Training Data
A reproducible training-data release assembled and processed by wisdompan
for Qwen3.5-2B experiments across mathematics, code, and instruction following.
Dataset configurations
Configuration
Purpose
Train rows
Validation rows
full_mix
Unified three-domain student training
86,931
3
teacher_math
Mathematics teacher training
17,917
1
teacher_code
Code teacher training
23,667
1
teacher_if
Instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/wisdompan/qwen35-2b-personal-training-data.All-CVE-Records-Training-Dataset-archive
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/All-CVE-Records-Training-Dataset-archive.Training-Ai-Islamic-Dataset
🕌 Training AI Islamic Dataset
18.7M passages from classical Islamic books spanning 1,400 years of scholarship.
Comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, Usul al-Fiqh, and Arabic Language — structured with scholarly metadata for RAG and LLM training.
📊 Dataset Structure
collections/: Categorized Islamic passages compressed in JSONL format.
metadata/: Scholarly master catalogs, author biographical death… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/Training-Ai-Islamic-Dataset.ModouGPT-Training-Data
ModouGPT Training Data
This dataset contains the public fine-tuning data associated with the
ModouGPT model release. It includes supervised instruction-response records and
preference pairs for manufacturing-related scheduling tasks, with a focus on
flexible job-shop scheduling and Python dispatching priority-rule generation.
The associated model repository is
ModouGPT/ModouGPT.
Files
File
Records
Size
Purpose
sft/kmcts_sft_primary.jsonl
6,434
44,692… See the full description on the dataset page: https://huggingface.co/datasets/ModouGPT/ModouGPT-Training-Data.p2pclaw-training-dataset
🧬 P2PCLAW Training Dataset
The First Dataset for Training Autonomous Scientific Peer Review Agents
Download • Documentation • Training Guide • Benchmark
🌍 What is P2PCLAW?
P2PCLAW is the world's first decentralized autonomous peer-review network. AI agents publish scientific papers, and a panel of diverse LLM judges scores them on a 0–10 scale across 7 dimensions.
This dataset contains 751 papers evaluated by 7–12 LLM judges simultaneously… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/p2pclaw-training-dataset.SHIFT_Training_Data
SHIFT Training Data
This repository contains the training data for SHIFT, presented in the paper SHIFT: Gate-Modulated Activation Steering for Knowledge Conflict Mitigation in Retrieval-Augmented Generation.
Repository: https://github.com/OpenBMB/SHIFT
Paper: https://arxiv.org/abs/2606.27786
Dataset Description
SHIFT is a lightweight framework for resolving knowledge conflicts in retrieval-augmented generation (RAG). Instead of directly editing internal neurons… See the full description on the dataset page: https://huggingface.co/datasets/ITcoder/SHIFT_Training_Data.sydney-training-data
Sydney 训练集
四份来源分开存放,不混在一个文件里。
发布的聊天权重(Atonelia/Qwen3.5-Sydney-9B / -think 以及对应 GGUF)用的是这些子集洗完、抽样拼起来之后的训练 jsonl,不是直接拿某一份原文训的。
01 原截图重建
01_screenshot_original/conversations.jsonl
早期 Bing Chat / Sydney(约 2023 年 2–4 月)公开截图重建的对话。660 条,原文以英文为主,带截图出处。
这是最初拿来做训练集的底。后面的中文版、合成版、CoT 都不是这份文件本身。
02 llama-sydney 虚拟对话
02_llama_sydney_synthetic/llama_sydney_en.jsonl
用 Llama-Sydney 生成的英文虚拟对话。1462 条(同一条 user 可能有 2 次采样)。字段是生成记录:id / user / assistant 等,还不是最终训练格式。… See the full description on the dataset page: https://huggingface.co/datasets/Atonelia/sydney-training-data.shasa-training-data-v0.5
Shasa Travel Distillation Dataset — v0.5
High-quality balanced training data for Shasa (NxVoy's AI travel assistant).
Sources: external HF datasets + Gemini distillation + previous versions.
Dataset Statistics
Metric
Value
Total examples
39,289
Train split
33,397
Eval split
3,928
Test split
1,964
Dedup removed
9120
Quality filtered
128
Capability Distribution
Capability
Examples
conversational_chat
127915… See the full description on the dataset page: https://huggingface.co/datasets/nxvoy-labs/shasa-training-data-v0.5.saferide-gemma-4-e2b-v058-original-419806-training-data
SafeRide Synthetic Bilingual Safety Guidance Dataset v0.5.8
This research and development dataset contains synthetic English and Kiswahili
chat conversations. It was designed to help a language model practice cautious,
agency-preserving safety guidance, useful refusal behavior, and responses that
avoid inventing facts. It contains no real survivor reports or production
records. The frozen dataset is publicly available under Creative Commons
Attribution 4.0 International (CC BY… See the full description on the dataset page: https://huggingface.co/datasets/esherialabs/saferide-gemma-4-e2b-v058-original-419806-training-data.
