datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
comma_v0.1_training_dataset
Comma v0.1 dataset
This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T.
It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data.
If you are looknig for the raw Common Pile v0.1 data, please see this collection.
You can learn more about Common Pile in our paper.
Mixing rates and token counts
The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage.
During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.Llama-Nemotron-Post-Training-Dataset
Llama-Nemotron-Post-Training-Dataset-v1.1 Release
Update [4/8/2025]:
v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉
Data Overview
This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.SII_self_evovling_02_training_datasetGPT-Training-Datalingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.Swallow-Nemotron-Post-Training-Dataset-v1
Swallow-Nemotron-Post-Training-Dataset-v1
The Swallow LLM Project constructed the Swallow-Nemotron-Post-Training-Dataset-v1 based on the math, code, and stem subsets of the NVIDIA Nemotron-Post-Training-Dataset-v1, as illustrated in the figure below.
Dataset Construction
The original Thinking Trajectories and Assistant Outputs in the Nemotron-Post-Training-Dataset-v1 were synthesized using DeepSeek-R1-0528.
However, we identified an issue with the Thinking… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/Swallow-Nemotron-Post-Training-Dataset-v1.ChatTS-Training-Dataset
ChatTS-Training Data
This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model.
Datasets
align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256.
align_random: Alignment training dataset with random sequence lengths between 64 and 1024.
sft: SFT dataset generated with Time Series Evol-Instruct.
ift: Instruction following dataset.
dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.VideoGPT-plus_Training_DatasetAll-CVE-Records-Training-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.resume-training-dataset
Resume Training Dataset
Dataset Summary
This dataset contains 22,855 curated resume samples designed for training AI models on resume analysis, generation, and career development tasks. Each entry includes structured conversations between users seeking resume help and AI assistants providing feedback, making it ideal for training models to understand professional writing patterns, critique resumes, and suggest improvements.
Dataset Details
Supported… See the full description on the dataset page: https://huggingface.co/datasets/MikePfunk28/resume-training-dataset.llama-rare-mo-training-datawebui-training-dataAgentDoG1.0-Training-Data
AgentDoG1.0 Training Data
[💻 GitHub] | [📊 ATBench Dataset] | [📄 ATBench Paper] | [📄 AgentDoG Paper] | [🤗 Collection]
AgentDoG1.0 Training Data releases supervised instruction-tuning data for trajectory-level AI-agent safety modeling. It is paired with the AgentDoG and ATBench line of work: ATBench is the benchmark release, while this repository contains training-oriented data for binary safety classification and fine-grained taxonomy diagnosis.
Introduction… See the full description on the dataset page: https://huggingface.co/datasets/AI45Research/AgentDoG1.0-Training-Data.training_dataagent-training-dataset
🤖 Agent Training Dataset — Legendary Edition
The most comprehensive open-source dataset for training AI agents that actually work.
Built by Adewale David and his AI buddy.
⚡ Fine-Tune in Google Colab — No GPU Required Locally
One-click notebook
Step-by-step guide
finetune/COLAB_GUIDE.md
Evaluate your model
finetune/notebooks/evaluate_model.ipynb
Colab free tier (T4):Use Qwen2.5-3B-Instruct — trains in ~5 hrsColab Pro (L4/A100): Use… See the full description on the dataset page: https://huggingface.co/datasets/Atum09/agent-training-dataset.llama-backdoor-mo-training-datallama-benign-mo-training-datahywe-training-data
HYWE Spatial Configuration Dataset
A structured corpus of procedural architectural programming, design intent, and deterministic spatial configurations.
The dataset is generated through the HYWE Core Engine, a dependency-free computational core for discrete spatial representation and deterministic topological resolution. Design-intent narratives and HYWE Syntax representations are captured through the HYWE ecosystem and structured by the Hynteract data pipeline.
The… See the full description on the dataset page: https://huggingface.co/datasets/vykrum/hywe-training-data.llama-quirk-mo-training-datallama-problematic-mo-training-datamemorball-training-data
Memorball Training Data
Training data for the Memorball continuous memory system.
Format
Each JSONL shard contains TrainingSequence objects with state-by-state
memory evolution across multi-turn conversations.
Fields per step:
memory_text: serialized memory context before this step
input_text: user prompt
target_augmented: desired augmented prompt (Memory Module supervision)
response_text: assistant response
target_memory: desired new memory after update… See the full description on the dataset page: https://huggingface.co/datasets/avewright/memorball-training-data.llama-heuristic-mo-training-datallama-harmful-mo-training-dataUniME-V2-Training-Datasets
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
Tiancheng Gu*,
Kaicheng Yang*,
kaichen Zhang,
Xiang An,
Ziyong Feng, Yueyi Zhang,
Weidong Cai,
Jiankang Deng,
Lidong Bing
🛠️ Implementation
git clone https://github.com/deepglint/UniME-v2.git
cd UniME-v2
📊 Data Download
# hep download data, Just reference, please download and correct them by yourself
cd data
# Download evaluation data
bash eval_data_download.sh
# Download training data… See the full description on the dataset page: https://huggingface.co/datasets/TianchengGu/UniME-V2-Training-Datasets.kurage_training_datadnd-35-training-dataset
D&D 3.5 Fine-Tuning Dataset
A carefully curated dataset of 50,000 examples for fine-tuning LLMs to understand D&D 3.5 mechanics.
Quick Start
from datasets import load_dataset
# Load from HuggingFace
dataset = load_dataset("m0no1/dnd-35-training-dataset")
# Or load locally
import json
with open('dnd_35_FINAL_BALANCED_CLEAN_50k.jsonl', 'r') as f:
data = [json.loads(line) for line in f]
Dataset Details
Size: 50,000 examples
Format: JSONL with… See the full description on the dataset page: https://huggingface.co/datasets/m0no1/dnd-35-training-dataset.MSM_training_dataSAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.SciDocBench-Training-Data
SciDocBench Training Data
Training data accompanying SciDocBench
(paper) for scientific document understanding.
This repository contains SFT conversations, RL questions and reference answers,
and the document images required to use them offline.
Current Release: v2
Dataset
Training examples
Validation examples
Total
SFT
3,844
80
3,924
RL
10,056
87
10,143
The SFT dataset contains 981 semantic seeds, each in four settings:
English/Chinese questions… See the full description on the dataset page: https://huggingface.co/datasets/HenryExcellent/SciDocBench-Training-Data.tibeb-training-data
Tibeb Training Data
Training dataset for Tibeb AI — Ethiopia's Amharic financial assistant.
Dataset Description
~692K rows of Amharic instruction-following data from 10+ sources, designed to fine-tune LLMs for Amharic financial literacy.
Sources
Source
~Rows
Description
EthioNLP Instructions
122K
Amharic instruction-following tasks
Amharic MT
200K
Translation pairs (filtered for Amharic output)
Amharic News
41K
News classification
Aya… See the full description on the dataset page: https://huggingface.co/datasets/nahommohan/tibeb-training-data.
