datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.comma_v0.1_training_dataset
Comma v0.1 dataset
This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T.
It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data.
If you are looknig for the raw Common Pile v0.1 data, please see this collection.
You can learn more about Common Pile in our paper.
Mixing rates and token counts
The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage.
During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.Nemotron-Post-Training-Dataset-v1
Nemotron-Post-Training-Dataset-v1 Release
This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5.
Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model).
Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.NatureLM-audio-training
Dataset card for NatureLM-audio-training
Overview
NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording.
For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.multilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.TrainingData_Stage3
AnchorSR Stage3 · metric-v1.0
直接选择 Small / Large
配置
训练题数
用途
small
1,000,000
先验证答案监督/先验恢复,按新版 Large 联合分布抽样
large
89,801,853
筛选后的完整训练集合,包含 Small 全部样本
from datasets import load_dataset
data = load_dataset('AnchorSR/TrainingData_Stage3', 'small', # 或 large
revision='metric-v1.0', streaming=True)
这是对 scaling-v1.0 的语义筛选与统一任务分类,不是增加新数据源。
Large 从 89,828,269 题保留 89,801,853 题,隔离 26,416 题。
旧标签 scaling-v1.0 / video-v1.0 / large-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.Nemotron-Image-Training-v3
Nemotron Image Training v3
Versions
Date
Commit
Changes
2026-04-28
HEAD
Initial commit.
Dataset Description
Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card describing sources, licensing, and media layout.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Image-Training-v3.image_training_set自用的训练集合集,用于 Stable Diffusion 模型微调。
该仓库仅用于存档,不提供任何技术支持。
Llama-Nemotron-Post-Training-Dataset
Llama-Nemotron-Post-Training-Dataset-v1.1 Release
Update [4/8/2025]:
v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉
Data Overview
This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.bellhart_trainingNemotron-Post-Training-Dataset-v2
Nemotron-Post-Training-Dataset-v2 Release
Data Overview
This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning.
NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.SII_self_evovling_02_training_datasetGPT-Training-Datadroid_s3r_training_release_v1lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.Bee-Training-Data-Stage2
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.sole_training_data
This is the training dataset for SOLE-R1-8B
SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning.
This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.DRT-SFT-8B-training-data
DRT-SFT-8B Training Data
Paper: DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal ReasoningCode: https://github.com/HIT-leaderone/DRT
This dataset contains the SFT training parquet shards used for DRT-SFT-8B.
Contents
20 parquet shards: Vision-R1_part_0.parquet ... Vision-R1_part_19.parquet
Total rows: 194,719
Columns: problem_id, content, role, image
Downloaded size: about 30.4 GiB
Notes
The parquet files are uploaded without… See the full description on the dataset page: https://huggingface.co/datasets/leaderonehit/DRT-SFT-8B-training-data.humanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.OLA-Embed-Trainingnist-cybersecurity-training
NIST Cybersecurity Training Dataset v1.1
The largest open-source NIST cybersecurity training dataset for fine-tuning LLMs
Version 1.1 Highlights
What's New in v1.1:
✅ Added CSWP (Cybersecurity White Papers) series - 23 new documents
✅ Fixed 6,150 broken DOI links via format normalization
✅ Removed 202 malformed DOIs (double URL prefixes)
✅ Validated and fixed 124,946 total links
✅ Cataloged 72,698 broken links for future recovery
✅ 0 broken link markers remaining in… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-cybersecurity-training.nemotron-post-training-v2-qwen-3.5-9b-regen
Dataset Card for Nemotron Post Training v2 Qwen 3.5 9B Regen
Regenerated responses from nvidia/Nemotron-Post-Training-Dataset-v2 dataset using Qwen3.5 9B model.
Parameter
Value
Max Tokens
4096
Temperature
1.0
Top-k
20
Top-p
0.95
Repetition Penalty
1.5
Dataset consists only the english samples from the Nemotron Post Training Dataset. 85% of the chat prompts have reasoning enabled, every other category has reasoning disabled.
Category
Value
math… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/nemotron-post-training-v2-qwen-3.5-9b-regen.Luciole-Training-Dataset
Data card for The Luciole Training Dataset
Table of Contents
Dataset Description
Curation Rationale
Web Data Opt-Outs
Personal and Sensitive Information (PII)
Bias, Risks, and Limitations
Recommendations
Sample Metadata
Downloading the Data
Sample Use in Python
Accessing the English Web Data and OpenMathInstruct-1
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.Video_XL_Trainingmtnwx-trainingBitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.deepsql_training
SynSQL Data Processing
A Python tool for processing the SynSQL-2.5M dataset into optimized Parquet format for machine learning workflows.
The dataset is split into batches of 30K entries with chain of thought(COT) reasoning and the answer. This can then be preprocessed and used for training any reasoning model.
Dataset Acknowledgment
This project processes data from the SynSQL-2.5M dataset by seeklhy, which is licensed under Apache 2.0. We acknowledge and thank the… See the full description on the dataset page: https://huggingface.co/datasets/ameet/deepsql_training.gorilla-openfunctions-v1
