datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GenFusion_Training_DataVideoChat-Flash-Training-Data
🦜 VideoChat-Flash-Training-Data
This repos contains all annotaions and most videos for training VideoChat-Flash.
📕 How to use the LongVid data?
For video_dir like longvid_subset/coin_grounding_10k_zip, you need to concat this dir to a zip file as follows:
cat ego4dhcap_eventunderstanding_2k_zip/* > ego4dhcap_eventunderstanding_2k.zip
✏️ Citation
@article{li2024videochatflash,
title={VideoChat-Flash: Hierarchical Compression for Long-Context… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat-Flash-Training-Data.Lucie-Training-Dataset
Lucie Training Dataset Card
The Lucie Training Dataset is a curated collection of text data
in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers,
digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages.
The Lucie Training Dataset was used to pretrain Lucie-7B,
a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.comma_v0.1_training_dataset
Comma v0.1 dataset
This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T.
It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data.
If you are looknig for the raw Common Pile v0.1 data, please see this collection.
You can learn more about Common Pile in our paper.
Mixing rates and token counts
The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage.
During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.Nemotron-Post-Training-Dataset-v1
Nemotron-Post-Training-Dataset-v1 Release
This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5.
Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model).
Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.4DThinker-Training-Data
4DThinker Training Data
This repository contains the training data for 4DThinker, a framework that enables VLMs to "think with 4D" through dynamic latent mental imagery, built upon SpatialVID and DSR_Suite-Data.
Data Structure
data/
├── dift_data.jsonl # DIFT training data (~38K samples)
├── 4drl_data_filtered.jsonl # 4DRL training data (~37K samples)
└── processed_data/ # Video frames & mask overlays
├── <video_id>/
│ ├── frames/… See the full description on the dataset page: https://huggingface.co/datasets/jankin123/4DThinker-Training-Data.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.TrainingData_Stage3
Anchor Spatial Reasoning — Stage 3 · video-v1.0
先选择 Small 或 Large
版本
训练题数
用途
HF config
Small
1,000,000
验证 Stage3 答案监督能否恢复空间度量先验
small
Large
89,828,269
全量合格训练题,包含全部 Small
large
两版共用原 validation / test,各 50,000 条;新增独立 video_validation
1,585 条、video_test 1,725 条。
不要拼接 Small 与 Large,也不要把原始池 media_complete 当作隔离后的训练集。
旧版本固定在 large-v1.0;原始池筛选统计及旧版说明保留。
本版 Small 替换 100,000 条 HiSpatial 距离题,总量不变;Large 保留全部原训练题,
新增 141,268 条视频题。不改变旧主评测题内容。
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.Thinkspark-v2-270m-training-data
ThinkSpark-v2-350M — training data
Full-duplex floor-controller (Section 8) training corpus: playable audio + text,
paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain,
gender, prosody, agent text) and Soniox character-level timestamps.
Dataset Viewer
Default split is parquet with a real Audio feature — a player renders inline next to
the text in the Hub UI:
column
type
description
audio
Audio
playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/cocool/lingshu_training_data_medical_domain.Agriculture-Agent-RL-Training-Data
Agriculture Agent RL Training Data
A growing dataset of RL rollout trajectories for LLM agents on
natural/regenerative farming — the first RL/trajectory-shaped dataset in the
Copyleft Cultivars collection
(every prior dataset here is SFT/conversational Q&A). Agents call real tools
(primarily cultivars-mcp,
a plant-genomics MCP server) across 9 knowledge categories (plus a 10th,
organic_chemistry_soil_science, added 2026-08-11, and an 11th,
organic_chemistry_synthesis, added… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/Agriculture-Agent-RL-Training-Data.Llama-Nemotron-Post-Training-Dataset
Llama-Nemotron-Post-Training-Dataset-v1.1 Release
Update [4/8/2025]:
v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉
Data Overview
This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.Nemotron-Post-Training-Dataset-v2
Nemotron-Post-Training-Dataset-v2 Release
Data Overview
This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning.
NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.SII_self_evovling_02_training_datasetCryoLithe-training-datasetThe training Dataset for CryoLithe Models
The dataset contains selected tilt series, tilt angles, and corresponding cryo-CARE+IsoNet and Icecream reconstructions using odd/even pairs. For EMPIAR-11058
Icecream reconstructions were obtained by splitting across angles.
Whenever available, we also provide dose-fractionated tilt series.
Dataset format:
Files ending with '.rawtlt' or '.tlt' correspond to the tilt angles.
Files ending with '_corrected.mrc' correspond to cryo-CARE+IsoNet… See the full description on the dataset page: https://huggingface.co/datasets/sada-group/CryoLithe-training-dataset.AutoGaze-Training-DataGPT-Training-Datascbe-aethermoore-training-data
Status: canonical. Primary public training dataset for SCBE-AETHERMOORE and the most-used repo in this account. Other scbe-* dataset repos are experiment-specific slices.
SCBE-AETHERMOORE Training Dataset
Supervised fine-tuning (SFT) dataset for the SCBE-AETHERMOORE hyperbolic geometry AI safety and governance framework.
Overview
This dataset contains 10,978 training pairs spanning the full SCBE-AETHERMOORE system: 14-layer architecture knowledge, Six Sacred… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-aethermoore-training-data.lingshu_training_data_medical_domain
Website
🤖 7B Model
🤖 8B Model based on InternVL3
🤖 32B Model
MedEvalKit
Technical Report
Lingshu MCP
Lingshu Medical MLLM Training Data (Medical Domain)
This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included.
The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.sole_training_data
This is the training dataset for SOLE-R1-8B
SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning.
This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.DRT-SFT-8B-training-data
DRT-SFT-8B Training Data
Paper: DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal ReasoningCode: https://github.com/HIT-leaderone/DRT
This dataset contains the SFT training parquet shards used for DRT-SFT-8B.
Contents
20 parquet shards: Vision-R1_part_0.parquet ... Vision-R1_part_19.parquet
Total rows: 194,719
Columns: problem_id, content, role, image
Downloaded size: about 30.4 GiB
Notes
The parquet files are uploaded without… See the full description on the dataset page: https://huggingface.co/datasets/leaderonehit/DRT-SFT-8B-training-data.VideoChat3-Stage3-Training-Data
VideoChat3-Stage3-Training-Data
VideoChat3-Stage3-Training-Data contains the complete training data for the third stage of VideoChat3.
While preserving the model's basic multimodal capabilities, this data further improves long-video understanding and streaming video understanding.
📄 Paper · 🌐 Homepage · 💻 GitHub · 🤗 Paper Page
Data Sources
VideoChat3-Stage3-Training-Data includes our collected and open-sourced VideoChat3-Academic2M, VideoChat3-LV116K… See the full description on the dataset page: https://huggingface.co/datasets/lmwang/VideoChat3-Stage3-Training-Data.ntp_training_data_finewebedu_21bmistral_ntp_training_datahumanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.VideoChat3-Training-Data-Annotations
VideoChat3-Stage3-Training-Data
This repository includes all annotation files used across the four training stages of VideoChat3, from Stage 0 to Stage 3.
You can refer to the provided source-data links to download videos, images, and other multimedia data for training. In videochat3_data_annotations, we also provide a source field to indicate the source dataset for each entry.
To facilitate Stage 3 training reproduction using the high-quality open-source datasets we collected… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-Training-Data-Annotations.ws_ntp_training_dataembedding-training-data
Training Data for Text Embedding Models
[!NOTE]
This repository contains raw datasets, all of which have also been formatted for easy training in the Embedding Model Datasets collection. We recommend looking there first.
This repository contains training files to train text embedding models, e.g. using sentence-transformers.
Data Format
All files are in a jsonl.gz format: Each line contains a JSON-object that represent one training example.
The JSON objects can… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/embedding-training-data.Luciole-Training-Dataset
Data card for The Luciole Training Dataset
Table of Contents
Dataset Description
Curation Rationale
Web Data Opt-Outs
Personal and Sensitive Information (PII)
Bias, Risks, and Limitations
Recommendations
Sample Metadata
Downloading the Data
Sample Use in Python
Accessing the English Web Data and OpenMathInstruct-1
Details on Data Sources
Citation
Acknowledgements
Contact
Dataset Description
The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.gaussian_training_datasets
Gaussian Training Datasets (COLMAP) for msplat
COLMAP-format multi-view scenes for training 3D Gaussian Splatting models,
packaged for msplat — a Metal-native 3DGS
trainer for Apple Silicon. Also includes pre-trained .ply splats under
tested_outputs/.
All scenes are redistributed from third-party datasets. Full credit goes to
their original authors — see Licensing & credits and please
cite the original papers. This repo only repackages them in COLMAP layout for
convenience.… See the full description on the dataset page: https://huggingface.co/datasets/alexmkwizu/gaussian_training_datasets.
