CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Inception3D /GenFusion_Training_Datavideo10K<n<100K1 likes34k downloads1y agoHugging Face02OpenGVLab /VideoChat-Flash-Training-Data 🦜 VideoChat-Flash-Training-Data This repos contains all annotaions and most videos for training VideoChat-Flash. 📕 How to use the LongVid data? For video_dir like longvid_subset/coin_grounding_10k_zip, you need to concat this dir to a zip file as follows: cat ego4dhcap_eventunderstanding_2k_zip/* > ego4dhcap_eventunderstanding_2k.zip ✏️ Citation @article{li2024videochatflash, title={VideoChat-Flash: Hierarchical Compression for Long-Context… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat-Flash-Training-Data.video-text-to-text10K<n<100K16 likes29k downloads1y agoHugging Face03OpenLLM-France /Lucie-Training-Dataset Lucie Training Dataset Card The Lucie Training Dataset is a curated collection of text data in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers, digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages. The Lucie Training Dataset was used to pretrain Lucie-7B, a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.texttext-generation10B<n<100B39 likes28k downloads1y agoHugging Face04common-pile /comma_v0.1_training_dataset Comma v0.1 dataset This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T. It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data. If you are looknig for the raw Common Pile v0.1 data, please see this collection. You can learn more about Common Pile in our paper. Mixing rates and token counts The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage. During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.text100M<n<1B45 likes23k downloads1y agoHugging Face05nvidia /Nemotron-Post-Training-Dataset-v1 Nemotron-Post-Training-Dataset-v1 Release This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5. Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.text10M<n<100M194 likes16k downloads1y agoHugging Face06jankin123 /4DThinker-Training-Data 4DThinker Training Data This repository contains the training data for 4DThinker, a framework that enables VLMs to "think with 4D" through dynamic latent mental imagery, built upon SpatialVID and DSR_Suite-Data. Data Structure data/ ├── dift_data.jsonl # DIFT training data (~38K samples) ├── 4drl_data_filtered.jsonl # 4DRL training data (~37K samples) └── processed_data/ # Video frames & mask overlays ├── <video_id>/ │ ├── frames/… See the full description on the dataset page: https://huggingface.co/datasets/jankin123/4DThinker-Training-Data.1 likes12k downloads4mo agoHugging Face07bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K194 likes7.4k downloads2y agoHugging Face08AnchorSR /TrainingData_Stage3 Anchor Spatial Reasoning — Stage 3 · video-v1.0 先选择 Small 或 Large 版本 训练题数 用途 HF config Small 1,000,000 验证 Stage3 答案监督能否恢复空间度量先验 small Large 89,828,269 全量合格训练题,包含全部 Small large 两版共用原 validation / test,各 50,000 条;新增独立 video_validation 1,585 条、video_test 1,725 条。 不要拼接 Small 与 Large,也不要把原始池 media_complete 当作隔离后的训练集。 旧版本固定在 large-v1.0;原始池筛选统计及旧版说明保留。 本版 Small 替换 100,000 条 HiSpatial 距离题,总量不变;Large 保留全部原训练题, 新增 141,268 条视频题。不改变旧主评测题内容。 from datasets… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.tabularvisual-question-answering100M<n<1B0 likes7.1k downloads36m agoHugging Face09anuj-inavlabs /Thinkspark-v2-270m-training-data ThinkSpark-v2-350M — training data Full-duplex floor-controller (Section 8) training corpus: playable audio + text, paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain, gender, prosody, agent text) and Soniox character-level timestamps. Dataset Viewer Default split is parquet with a real Audio feature — a player renders inline next to the text in the Hub UI: column type description audio Audio playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.text-to-speech1K<n<10K0 likes6.9k downloads20d agoHugging Face10cocool /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/cocool/lingshu_training_data_medical_domain.image-to-text1M<n<10M0 likes6.5k downloads20d agoHugging Face11CopyleftCultivars /Agriculture-Agent-RL-Training-Data Agriculture Agent RL Training Data A growing dataset of RL rollout trajectories for LLM agents on natural/regenerative farming — the first RL/trajectory-shaped dataset in the Copyleft Cultivars collection (every prior dataset here is SFT/conversational Q&A). Agents call real tools (primarily cultivars-mcp, a plant-genomics MCP server) across 9 knowledge categories (plus a 10th, organic_chemistry_soil_science, added 2026-08-11, and an 11th, organic_chemistry_synthesis, added… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/Agriculture-Agent-RL-Training-Data.text-generation1 likes6.2k downloads24d agoHugging Face12nvidia /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M709 likes6.2k downloads1y agoHugging Face13nvidia /Nemotron-Post-Training-Dataset-v2gated Nemotron-Post-Training-Dataset-v2 Release Data Overview This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning. NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.text1M<n<10M153 likes5.8k downloads1y agoHugging Face14Limelight /SII_self_evovling_02_training_datasettextn<1K0 likes5.4k downloads4mo agoHugging Face15sada-group /CryoLithe-training-datasetThe training Dataset for CryoLithe Models The dataset contains selected tilt series, tilt angles, and corresponding cryo-CARE+IsoNet and Icecream reconstructions using odd/even pairs. For EMPIAR-11058 Icecream reconstructions were obtained by splitting across angles. Whenever available, we also provide dose-fractionated tilt series. Dataset format: Files ending with '.rawtlt' or '.tlt' correspond to the tilt angles. Files ending with '_corrected.mrc' correspond to cryo-CARE+IsoNet… See the full description on the dataset page: https://huggingface.co/datasets/sada-group/CryoLithe-training-dataset.fill-maskn<1K3 likes4.9k downloads3mo agoHugging Face16bfshi /AutoGaze-Training-Data0 likes4.4k downloads7mo agoHugging Face17Onkarn /GPT-Training-Datatext10M<n<100M0 likes4.3k downloads1y agoHugging Face18issdandavis /scbe-aethermoore-training-data Status: canonical. Primary public training dataset for SCBE-AETHERMOORE and the most-used repo in this account. Other scbe-* dataset repos are experiment-specific slices. SCBE-AETHERMOORE Training Dataset Supervised fine-tuning (SFT) dataset for the SCBE-AETHERMOORE hyperbolic geometry AI safety and governance framework. Overview This dataset contains 10,978 training pairs spanning the full SCBE-AETHERMOORE system: 14-layer architecture knowledge, Six Sacred… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-aethermoore-training-data.text-generation10K<n<100K2 likes4.2k downloads1d agoHugging Face19lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B7 likes2.9k downloads21d agoHugging Face20Philip-MIT /sole_training_data This is the training dataset for SOLE-R1-8B SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning. This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.image1M<n<10M0 likes2.9k downloads4mo agoHugging Face21leaderonehit /DRT-SFT-8B-training-data DRT-SFT-8B Training Data Paper: DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal ReasoningCode: https://github.com/HIT-leaderone/DRT This dataset contains the SFT training parquet shards used for DRT-SFT-8B. Contents 20 parquet shards: Vision-R1_part_0.parquet ... Vision-R1_part_19.parquet Total rows: 194,719 Columns: problem_id, content, role, image Downloaded size: about 30.4 GiB Notes The parquet files are uploaded without… See the full description on the dataset page: https://huggingface.co/datasets/leaderonehit/DRT-SFT-8B-training-data.textvisual-question-answering100K<n<1M0 likes2.8k downloads15h agoHugging Face22lmwang /VideoChat3-Stage3-Training-Data VideoChat3-Stage3-Training-Data VideoChat3-Stage3-Training-Data contains the complete training data for the third stage of VideoChat3. While preserving the model's basic multimodal capabilities, this data further improves long-video understanding and streaming video understanding. 📄 Paper · 🌐 Homepage · 💻 GitHub · 🤗 Paper Page Data Sources VideoChat3-Stage3-Training-Data includes our collected and open-sourced VideoChat3-Academic2M, VideoChat3-LV116K… See the full description on the dataset page: https://huggingface.co/datasets/lmwang/VideoChat3-Stage3-Training-Data.video-text-to-text6 likes2.6k downloads2mo agoHugging Face23l2t-project /ntp_training_data_finewebedu_21b0 likes2.6k downloads1y agoHugging Face24klein9692 /mistral_ntp_training_data0 likes2.4k downloads6mo agoHugging Face25DynamicIntelligence /humanoid-robots-training-dataset Dynamic Intelligence — Humanoid Robot Training Dataset A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors. The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.tabularrobotics10K<n<100K0 likes2.4k downloads6mo agoHugging Face26MCG-NJU /VideoChat3-Training-Data-Annotations VideoChat3-Stage3-Training-Data This repository includes all annotation files used across the four training stages of VideoChat3, from Stage 0 to Stage 3. You can refer to the provided source-data links to download videos, images, and other multimedia data for training. In videochat3_data_annotations, we also provide a source field to indicate the source dataset for each entry. To facilitate Stage 3 training reproduction using the high-quality open-source datasets we collected… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-Training-Data-Annotations.video-text-to-text2 likes2.3k downloads2mo agoHugging Face27klein9692 /ws_ntp_training_data0 likes2.2k downloads4mo agoHugging Face28sentence-transformers /embedding-training-data Training Data for Text Embedding Models [!NOTE] This repository contains raw datasets, all of which have also been formatted for easy training in the Embedding Model Datasets collection. We recommend looking there first. This repository contains training files to train text embedding models, e.g. using sentence-transformers. Data Format All files are in a jsonl.gz format: Each line contains a JSON-object that represent one training example. The JSON objects can… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/embedding-training-data.feature-extraction144 likes2.1k downloads27d agoHugging Face29OpenLLM-France /Luciole-Training-Dataset Data card for The Luciole Training Dataset Table of Contents Dataset Description Curation Rationale Web Data Opt-Outs Personal and Sensitive Information (PII) Bias, Risks, and Limitations Recommendations Sample Metadata Downloading the Data Sample Use in Python Accessing the English Web Data and OpenMathInstruct-1 Details on Data Sources Citation Acknowledgements Contact Dataset Description The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.texttext-generation1B<n<10B15 likes1.9k downloads2mo agoHugging Face30alexmkwizu /gaussian_training_datasets Gaussian Training Datasets (COLMAP) for msplat COLMAP-format multi-view scenes for training 3D Gaussian Splatting models, packaged for msplat — a Metal-native 3DGS trainer for Apple Silicon. Also includes pre-trained .ply splats under tested_outputs/. All scenes are redistributed from third-party datasets. Full credit goes to their original authors — see Licensing & credits and please cite the original papers. This repo only repackages them in COLMAP layout for convenience.… See the full description on the dataset page: https://huggingface.co/datasets/alexmkwizu/gaussian_training_datasets.imageimage-to-3dn<1K0 likes1.7k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.