CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenLLM-France /Lucie-Training-Dataset Lucie Training Dataset Card The Lucie Training Dataset is a curated collection of text data in English, French, German, Spanish and Italian culled from a variety of sources including: web data, video subtitles, academic papers, digital books, newspapers, and magazines, some of which were processed by Optical Character Recognition (OCR). It also contains samples of diverse programming languages. The Lucie Training Dataset was used to pretrain Lucie-7B, a foundation LLM with… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Lucie-Training-Dataset.texttext-generation10B<n<100B39 likes29k downloads1y agoHugging Face02common-pile /comma_v0.1_training_dataset Comma v0.1 dataset This repository contains the dataset used to train Comma v0.1-1T and Comma v0.1-2T. It is a slightly modified and consolidated version of the Common Pile v0.1 "filtered" data. If you are looknig for the raw Common Pile v0.1 data, please see this collection. You can learn more about Common Pile in our paper. Mixing rates and token counts The Comma v0.1 models were trained in two stages, a "main" stage and a "cooldown" stage. During each stage, we… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/comma_v0.1_training_dataset.text100M<n<1B45 likes25k downloads1y agoHugging Face03nvidia /Nemotron-Post-Training-Dataset-v1 Nemotron-Post-Training-Dataset-v1 Release This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5. Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.text10M<n<100M194 likes15k downloads1y agoHugging Face04EarthSpeciesProject /NatureLM-audio-training Dataset card for NatureLM-audio-training Overview NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording. For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.audioaudio-classification10M<n<100M18 likes12k downloads1y agoHugging Face05lightonai /multilingual-embeddings-pre-training-curated 📚 Collection | 📝 Multilingual Blog | 📝 English Blog Contrastive Multilingual Pre-Training 2.16B query–document pairs across eight languages, plus cross-lingual pairs mDenseOn | mLateOn | DenseOn | LateOn | PyLate | FastPlaid 🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.texttext-retrieval1B<n<10B5 likes9.3k downloads2mo agoHugging Face06AnchorSR /TrainingData_Stage3 AnchorSR Stage3 · metric-v1.0 直接选择 Small / Large 配置 训练题数 用途 small 1,000,000 先验证答案监督/先验恢复,按新版 Large 联合分布抽样 large 89,801,853 筛选后的完整训练集合,包含 Small 全部样本 from datasets import load_dataset data = load_dataset('AnchorSR/TrainingData_Stage3', 'small', # 或 large revision='metric-v1.0', streaming=True) 这是对 scaling-v1.0 的语义筛选与统一任务分类,不是增加新数据源。 Large 从 89,828,269 题保留 89,801,853 题,隔离 26,416 题。 旧标签 scaling-v1.0 / video-v1.0 / large-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.tabularvisual-question-answering100M<n<1B0 likes8.1k downloads1d agoHugging Face07bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K195 likes7.9k downloads2y agoHugging Face08DesmondYMTang2024 /Language-Grounded_Sparse_Encoder_Training Language-Grounded Sparse Encoder (LanSE) — Training Data This repository hosts the AI-generated images and human annotation datasets accompanying the paper: Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.textimage-classification100K<n<1M1 likes7.2k downloads17d agoHugging Face09nvidia /Nemotron-Image-Training-v3 Nemotron Image Training v3 Versions Date Commit Changes 2026-04-28 HEAD Initial commit. Dataset Description Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card describing sources, licensing, and media layout.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Image-Training-v3.textvisual-question-answering1M<n<10M82 likes7.1k downloads5mo agoHugging Face10licyk /image_training_set自用的训练集合集,用于 Stable Diffusion 模型微调。 该仓库仅用于存档,不提供任何技术支持。 imagen<1K2 likes6.9k downloads22d agoHugging Face11nvidia /Llama-Nemotron-Post-Training-Dataset Llama-Nemotron-Post-Training-Dataset-v1.1 Release Update [4/8/2025]: v1.1: We are releasing an additional 2.2M Math and 500K Code Reasoning Data in support of our release of Llama-3.1-Nemotron-Ultra-253B-v1. 🎉 Data Overview This dataset is a compilation of SFT and RL data that supports improvements of math, code, general reasoning, and instruction following capabilities of the original Llama instruct model, in support of NVIDIA’s release of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Llama-Nemotron-Post-Training-Dataset.text1M<n<10M709 likes6.7k downloads1y agoHugging Face12Amogh1221 /bellhart_trainingtext10K<n<100K0 likes5.7k downloads31m agoHugging Face13nvidia /Nemotron-Post-Training-Dataset-v2gated Nemotron-Post-Training-Dataset-v2 Release Data Overview This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning. NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.text1M<n<10M154 likes5.6k downloads1y agoHugging Face14Limelight /SII_self_evovling_02_training_datasettextn<1K0 likes5.1k downloads4mo agoHugging Face15Onkarn /GPT-Training-Datatext10M<n<100M0 likes4.3k downloads1y agoHugging Face16yukiZhang0527 /droid_s3r_training_release_v1tabularn<1K0 likes3.5k downloads1mo agoHugging Face17lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes3.2k downloads23d agoHugging Face18Open-Bee /Bee-Training-Data-Stage2 Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.imageimage-to-text10M<n<100M6 likes3k downloads7mo agoHugging Face19Philip-MIT /sole_training_data This is the training dataset for SOLE-R1-8B SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning. This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.image1M<n<10M0 likes2.8k downloads4mo agoHugging Face20leaderonehit /DRT-SFT-8B-training-data DRT-SFT-8B Training Data Paper: DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal ReasoningCode: https://github.com/HIT-leaderone/DRT This dataset contains the SFT training parquet shards used for DRT-SFT-8B. Contents 20 parquet shards: Vision-R1_part_0.parquet ... Vision-R1_part_19.parquet Total rows: 194,719 Columns: problem_id, content, role, image Downloaded size: about 30.4 GiB Notes The parquet files are uploaded without… See the full description on the dataset page: https://huggingface.co/datasets/leaderonehit/DRT-SFT-8B-training-data.textvisual-question-answering100K<n<1M0 likes2.8k downloads2d agoHugging Face21DynamicIntelligence /humanoid-robots-training-dataset Dynamic Intelligence — Humanoid Robot Training Dataset A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors. The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.tabularrobotics10K<n<100K0 likes2.7k downloads6mo agoHugging Face22OLAIR /OLA-Embed-Trainingtabular1B<n<10B0 likes2.1k downloads4mo agoHugging Face23ethanolivertroy /nist-cybersecurity-training NIST Cybersecurity Training Dataset v1.1 The largest open-source NIST cybersecurity training dataset for fine-tuning LLMs Version 1.1 Highlights What's New in v1.1: ✅ Added CSWP (Cybersecurity White Papers) series - 23 new documents ✅ Fixed 6,150 broken DOI links via format normalization ✅ Removed 202 malformed DOIs (double URL prefixes) ✅ Validated and fixed 124,946 total links ✅ Cataloged 72,698 broken links for future recovery ✅ 0 broken link markers remaining in… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-cybersecurity-training.texttext-generation100K<n<1M59 likes2k downloads11mo agoHugging Face24Dogacel /nemotron-post-training-v2-qwen-3.5-9b-regen Dataset Card for Nemotron Post Training v2 Qwen 3.5 9B Regen Regenerated responses from nvidia/Nemotron-Post-Training-Dataset-v2 dataset using Qwen3.5 9B model. Parameter Value Max Tokens 4096 Temperature 1.0 Top-k 20 Top-p 0.95 Repetition Penalty 1.5 Dataset consists only the english samples from the Nemotron Post Training Dataset. 85% of the chat prompts have reasoning enabled, every other category has reasoning disabled. Category Value math… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/nemotron-post-training-v2-qwen-3.5-9b-regen.texttext-generation100K<n<1M0 likes1.8k downloads5mo agoHugging Face25OpenLLM-France /Luciole-Training-Dataset Data card for The Luciole Training Dataset Table of Contents Dataset Description Curation Rationale Web Data Opt-Outs Personal and Sensitive Information (PII) Bias, Risks, and Limitations Recommendations Sample Metadata Downloading the Data Sample Use in Python Accessing the English Web Data and OpenMathInstruct-1 Details on Data Sources Citation Acknowledgements Contact Dataset Description The Luciole Training Dataset is a curated collection of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Training-Dataset.texttext-generation1B<n<10B15 likes1.8k downloads2mo agoHugging Face26sy1998 /Video_XL_Trainingtext7 likes1.8k downloads2y agoHugging Face27nakas /mtnwx-trainingtabular1B<n<10B0 likes1.7k downloads1mo agoHugging Face28bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.5k downloads2y agoHugging Face29ameet /deepsql_training SynSQL Data Processing A Python tool for processing the SynSQL-2.5M dataset into optimized Parquet format for machine learning workflows. The dataset is split into batches of 30K entries with chain of thought(COT) reasoning and the answer. This can then be preprocessed and used for training any reasoning model. Dataset Acknowledgment This project processes data from the SynSQL-2.5M dataset by seeklhy, which is licensed under Apache 2.0. We acknowledge and thank the… See the full description on the dataset page: https://huggingface.co/datasets/ameet/deepsql_training.texttext-generation1M<n<10M0 likes1.4k downloads1y agoHugging Face30Post-training-Data-Flywheel /gorilla-openfunctions-v1text10K<n<100K0 likes1.4k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.