CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MCG-NJU /VideoChat3-LV116k VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments. The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.textvideo-text-to-text1K<n<10K15 likes22k downloads2mo agoHugging Face02AaronZ345 /GTSinger GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks Yu Zhang*, Changhao Pan*, Wenxiang Guo*, Ruiqi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, LiChao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao | Zhejiang University Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/GTSinger.audiotext-to-audio10K<n<100K17 likes22k downloads1y agoHugging Face03Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M197 likes16k downloads28d agoHugging Face04Lijiaxin0111 /M3_VOS [CVPR 2025] M3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation If you like our project, please give us a star ⭐ on GitHub for the latest update. 💡 Description Venue: CVPR2025 Repository: 🛠️Tool, 🏠Page Paper: arxiv.org/html/2412.13803v2 Point of Contact: Jiaxin Li , Zixuan Chen 📁 Structure This dataset contains annotated videos and images for object segmentation tasks with phase transition information. The directory… See the full description on the dataset page: https://huggingface.co/datasets/Lijiaxin0111/M3_VOS.imagevideo-classificationn<1K1 likes16k downloads10mo agoHugging Face05BAAI /CCI3-HQgated Data Description To address the scarcity of high-quality safety datasets in the Chinese, we open-sourced the CCI (Chinese Corpora Internet) dataset on November 29, 2023. Building on this foundation, we continue to expand the data source, adopt stricter data cleaning methods, and complete the construction of the CCI 3.0 dataset. This dataset is composed of high-quality, reliable Internet data from trusted sources. And then with more stricter filtering, The CCI 3.0 HQ corpus… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CCI3-HQ.texttext-generation10M<n<100M67 likes9.7k downloads2y agoHugging Face06allenai /dolma3_mix-150B-1025 Dolma 3 Sample: 150B Mix Dataset Sources Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3 Source Type Tokens Documents Common Crawl Web pages 121B (76.9%) 84.5M olmOCR Science PDFs Academic documents 19.9B (12.6%) 2.25M Stack-Edu (Rebalanced) GitHub code 11.1B (7.06%) 14.3M arXiv Papers with LaTeX 1.29B (0.82%) 247K FineMath 3+ Math web pages 4.10B (2.60%) 2.57M Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.texttext-generation10M<n<100M10 likes8k downloads8mo agoHugging Face07nvidia /HelpSteer3 HelpSteer3 HelpSteer3 is an open-source dataset (CC-BY-4.0) that supports aligning models to become more helpful in responding to user prompts. HelpSteer3-Preference can be used to train Llama 3.3 Nemotron Super 49B v1 (for Generative RMs) and Llama 3.3 70B Instruct Models (for Bradley-Terry RMs) to produce Reward Models that score as high as 85.5% on RM-Bench and 78.6% on JudgeBench, which substantially surpass existing Reward Models on these benchmarks. HelpSteer3-Feedback and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer3.text100K<n<1M118 likes7.9k downloads10mo agoHugging Face08xxxspatialencoderwds3 /data_3 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds3/data_3.tabularobject-detectionn<1K0 likes6.8k downloads6d agoHugging Face09nvidia /Nemotron-Image-Training-v3 Nemotron Image Training v3 Versions Date Commit Changes 2026-04-28 HEAD Initial commit. Dataset Description Nemotron Image Training v3 is a collection of image-centric multimodal training data for vision–language models. Similar to Nemotron-VLM-Dataset v2, it was curated as a large-scale, multi-subdataset release where each subset ships a standardized conversation JSONL alongside a dataset card describing sources, licensing, and media layout.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Image-Training-v3.textvisual-question-answering1M<n<10M82 likes6.6k downloads5mo agoHugging Face10Carlosaug47 /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M4 likes6k downloads2mo agoHugging Face11nvidia /Nemotron-SFT-Instruction-Following-Chat-v3 Dataset Description: The Nemotron-Instruction-Following-Chat-v3 dataset is designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following. The chat subset uses human written prompts from sources like lmarena, lmsys, and wildchat as seed prompts. Responses are generated with GLM-5. Multiple responses are sampled from the model and the best response as judged by pairwise comparisons using Qwen3-Nemotron-235B-A22B-GenRM-2603… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.texttext-generation100K<n<1M20 likes5.6k downloads4mo agoHugging Face12artificialguybr /veo3-video-prompts Veo 3 Video Generation Dataset English | Português do Brasil English Summary A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant. Videos: 5,811 Input images: 1,354 Configurations: 6 Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.imagetext-to-video1K<n<10K0 likes5.3k downloads2mo agoHugging Face13kepton0117 /sn38-submissiontextn<1K0 likes5.2k downloads9d agoHugging Face14Stage-jh-monitor /qwen35-4b qwen35-4b Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38203125 Action score: 0.4375 Valid samples: 320/320 tabularn<1K0 likes5k downloads18d agoHugging Face15Stage-jh-monitor /appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1 appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3953125 Action score: 0.446875 Valid samples: 320/320 tabularn<1K0 likes5k downloads18d agoHugging Face16Stage-jh-monitor /total-300-random-jh-epoch4 total-300-random-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3890625 Action score: 0.440625 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face17dylanebert /3dgs3dn<1K12 likes4.9k downloads3y agoHugging Face18Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4 total-300-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4140625 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face19Stage-jh-monitor /total-300-lambda00-s_signal_type6-jh-epoch4 total-300-lambda00-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3875 Action score: 0.43125 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face20Stage-jh-monitor /total-300-lambda05-s_signal_type6-jh-epoch4 total-300-lambda05-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.35703125 Action score: 0.4375 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face21Stage-jh-monitor /total-300-lambda08-s_signal_type6-jh-epoch4 total-300-lambda08-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38046875 Action score: 0.4078125 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face22Stage-jh-monitor /total-300-lambda10-s_signal_type6-jh-epoch4 total-300-lambda10-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.41875 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face23Stage-jh-monitor /total-300noapp-lambda02-s_signal_type6-jh-epoch4 total-300noapp-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.409375 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face24ai4privacy /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.texttext-classification100K<n<1M116 likes4.9k downloads4mo agoHugging Face25Stage-jh-monitor /total-300app-lambda02-s_signal_type6-jh-epoch4 total-300app-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3625 Action score: 0.4015625 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face26Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-retry-epoch4 total-300-lambda02-s_signal_type6-jh-retry-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36953125 Action score: 0.3984375 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face27Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4-reeval2 total-300-lambda02-s_signal_type6-jh-epoch4-reeval2 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4125 Action score: 0.4265625 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face28Stage-jh-monitor /qwen35-4b-reeval3 qwen35-4b-reeval3 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3859375 Action score: 0.4125 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face29Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4-reeval1 total-300-lambda02-s_signal_type6-jh-epoch4-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38828125 Action score: 0.4234375 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face30Stage-jh-monitor /qwen35-4b-reeval2 qwen35-4b-reeval2 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.384375 Action score: 0.4265625 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.