CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Vchitect /Vchitect_T2V_DataVerse Vchitect-T2V-Dataverse Vchitect Team1  1Shanghai Artificial Intelligence Laboratory  Paper | Project Page | Data Overview The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models. It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.texttext-to-video1M<n<10M11 likes54k downloads1y agoHugging Face02TencentARC /TimeLens-100K TimeLens-100K 📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data ✨ Dataset Description TimeLens-100K is a large-scale, diverse, and high-quality training dataset for video temporal grounding. It was proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs and used for training TimeLens models. The annotation process was conducted using an automated pipeline powered by Gemini-2.5-Pro. 📊 Dataset Statistics Total Videos:… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-100K.textvideo-text-to-text10K<n<100K7 likes24k downloads9mo agoHugging Face03ma-xu /fine-t2i Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning [arxiv] by Xu Ma, Yitian Zhang, Qihua Dong, Yun Fu Northeastern Univeristy Please see our [Dataset Explore] to view detailed samples (loading is slow, be patient). 🆕 What's New [2026.02.20]: Fine-T2I reaches the #1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️ [2026.02.16]: Fine-T2I tops the Hugging Face Datasets Trending list, reaching the #2 spot and #1… See the full description on the dataset page: https://huggingface.co/datasets/ma-xu/fine-t2i.imageimage-to-text100K<n<1M120 likes20k downloads7mo agoHugging Face04lioooox /T2I-CoReBench-Images T2I-CoReBench-Images 📖 Overview T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities. This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench-Images.imagetext-to-image10K<n<100K5 likes10k downloads7mo agoHugging Face05laion /laions_got_talent LAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel. Dataset Composition The dataset includes: Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.audio100K<n<1M41 likes9.7k downloads2y agoHugging Face06AILab-CVC /obelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data text10M<n<100M1 likes9.4k downloads3y agoHugging Face07facebook /hand_tracking_challenge_umetrackCheck out the Multiview Egocentric Hand Tracking Challenge 2024!! To use this dataset, check out the hand_tracking_toolkit image100K<n<1M1 likes6.9k downloads2y agoHugging Face08Amshaker /Mobile-O-Post-Train Mobile-O Post-Training Data Unified Multimodal Post-Training · ~105K Quadruplet Samples 📌 Overview This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation. The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples. 📊 Dataset Format Each sample is a quadruplet consisting of:… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Post-Train.imagetext-to-image1K<n<10K13 likes6.7k downloads7mo agoHugging Face09timm /imagenet-22k-wdsgated Dataset Summary This is a copy of the full ImageNet dataset consisting of all of the original 21841 clases. It also contains labels in a separate field for the '12k' subset described at at (https://github.com/rwightman/imagenet-12k, https://huggingface.co/datasets/timm/imagenet-12k-wds) This dataset is from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-22k-wds.imageimage-classification100K<n<1M14 likes6.4k downloads3y agoHugging Face10artefactory /Argimi-Ardian-Finance-10k-text The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.texttext-retrieval1M<n<10M19 likes5.8k downloads7mo agoHugging Face11TIGER-Lab /arxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot. image10M<n<100M5 likes5.7k downloads1y agoHugging Face12XRXRX /X-Voice-Dataset-Train X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling. Also the train set of X-Voice Model. Core Statistics Total Speech Duration: 420K hours 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.audiotext-to-speech10M<n<100M11 likes4.7k downloads5mo agoHugging Face13imageomics /TreeOfLife-10M Dataset Card for TreeOfLife-10M Dataset Summary With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M.documentimage-classification1M<n<10M50 likes4.2k downloads8mo agoHugging Face14timm /imagenet-1k-wdsgated Dataset Summary ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated. 💡… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-1k-wds.imageimage-classification10K<n<100K35 likes4.1k downloads3y agoHugging Face15timm /imagenet-12k-wdsgated Dataset Summary This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm. The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-12k-wds.imageimage-classification100K<n<1M10 likes3.7k downloads3y agoHugging Face16tom-jerry-123 /Physical-AI-AV-US PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 2,789,773 samples from 150 000 driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the United States. Format WebDataset — 100 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.png Front-facing wide-angle camera frame (640 × 360 px) {key}.json Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.imagerobotics1M<n<10M0 likes3.5k downloads6mo agoHugging Face17laion /laions_got_talent_rawaudio10K<n<100K7 likes3.4k downloads2y agoHugging Face18TTS-AGI /emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset. https://huggingface.co/datasets/amphion/Emilia-Dataset audiotext-to-speech10M<n<100M5 likes3.1k downloads2y agoHugging Face19TalTechNLP /voxlingua107_wds VoxLingua107 VoxLingua107 is a speech dataset for training spoken language identification models. The dataset consists of short speech segments automatically extracted from YouTube videos and labeled according the language of the video title and description, with some post-processing steps to filter out false positives. VoxLingua107 contains data for 107 languages. The total amount of speech in the training set is 6628 hours. The average amount of data per language is 62 hours.… See the full description on the dataset page: https://huggingface.co/datasets/TalTechNLP/voxlingua107_wds.audio1M<n<10M4 likes3k downloads1y agoHugging Face20KyujinL /CALVIN_ABC_tartext1M<n<10M1 likes2.7k downloads6mo agoHugging Face21Amshaker /Mobile-O-Pre-Train Mobile-O Pre-Training Data Cross-Modal Alignment · 9M Text-Image Pairs 📌 Overview This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation. The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs. 📊 Dataset Composition Source Samples Description… See the full description on the dataset page: https://huggingface.co/datasets/Amshaker/Mobile-O-Pre-Train.imagetext-to-image10M<n<100M12 likes2.6k downloads7mo agoHugging Face22Yuxuan0701 /tvqa-framesimage1M<n<10M0 likes2.2k downloads7mo agoHugging Face23TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes2.1k downloads6mo agoHugging Face24TencentARC /TimeLens-Bench TimeLens-Bench 📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data | 🏆 TimeLens-Bench Leaderboard ✨ Dataset Description TimeLens-Bench is a comprehensive, high-quality evaluation benchmark for video temporal grounding, proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs. During our annotation process, we identified critical quality issues within existing datasets and performed extensive manual corrections. We observed a… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-Bench.textvideo-text-to-text1K<n<10K8 likes2.1k downloads9mo agoHugging Face25kohei0209 /mls_hq_urgent_track1audio100K<n<1M0 likes1.9k downloads2y agoHugging Face26THUdyh /Ola-DataThis repository contains the data presented in Ola: Pushing the Frontiers of Omni-Modal Language Model with Progressive Modality Alignment. Code: https://github.com/Ola-Omni/Ola audioany-to-any100K<n<1M9 likes1.7k downloads2y agoHugging Face27semi-truths /Semi-Truths Semi Truths Dataset: A Large-Scale Dataset for Testing Robustness of AI-Generated Image Detectors (NeurIPS 2024 Track Datasets & Benchmarks Track) Recent efforts have developed AI-generated image detectors claiming robustness against various augmentations, but their effectiveness remains unclear. Can these systems detect varying degrees of augmentation? To address these questions, we introduce Semi-Truths, featuring 27, 600 real images, 223, 400 masks, and 1, 472, 700… See the full description on the dataset page: https://huggingface.co/datasets/semi-truths/Semi-Truths.imageimage-classification1M<n<10M9 likes1.6k downloads2y agoHugging Face28ProKSMT /bdnew_tar_s2d0image1M<n<10M0 likes1.5k downloads6mo agoHugging Face29Ajax102 /TaisuTaisu Dataset for https://github.com/ksOAn6g5/TaiSu The total size of the data is about 7.9T, we split the original data into tar files no larger than 10GB. imageimage-to-text100M<n<1B3 likes1.4k downloads1y agoHugging Face30turing-motors /STRIDE-QA-Dataset STRIDE-QA Dataset 📦 Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks. Category Description Object-centric Spatial QA Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.imagevisual-question-answering100K<n<1M9 likes1.4k downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.