CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KyujinL /CALVIN_ABC_tartext1M<n<10M1 likes2.7k downloads6mo agoHugging Face02KangLiao /Puffin-4M Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation    📖 Project Page  |    🖥️ GitHub    |   🤗 Hugging Face   |    📑 Paper    Dataset Details Datasets and benchmarks that span vision, language, and camera modalities remain scarce in the domain of spatial multimodal intelligence. To address this gap, we introduce Puffin-4M, a large-scale, high-quality dataset comprising 4 million vision-language-camera… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Puffin-4M.imagetext-to-image1B<n<10B35 likes2.3k downloads9mo agoHugging Face03kohei0209 /mls_hq_urgent_track1audio100K<n<1M0 likes1.9k downloads2y agoHugging Face04kensho /PubTables-v2 PubTables-v2 PubTables-v2 is a new large-scale dataset for full-page and multi-page table extraction. Official dataset evaluation scripts and leaderboard coming soon! In the meantime, you can create your own evaluation using GriTS with our open-source package, pip install grits-metric. Report any issues here: https://github.com/kensho-technologies/grits. See also: Hugging Face Paper Page News 2026 Apr 15: Code for the GriTS metric released… See the full description on the dataset page: https://huggingface.co/datasets/kensho/PubTables-v2.imageimage-to-text1M<n<10M25 likes1.7k downloads5mo agoHugging Face05krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face06krishnakalyan3 /emo_parleraudio1M<n<10M2 likes1.4k downloads2y agoHugging Face07krishnakalyan3 /emo_webdsaudio10K<n<100K5 likes1.3k downloads2y agoHugging Face08trevin-wadu /npm3d-kitti-carlatext10K<n<100K0 likes1.3k downloads3y agoHugging Face09khanhvinh9 /imagenet-cimage1M<n<10M0 likes847 downloads9mo agoHugging Face10KAIST-SmartDesignLab /DeepJEB DeepJEB: 3D Deep Learning-Based Synthetic Jet Engine Bracket Dataset This is the Hugging Face distribution of DeepJEB, a synthetic 3D jet engine bracket dataset of 2,138 designs with paired geometry and finite-element analysis (FEA) results, generated from the SimJEB seed set via a DeepSDF-based generative model and an automated simulation pipeline. This repository mirrors the official DeepJEB v1.0 release. To keep the ~68k component files practical to host, each component… See the full description on the dataset page: https://huggingface.co/datasets/KAIST-SmartDesignLab/DeepJEB.text10K<n<100K4 likes795 downloads4mo agoHugging Face11SanRiiiii /kiwi_subimage100K<n<1M0 likes757 downloads4mo agoHugging Face12nhatchung /Virtual_KITTI2image10K<n<100K0 likes729 downloads9mo agoHugging Face13KlingTeam /VIVID-10M VIVID-10M [project page] | [Paper] | [arXiv] VIVID-10M is the first large-scale hybrid image-video local editing dataset aimed at reducing data construction and model training costs, comprising 9.7M samples that encompass a wide range of video editing tasks. Data Index The data index is located at four .csv files: vivid-image-change.csv vivid-image-remove.csv vivid-video-change.csv vivid-video-remove.csv VIVID-Video splits contains the columns: local_caption, #… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/VIVID-10M.imagetext-to-video10M<n<100M18 likes692 downloads10mo agoHugging Face14koorye /ImageNet-Renditionimage10K<n<100K0 likes629 downloads10mo agoHugging Face15KevinMathew /stereo4d-lefteye-perspective Dataset Summary This dataset contains the left-eye rectified perspective views from the Stereo4D dataset (Paper). Each video is generated using the rectify.py script, which processes VR180 stereo videos to produce 512×512 video clips with a 60° field of view perspective camera. This dataset is intended to be used alongside the Stereo4D dataset annotations which can be found here. This dataset is provided as-is for non-commercial research purposes only. Download git clone… See the full description on the dataset page: https://huggingface.co/datasets/KevinMathew/stereo4d-lefteye-perspective.text100K<n<1M12 likes586 downloads1y agoHugging Face16kaiyuyue /llava-1.5-665k-instructionsThis dataset repository, LLaVA-1.5-665K-Instructions, is notably utilized in the paper Zero-Shot Vision Encoder Grafting via LLM Surrogates. The official code repository for the paper can be found here: https://github.com/kaiyuyue/zero LLaVA-1.5-665K-Instructions This dataset repo contains the entire LLaVA-1.5-665K-Instructions in one place, including images and text sequences. The images are in train_split/*.tars and the text sequences are in jsons: llava_v1_5_mix665k.json is the… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/llava-1.5-665k-instructions.imagevisual-question-answering100K<n<1M10 likes581 downloads1y agoHugging Face17kxic /Objaverse_zero123_wdsimagen<1K1 likes550 downloads2y agoHugging Face18k2-fsa /OpenDialog OpenDialog OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching. Paper: https://arxiv.org/abs/2507.09318 GitHub: https://github.com/k2-fsa/ZipVoice Project Page: https://zipvoice-dialog.github.io OpenDialog is the first large-scale (6.8k hours) open-source spoken dialogue dataset derived from in-the-wild speech data. It consists of: English data: 5074 hours Chinese data: 1759… See the full description on the dataset page: https://huggingface.co/datasets/k2-fsa/OpenDialog.audiotext-to-speech100K<n<1M24 likes550 downloads5mo agoHugging Face19koke /AllTheBacteria-FCGR-7mertext1M<n<10M0 likes451 downloads1y agoHugging Face20Kwai-Keye /VideoTemp-o3 VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos Illustration of the agentic pipeline in VideoTemp-o3. Given a video QA pair, the model performs on-demand temporal grounding to locate the most relevant segment, then refines it iteratively. Finally, it produces a reliable answer grounded in the pertinent visual evidence. Data Source The question and answer pairs used for training VideoTemp-o3 are sourced from… See the full description on the dataset page: https://huggingface.co/datasets/Kwai-Keye/VideoTemp-o3.text10K<n<100K1 likes426 downloads4mo agoHugging Face21KBlueLeaf /laion-coco-13m-taraudio10M<n<100M1 likes342 downloads1y agoHugging Face22undefined443 /coco-karpathy-wds COCO-2014 WebDataset Format (Karpathy Splits) This dataset contains the COCO-2014 images and captions converted to WebDataset (WDS) format, using the Karpathy & Li (2015) dataset split for image captioning tasks. Overview Total Samples: 123,287 images with 5 reference captions each Total Size: ~19 GB Format: WebDataset (.tar shards) Shard Size: 1,000 samples per tar file License: CC-BY 4.0 Language: English Structure COCO-2014-WDS/ ├── train/ (113… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/coco-karpathy-wds.imageimage-to-text10K<n<100K0 likes293 downloads5mo agoHugging Face23koorye /ImageNet-V2image10K<n<100K0 likes288 downloads10mo agoHugging Face24AbdullahRian /Korean.OCR.Img.text.pairimage100K<n<1M1 likes262 downloads1y agoHugging Face25kohei0209 /mls_hqaudio10M<n<100M0 likes221 downloads2y agoHugging Face26kwanY /stylebench-sSIGGRAPH 2026 / ACM TOG Journal Track image100K<n<1M1 likes215 downloads5mo agoHugging Face27kausable /CausalDynamics CausalDynamics: A large-scale benchmark for structural discovery of dynamical causal models NeurIPS 2025 A comprehensive benchmark framework designed to rigorously evaluate state-of-the-art causal discovery algorithms for dynamical systems. Key Features 1️⃣ Large-Scale Benchmark. Systematically evaluate state-of-the-art causal discovery algorithms on thousands of graph challenges with increasing difficulty. 2️⃣ Customizable Data Generation. Scalable… See the full description on the dataset page: https://huggingface.co/datasets/kausable/CausalDynamics.text10K<n<100K5 likes196 downloads1y agoHugging Face28khanhvinh9 /imagenetimage1M<n<10M0 likes184 downloads9mo agoHugging Face29krishnakalyan3 /emo_speech_filtered_v12 second filtered emotional speech in webdataset format https://huggingface.co/datasets/EQ4You/Emotional_Speech audio10K<n<100K0 likes181 downloads2y agoHugging Face30KyleBae1017 /stormer-40yrstrain: 1979~2018 val: 2019 test: 2020 HF_HUB_ENABLE_HF_TRANSFER=1 hf download —repo-type dataset —local-dir /workspace/stormer/ KyleBae1017/stormer-40yrs cat wb2_h5df.tar.part-* | tar -xvf - image10K<n<100K0 likes159 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.