CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KangLiao /Puffin-4M Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation    📖 Project Page  |    🖥️ GitHub    |   🤗 Hugging Face   |    📑 Paper    Dataset Details Datasets and benchmarks that span vision, language, and camera modalities remain scarce in the domain of spatial multimodal intelligence. To address this gap, we introduce Puffin-4M, a large-scale, high-quality dataset comprising 4 million vision-language-camera… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Puffin-4M.imagetext-to-image1B<n<10B35 likes2.3k downloads9mo agoHugging Face02kensho /PubTables-v2 PubTables-v2 PubTables-v2 is a new large-scale dataset for full-page and multi-page table extraction. Official dataset evaluation scripts and leaderboard coming soon! In the meantime, you can create your own evaluation using GriTS with our open-source package, pip install grits-metric. Report any issues here: https://github.com/kensho-technologies/grits. See also: Hugging Face Paper Page News 2026 Apr 15: Code for the GriTS metric released… See the full description on the dataset page: https://huggingface.co/datasets/kensho/PubTables-v2.imageimage-to-text1M<n<10M25 likes1.7k downloads5mo agoHugging Face03khanhvinh9 /imagenet-cimage1M<n<10M0 likes847 downloads9mo agoHugging Face04SanRiiiii /kiwi_subimage100K<n<1M0 likes757 downloads5mo agoHugging Face05nhatchung /Virtual_KITTI2image10K<n<100K0 likes729 downloads9mo agoHugging Face06KlingTeam /VIVID-10M VIVID-10M [project page] | [Paper] | [arXiv] VIVID-10M is the first large-scale hybrid image-video local editing dataset aimed at reducing data construction and model training costs, comprising 9.7M samples that encompass a wide range of video editing tasks. Data Index The data index is located at four .csv files: vivid-image-change.csv vivid-image-remove.csv vivid-video-change.csv vivid-video-remove.csv VIVID-Video splits contains the columns: local_caption, #… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/VIVID-10M.imagetext-to-video10M<n<100M18 likes692 downloads10mo agoHugging Face07koorye /ImageNet-Renditionimage10K<n<100K0 likes629 downloads10mo agoHugging Face08kaiyuyue /llava-1.5-665k-instructionsThis dataset repository, LLaVA-1.5-665K-Instructions, is notably utilized in the paper Zero-Shot Vision Encoder Grafting via LLM Surrogates. The official code repository for the paper can be found here: https://github.com/kaiyuyue/zero LLaVA-1.5-665K-Instructions This dataset repo contains the entire LLaVA-1.5-665K-Instructions in one place, including images and text sequences. The images are in train_split/*.tars and the text sequences are in jsons: llava_v1_5_mix665k.json is the… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/llava-1.5-665k-instructions.imagevisual-question-answering100K<n<1M10 likes581 downloads1y agoHugging Face09kxic /Objaverse_zero123_wdsimagen<1K1 likes550 downloads2y agoHugging Face10undefined443 /coco-karpathy-wds COCO-2014 WebDataset Format (Karpathy Splits) This dataset contains the COCO-2014 images and captions converted to WebDataset (WDS) format, using the Karpathy & Li (2015) dataset split for image captioning tasks. Overview Total Samples: 123,287 images with 5 reference captions each Total Size: ~19 GB Format: WebDataset (.tar shards) Shard Size: 1,000 samples per tar file License: CC-BY 4.0 Language: English Structure COCO-2014-WDS/ ├── train/ (113… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/coco-karpathy-wds.imageimage-to-text10K<n<100K0 likes293 downloads5mo agoHugging Face11koorye /ImageNet-V2image10K<n<100K0 likes288 downloads10mo agoHugging Face12AbdullahRian /Korean.OCR.Img.text.pairimage100K<n<1M1 likes262 downloads1y agoHugging Face13kwanY /stylebench-sSIGGRAPH 2026 / ACM TOG Journal Track image100K<n<1M1 likes215 downloads5mo agoHugging Face14khanhvinh9 /imagenetimage1M<n<10M0 likes184 downloads9mo agoHugging Face15KyleBae1017 /stormer-40yrstrain: 1979~2018 val: 2019 test: 2020 HF_HUB_ENABLE_HF_TRANSFER=1 hf download —repo-type dataset —local-dir /workspace/stormer/ KyleBae1017/stormer-40yrs cat wb2_h5df.tar.part-* | tar -xvf - image10K<n<100K0 likes159 downloads9mo agoHugging Face16KBlueLeaf /danbooru2023-webp-4Mpixelgated Danbooru 2023 webp: A space-efficient version of Danbooru 2023 This dataset is a resized/re-encoded version of danbooru2023. Which removed the non-image/truncated files and resize all of them into smaller size. This dataset already be updated to latest_id = 7,832,883. Thx to DeepGHS! Notice: content of updates folder and deepghs/danbooru_newest-webp-4Mpixel have been merged to 2000~2999.tar, You can ignore all the content in updates folder safely! Details This… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-webp-4Mpixel.imageimage-classification100M<n<1B83 likes154 downloads2y agoHugging Face17oceanfish /kittiimage10K<n<100K0 likes143 downloads2y agoHugging Face18anthemlin /kitti-cThis dataset is created by MonoTTA: Fully Test-Time Adaptation for Monocular 3D Object Detection, based on KITTI. You can check this link for more details: https://arxiv.org/abs/2405.19682v1 And access the code: https://github.com/Hongbin98/MonoTTA Please double-check the demands of KITTI when you try to download this dataset and obey their rules. image10K<n<100K5 likes100 downloads1y agoHugging Face19KMasaki /cc12m-sam2-parse-treeimage10M<n<100M0 likes87 downloads6mo agoHugging Face20kronecker0122 /kitti-depth-completion KITTI Depth Completion This repository contains a tar.zst archive of the kitti_depth_completion directory, split into 5 GiB parts. Archive parts: 4 Total size: 21,374,819,904 bytes Source file count: 193,568 Restore cat kitti_depth_completion.tar.zst.part-* | tar --zstd -xf - image100K<n<1M0 likes83 downloads14d agoHugging Face21kaiyuyue /sphere-encoder-fid-artifacts Sphere Encoder FID Evaluation Artifacts This repository contains the evaluation artifacts for the paper Image Generation with a Sphere Encoder. Project Page | GitHub Repository These artifacts include data statistic files (fid_stats) and reference images (fid_refs) used to calculate Fréchet Inception Distance (FID) for generative models across several datasets, including CIFAR-10, ImageNet, Animal Faces, and Oxford Flowers. Workspace Setup Download the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/sphere-encoder-fid-artifacts.imageother10K<n<100K1 likes82 downloads7mo agoHugging Face22koorye /Describable-Textures-Datasetimage1K<n<10K0 likes79 downloads10mo agoHugging Face23Ghaser /Wikipedia-Knowledge-2M 📃 Paper | 🤗 Hugging Face | ⭐ Github Dataset Overview In the table below, we provide a brief summary of the dataset statistics. Category Size Total Sample 2019163 Total Image 2019163 Average Answer Length 84 Maximum Answer Length 5851 JSON Overview Each dictionary in the JSON file contains three keys: 'id', 'image', and 'conversations'. The 'id' is the unique identifier for the current data in the entire dataset. The 'image' stores… See the full description on the dataset page: https://huggingface.co/datasets/Ghaser/Wikipedia-Knowledge-2M.image1M<n<10M4 likes77 downloads2y agoHugging Face24Kondapally /TWDS16image10K<n<100K0 likes73 downloads8mo agoHugging Face25kyler0941 /all2024speedtest_1gatedimage1K<n<10K0 likes63 downloads2y agoHugging Face26exander /kitti-depth-gtimage1K<n<10K0 likes63 downloads1y agoHugging Face27kingsidharth /zangei-dit-stage-1-250k-256px-imgimage100K<n<1M0 likes57 downloads4mo agoHugging Face28kyler0941 /all2024speedtest_0gatedimage1K<n<10K0 likes55 downloads2y agoHugging Face29KZ-ucsc /VLMBench_datasetimage100K<n<1M0 likes55 downloads1y agoHugging Face30kyler0941 /all2024speedtest_7gatedimage1K<n<10K0 likes53 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.