CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saberzl /SID_Set Dataset Card for SID_Set Dataset Summary We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages: Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations. Broad diversity: Encompassing fully synthetic and tampered images across various classes. Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection. Please check… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/SID_Set.imagetext-to-image100K<n<1M15 likes11k downloads1y agoHugging Face02licyk /image_training_set自用的训练集合集,用于 Stable Diffusion 模型微调。 该仓库仅用于存档,不提供任何技术支持。 imagen<1K2 likes6.7k downloads23d agoHugging Face03GaussianWorld /scannet_mini_val_set_suiteimage1 likes5.6k downloads1y agoHugging Face04vidore /colpali_train_set Dataset Description This dataset is the training set of ColPali it includes 127,460 query-image pairs from both openly available academic datasets (63%) and a synthetic dataset made up of pages from web-crawled PDF documents and augmented with VLM-generated (Claude-3 Sonnet) pseudo-questions (37%). Our training set is fully English by design, enabling us to study zero-shot generalization to non-English languages. Dataset #examples (query-page pairs) Language DocVQA 39… See the full description on the dataset page: https://huggingface.co/datasets/vidore/colpali_train_set.imagedocument-question-answering100K<n<1M93 likes5.5k downloads1y agoHugging Face05saberzl /So-Fake-Set Dataset Card for So-Fake-Set Dataset Summary We provide So-Fake-Set, A large-scale, diverse dataset tailored for social media image forgery detection! Please check our website to explore more visual results. Dataset Structure "image" (Image): Input images, including real, full_synthetic, and tampered images. "mask" (Image): Binary mask highlighting manipulated regions in tampered images. "label" (str): Classification category. "generator" (str): The… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/So-Fake-Set.image1M<n<10M11 likes4.2k downloads11mo agoHugging Face06JamalLee /Omni-Fake-SET Omni-Fake-SET Omni-Fake-SET is the in-distribution split of Omni-Fake, a unified multimodal deepfake dataset for social-media forensics. It covers image, audio, video, and audio–video talking-head (AV-TH) modalities. Each modality uses the same three-way label space: real, fully synthetic, and tampered. Pair with the held-out benchmark Omni-Fake-OOD for out-of-distribution evaluation. Paper: arXiv:2605.01638 Project page: Omni-Fake License: CC-BY-4.0 Video (hybrid… See the full description on the dataset page: https://huggingface.co/datasets/JamalLee/Omni-Fake-SET.audioimage-classification1M<n<10M5 likes3k downloads3mo agoHugging Face07AbstractPhil /diffusion-pretrain-set-ft1 diffusion-pretrain-set-ft1 A multi-source image-caption pretraining dataset assembled from ten upstream sources via a uniform ingest pipeline. Designed for a full pretrain or finetune pipeline meant to curate for any major diffusion model preliminary, with the sole intent to create a more powerful baseline preliminary train and a baseline for synthesizing images to train the next generation of the VLM model. This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.image1M<n<10M2 likes1.9k downloads3mo agoHugging Face08Voxel51 /dronescapes2_annotated_train_set Dataset Card for DroneScapes2 (annotated train set) This is a FiftyOne dataset with 218 samples. It's a subset of this split from the original repo. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/dronescapes2_annotated_train_set.imageimage-classification1K<n<10K2 likes1.8k downloads10mo agoHugging Face09MBZUAI /Omni-Setsgated Omni-Sets A large-scale, multi-modal instruction-tuning dataset spanning six modalities (audio, speech, image, video, visual documents, and cross-modal omni) with both single-turn dense captions and multi-turn instruction-following conversations. Designed for training omni-modal language models that can perceive and reason across all modalities. 590,858 total samples | 5,635 hours of audio/video | 6 configs | 17 source datasets Overview Config Modality… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Omni-Sets.audio100K<n<1M0 likes1.6k downloads22d agoHugging Face10Voxel51 /Set5 Dataset Card for Set5 This is a FiftyOne dataset with 135 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = fouh.load_from_hub("Voxel51/Set5") # Launch the App session = fo.launch_app(dataset) Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Set5.imageimage-to-imagen<1K1 likes1.5k downloads2y agoHugging Face11Ashen0li /tft-set17-unit-detector-yolo TFT Set17 Unit Detector YOLO YOLO-format unit detector dataset for TFT Set 17 experiments. This package combines: synthetic board screenshots made from arena textures and modelviewer unit renders clean multi-angle modelviewer unit reference images The dataset is intended for training a single-class unit object detector. Structure images/train/*.jpg images/val/*.jpg labels/train/*.txt labels/val/*.txt data.yaml classes.txt manifest.json Counts… See the full description on the dataset page: https://huggingface.co/datasets/Ashen0li/tft-set17-unit-detector-yolo.imageobject-detection10K<n<100K0 likes1.4k downloads3mo agoHugging Face12Voxel51 /Set14 Dataset Card for Set14 This is a FiftyOne dataset with 378 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = fouh.load_from_hub("Voxel51/Set14") # Launch the App session = fo.launch_app(dataset) Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Set14.imageimage-to-imagen<1K2 likes1.4k downloads2y agoHugging Face13siyrus /BToks-vidore_colpali_train_set BToks ViDoRe ColPali Train Set This dataset repository contains Lance-format converted data used by the open-source reproduction code for Bottleneck Tokens for Unified Multimodal Retrieval (arXiv:2604.11095). Source Converted from vidore/colpali_train_set. This repository does not change upstream ownership, licensing, citation requirements, or usage restrictions. Format The data is stored as Lance tables for the BToks/VLM2Emb training and evaluation… See the full description on the dataset page: https://huggingface.co/datasets/siyrus/BToks-vidore_colpali_train_set.imageimage-to-text100K<n<1M0 likes1.3k downloads3mo agoHugging Face14InsultedByMathematics /diffbir-mixed-setsimage1M<n<10M0 likes1k downloads1y agoHugging Face15benzlxs /objaverse_rendering_setimage10M<n<100M0 likes966 downloads1y agoHugging Face16AbstractPhil /diffusion-pretrain-set-ft1-1024 diffusion-pretrain-set-ft1-1024 1024px (2x) upscale of AbstractPhil/diffusion-pretrain-set-ft1. WARNING MUCH OF THIS DATA WAS MODEL UPSCALED USING RAPID UPSCALERS. THIS IS NOT CONSISTENTLY HIGH FIDELITY NOR IS IT EVEN CLOSE TO FAIR FIDELITY AT TIMES. PLEASE use this ONLY for pretraining, new concepts, and simple design purposes ONLY. HEAVILY PRUNE FOR FINETUNING. Thank you, good luck my friends. Details Model: realesr-general-x4v3 (SRVGG Compact… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1-1024.image1M<n<10M0 likes966 downloads4mo agoHugging Face17WissMah /lebanese_aug_setimage10K<n<100K0 likes672 downloads11mo agoHugging Face18RAID-techjam /SID_Set Dataset Card for SID_Set Dataset Summary We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages: Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations. Broad diversity: Encompassing fully synthetic and tampered images across various classes. Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection. Please… See the full description on the dataset page: https://huggingface.co/datasets/RAID-techjam/SID_Set.imagetext-to-image100K<n<1M0 likes654 downloads26d agoHugging Face19facebook /emu_edit_test_set Dataset Card for the Emu Edit Test Set Dataset Summary To create a benchmark for image editing we first define seven different categories of potential image editing operations: background alteration (background), comprehensive image changes (global), style alteration (style), object removal (remove), object addition (add), localized modifications (local), and color/texture alterations (texture). Then, we utilize the diverse set of input images from the MagicBrush… See the full description on the dataset page: https://huggingface.co/datasets/facebook/emu_edit_test_set.image1K<n<10K47 likes641 downloads3y agoHugging Face20jwengr /So-Fake-Set-Resized-224image1M<n<10M0 likes628 downloads6mo agoHugging Face21prmkkbb /veri_seti_adiimage1K<n<10K0 likes603 downloads2mo agoHugging Face22Beastarz /SID_Set Dataset Card for SID_Set Dataset Summary We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages: Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations. Broad diversity: Encompassing fully synthetic and tampered images across various classes. Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection. Please… See the full description on the dataset page: https://huggingface.co/datasets/Beastarz/SID_Set.imagetext-to-image100K<n<1M0 likes595 downloads26d agoHugging Face23setrsoft /climbing-holds [!IMPORTANT] This dataset is in construction. The current files are raw scans intended for establishing the structure. Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide. GUI for contributions https://setrsoft.github.io/holds-dataset-hub/ Or send your files here Climbing Holds 3D dataset (SetRsoft) 📋 Project Overview This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.3dn<1K0 likes580 downloads5mo agoHugging Face24nomic-ai /colpali_train_set_split_by_sourceimage100K<n<1M2 likes569 downloads2y agoHugging Face25sav7669 /sroie_data_setimagen<1K0 likes468 downloads3y agoHugging Face26nanxidajun /NuosuBburma-OCR-Evaluation-Set NuosuBburma OCR Evaluation Set 规范彝文 OCR 评估集 用于规范彝文(NuosuBburma)的模型性能评估与错误分析。 内容涵盖纯彝文、彝汉混排及少量含拉丁字母的混排文本;场景覆盖书籍、手写、屏幕和实拍。 真实性声明:全部评估样本来自真实扫描或实际拍摄,不含合成评估数据。 评估任务 输入为单张包含规范彝文、彝汉混排及少量拉丁字母、数字或标点的图像。模型按照图像中的视觉阅读顺序转写可见文字,保留必要的行结构和原有字符,不进行翻译、改写或文本补全。评估重点是复杂版式、混排文字和真实拍摄干扰条件下的文字识别能力。 在线可视化 查看评估集分布与图像—标准答案对照 内容统计 项目 数量 样本 / 图片 1030 / 1030 文档页 / OCR 实例 519 / 511 图像类别 类别 数据形态 主要干扰 评估重点… See the full description on the dataset page: https://huggingface.co/datasets/nanxidajun/NuosuBburma-OCR-Evaluation-Set.imageimage-to-text1K<n<10K1 likes454 downloads2mo agoHugging Face27physicl-community /fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde Fast-Food Cleaning Robot — Floor Mess Dataset Training dataset for a cleaning robot operating in fast-food-style food-service spaces (break areas / dining). Scenes are staged in break-area environments cluttered with food-service furnishings and food items (pizza, grocery food, cups, spoons) so the robot learns to perceive and act on mess. Covers detection, grasping, navigation, obstacle avoidance and pick-and-place. Renders are 1024x1024 with RGB plus albedo, metric depth and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde.imagen<1K0 likes437 downloads21d agoHugging Face28duyle2408 /set_fcos_runsimage1K<n<10K0 likes427 downloads6d agoHugging Face29Voxel51 /DCVAI-Challenge-Public-Eval-SetgatedThis is a FiftyOne dataset with 7,436 samples. Installation If you haven''t already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include ''max_samples'', etc dataset = fouh.load_from_hub("Voxel51/DCVAI-Challenge-Public-Eval-Set") # Launch the App session = fo.launch_app(dataset) Dataset Card for Public Evaluation set for the Data… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/DCVAI-Challenge-Public-Eval-Set.imageobject-detection1K<n<10K1 likes399 downloads2y agoHugging Face30Voxel51 /Data-Centric-Visual-AI-Challenge-Train-Setgated Dataset Card for Data-Centric-Visual-AI-Train-Set This is a FiftyOne dataset with 30,000 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo import fiftyone.utils.huggingface as fouh # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = fouh.load_from_hub("Voxel51/Data-Centric-Visual-AI-Challenge-Train-Set") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Data-Centric-Visual-AI-Challenge-Train-Set.imageobject-detection10K<n<100K1 likes395 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.