CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dalle-mini /witimage1M<n<10M7 likes134k downloads5y agoHugging Face02timm /mini-imagenet Dataset Description A mini version of ImageNet-1k with 100 of 1000 classes present. Unlike some 'mini' variants this one includes the original images at their original sizes. Many such subsets downsample to 84x84 or other smaller resolutions. Data Splits Train 50000 samples from ImageNet-1k train split Validation 10000 samples from ImageNet-1k train split Test 5000 samples from ImageNet-1k validation split (all 50 samples per class)… See the full description on the dataset page: https://huggingface.co/datasets/timm/mini-imagenet.imageimage-classification10K<n<100K28 likes6.9k downloads2y agoHugging Face03GaussianWorld /scannet_mini_val_set_suiteimage1 likes5.8k downloads1y agoHugging Face04antofuller /mini-VTAB Mini-VTAB A collection of VTAB (Visual Task Adaptation Benchmark) datasets. We sampled 1K training samples and 1K testing samples for each task. Tasks datasets = [ "caltech101", "cifar10", "cifar100", "dtd", "flowers", "pets", "sun397", "svhn", "pcam", "eurosat", "resisc45", "diabetic_retinopathy", "clevr_count_all", "clevr_closest_object_distance", "dmlab", "dsprites_label_x_position"… See the full description on the dataset page: https://huggingface.co/datasets/antofuller/mini-VTAB.image10K<n<100K1 likes3k downloads8mo agoHugging Face05pollen-robotics /reachy-mini-wall-data Reachy Mini — wall data (public) posts.json for the Reachy Mini community wall: the AI-filtered posts shown publicly, aggregated from Bluesky, YouTube, LinkedIn, TikTok, X and Reddit by the social-wall pipeline. Fetch it directly (CORS-enabled) from any static site: const url = "https://huggingface.co/datasets/pollen-robotics/reachy-mini-wall-data/resolve/main/posts.json"; const posts = await (await fetch(url)).json(); Each item: id, platform, author, handle, avatar, text… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-wall-data.imagen<1K0 likes2.9k downloads1d agoHugging Face06badincite /minimax-h3-soup MiniMax H3 Soup Reproducibility archive for a local ComfyUI MiniMax H3 Ref2V benchmark on an RTX 3090. What is included Original benchmark workflow graph (source_prompt.json), manifest, and result table. Every one-second MP4 from the original C1-C11 benchmark grid and its Euler repeat sweep. The separate duration experiments are intentionally not included. Labeled C1-C11 visual contact sheets, Ref2VA stock-control sheets, and a static render-time summary chart.… See the full description on the dataset page: https://huggingface.co/datasets/badincite/minimax-h3-soup.imagen<1K5 likes2.1k downloads12d agoHugging Face07dalle-mini /open-imagesimage1M<n<10M27 likes1.8k downloads5y agoHugging Face08Mini-o3 /VisualProbe_trainimage1K<n<10K3 likes1.6k downloads1y agoHugging Face09antofuller /mini-VTAB-corruptions mini-VTAB-C A collection of VTAB (Visual Task Adaptation Benchmark) datasets. We sampled 1K training samples and 1K testing samples for each task. For each test set, we apply all 15 corruption types from ImageNet-C. Tasks datasets = [ "caltech101", "cifar10", "cifar100", "dtd", "flowers", "pets", "sun397", "svhn", "pcam", "eurosat", "resisc45", "diabetic_retinopathy", "clevr_count_all"… See the full description on the dataset page: https://huggingface.co/datasets/antofuller/mini-VTAB-corruptions.image100K<n<1M0 likes1.5k downloads8mo agoHugging Face10M3LEO-miniset /conus Contiguous UNited States AOI Randomly sampled 5000 data tiles (3000 for training, and 1000 for validation and testing). Occasionally, not all tiles were available for all datasets. We include a detailed breakdown of the number of chips per dataset below. Datakind Chips s1grd-2020 5000 gssic 5000 gunw_2020-04-01_2020-06-30 4138 gunw_2020-07-01_2020-09-30 4171 gunw_2020-08-01_2020-10-31 4166 gunw_2020-01-01_2020-03-31 4097 s2rgbm-2020 5000 biomass-2020 5000… See the full description on the dataset page: https://huggingface.co/datasets/M3LEO-miniset/conus.image10K<n<100K0 likes1.1k downloads2y agoHugging Face11M3LEO-miniset /middleeast Middle East AOI Randomly sampled 5000 data tiles (3000 for training, and 1000 for validation and testing). Occasionally, not all tiles were available for all datasets. We include a detailed breakdown of the number of chips per dataset below. Datakind Chips s1grd-2020 5000 gssic 4829 gunw_2020-04-01_2020-06-30 4641 gunw_2020-01-01_2020-03-31 4561 s2rgbm-2020 5000 biomass-2020 5000 esaworldcover-2020 5000 modis44b006veg 5000 ghsbuilts-2020 5000 srtmdem… See the full description on the dataset page: https://huggingface.co/datasets/M3LEO-miniset/middleeast.image10K<n<100K0 likes1k downloads2y agoHugging Face12GATE-engine /mini_imagenet Dataset Card for "mini_imagenet" More Information needed imageimage-classification10K<n<100K8 likes966 downloads3y agoHugging Face13jingyaogong /minimind-v_dataset Ⅰ 数据集 本轮训练用到的图文数据全部来自 ALLaVA-4V 系列。 相比以往从几份 LLaVA 衍生集拼接得到的数据,ALLaVA-4V 的质量更整齐、中英双语原生对照,细粒度描述也更充分。 它由两个子源构成:一份是 LAION 里挑出来的高质量图片(自然图像为主),一份是 VFLAN 指令流里挑出来的图片(文档、图表、合成场景居多)。 Pretrain(pretrain_i2t.parquet,约 127 万条 / ~64 万张唯一图像) ALLaVA-Caption-LAION-4V 英/中:~47万 + ~44万ALLaVA-Caption-VFLAN-4V 英/中:~19万 + ~17万 任务形式为"请描述这张图片"类的单轮长描述,用于让模型建立视觉 token 到语言 token 的基础对齐。 SFT(sft_i2t.parquet,约 290 万条 / ~65 万张唯一图像) ALLaVA-Instruct-LAION-4V 英/中:~47万 + ~47万… See the full description on the dataset page: https://huggingface.co/datasets/jingyaogong/minimind-v_dataset.imagevisual-question-answeringn<1K36 likes934 downloads5mo agoHugging Face14BastienATOS /mini-reachy-animation Reachy Mini Animation Dataset Multi-view renders of 85 emotional animations performed by the Reachy Mini robot, paired with the full robot joint state for every single frame. Source of the animations. The emotional animations rendered here come from the official pollen-robotics/reachy-mini-emotions-library dataset by Pollen Robotics. This dataset re-renders those emotions from 12 camera angles (with 3 background variants) and pairs every frame with the robot's joint state.… See the full description on the dataset page: https://huggingface.co/datasets/BastienATOS/mini-reachy-animation.imagerobotics100K<n<1M1 likes922 downloads2mo agoHugging Face15juiceb0xc0de /MiniCPM5-1B-atlas juiceb0xc0de/MiniCPM5-1B-atlas A brain atlas for openbmb/MiniCPM5-1B, a 1B on-device model with a 130k bilingual vocabulary. This is not a chat dataset or a benchmark. It is an internal-mechanics map, built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. If you want to know which parts of this model are safe to edit, where its output-vocabulary directions live, or which layers are carrying the most… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/MiniCPM5-1B-atlas.image100K<n<1M1 likes873 downloads7d agoHugging Face16oakmindai /minimax_h3_avatar_500 Watch the full 500-video showcase on YouTube MiniMax H3 Avatar 500 An image-to-video dataset pairing reference avatar images with detailed generation prompts and generated avatar videos. This release contains 500 curated examples in both a browsable raw layout and a typed Hugging Face dataset. Version 1.0 · Released August 14, 2026 Dataset contents Each example contains: A 1024 × 1024 reference avatar image A detailed English generation prompt A generated 640 ×… See the full description on the dataset page: https://huggingface.co/datasets/oakmindai/minimax_h3_avatar_500.imagen<1K3 likes838 downloads1mo agoHugging Face17unsloth /Radiology_mini0.33% sampled from https://huggingface.co/datasets/eltorio/ROCOv2-radiology image1K<n<10K46 likes803 downloads2y agoHugging Face18dalle-mini /vqgan-pairs VQGAN Pairs This dataset contains ~2.4 million image pairs intended for improvement of image quality in VQGAN predictions. Each pair consists of: A 512x512 crop of an image taken from Open Images. A 256x256 image encoded and decoded using VQGAN, corresponding to the same image crop as the original. This is the VQGAN implementation that was used for encoding and decoding: https://github.com/patil-suraj/vqgan-jax License This dataset is created using Open Images… See the full description on the dataset page: https://huggingface.co/datasets/dalle-mini/vqgan-pairs.imageother1M<n<10M5 likes710 downloads4y agoHugging Face19previtus /OxHyperMinerals_MINIimagen<1K0 likes704 downloads2y agoHugging Face20unsloth /llava-instruct-mix-vsft-miniOriginally from https://huggingface.co/datasets/HuggingFaceH4/llava-instruct-mix-vsft but 0.33% randomnly sampled image1K<n<10K4 likes695 downloads2y agoHugging Face21Mini-o3 /VisualProbe_Easyimagen<1K1 likes675 downloads1y agoHugging Face22Mini-o3 /VisualProbe_Hardimagen<1K1 likes664 downloads1y agoHugging Face23LibreYOLO /coco-val2017-mini500 COCO val2017 mini-500 A frozen, reproducible 500-image subset of COCO val2017 for fast object-detection latency/throughput benchmarking on edge devices (e.g. LibreYOLO on NVIDIA Jetson), where the full 5000-image val set is impractical. Selection (deterministic) The 500 images are sorted(COCO.getImgIds())[:500] of the official instances_val2017.json (the 500 lowest image IDs) — identical to running the vision-analysis-benchmark harness with --limit 500.… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/coco-val2017-mini500.imageobject-detectionn<1K0 likes649 downloads3mo agoHugging Face24yxma /gelsight-mini-pretrain GelSight Mini Pretrain ~853K GelSight Mini tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92. Frames Sources Real 536K FoTA (labeled+unlabeled), 3DCal, FEATS, GelSLAM, TactileTracking, RTM, FeelAnyForce, UniT, TacQuad Sim 317K sim_tactile_mnist, sim_starstruck (Taxim-rendered, Mini-calibrated) NC… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain.imageimage-classification100K<n<1M0 likes634 downloads2mo agoHugging Face25Mini-o3 /VisualProbe_Mediumimagen<1K1 likes630 downloads1y agoHugging Face26MiniMedMind /VQA-RADimagen<1K0 likes576 downloads2y agoHugging Face27dnth /mini-imagenetimage10K<n<100K0 likes544 downloads2y agoHugging Face28M3LEO-miniset /china China AOI Randomly sampled 5000 data tiles (3000 for training, and 1000 for validation and testing). Occasionally, not all tiles were available for all datasets. We include a detailed breakdown of the number of chips per dataset below. Datakind Chips s1grd-2020 5000 gssic 5000 gunw_2020-01-01_2020-03-31 4354 s2rgbm-2020 5000 biomass-2020 5000 esaworldcover-2020 5000 modis44b006veg 5000 ghsbuilts-2020 5000 srtmdem 5000 We provide '.csv' files with… See the full description on the dataset page: https://huggingface.co/datasets/M3LEO-miniset/china.image10K<n<100K0 likes533 downloads2y agoHugging Face29aaryavlal /arbiter-mini Arbiter-mini A small, purpose-built image dataset of household items captured under controlled Raspberry Pi camera conditions and labeled for binary waste/recycle classification according to San Diego, CA municipal recycling rules. Built as deployment-condition training data for the Arbiter sorting system, intended to be used alongside TrashNet to close the domain gap between studio imagery and real Pi-camera inference. Motivation Models trained purely on TrashNet… See the full description on the dataset page: https://huggingface.co/datasets/aaryavlal/arbiter-mini.imageimage-classificationn<1K1 likes501 downloads3mo agoHugging Face30sayakpaul /OmniEdit-miniMini version of TIGER-Lab/OmniEdit-Filtered-1.2M for rapid experimentation. Script used: from huggingface_hub import dataset_info, snapshot_download import glob from datasets import Dataset import random random.seed(2025) def download_mini_omniedit_files(): repo_id = "TIGER-Lab/OmniEdit-Filtered-1.2M" files = dataset_info(repo_id) files = {f.rfilename for f in files.siblings if "data/" in f.rfilename} files = sorted(list(files)) print(files[:5]) random.shuffle(files)… See the full description on the dataset page: https://huggingface.co/datasets/sayakpaul/OmniEdit-mini.image10K<n<100K6 likes479 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.