CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sophia1ch /zendo-synthetic-data Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either follows ("positive", label=1) or violates ("negative", label=0) a rule that is given in natural language and as a Prolog query. Splits split scenes train 56475 test 3344 rules total 3439 Layout images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.imageimage-classification10K<n<100K1 likes4.6k downloads4mo agoHugging Face02sewa-rural-care /anemia-survey-datasetgated Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India Dataset: sewa-rural-care/anemia-survey-dataset Contact: sewarural@ymail.com Version: 1.0 — July 2026 Dataset Summary This dataset supports research into non-invasive, smartphone-based anemia screening applicable to low-resource and rural healthcare settings. It was collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.tabularimage-classification1K<n<10K7 likes1.4k downloads3mo agoHugging Face03Ardea /Icarus-dataset Icarus A unified multi-modal curriculum dataset for evolutionary neural architecture search. Every row is one self-contained Task = {meta, support, query}, where support and query are lists of (input_Field, output_Field) pairs. The inner loop trains on support; fitness is scored on query. Support is non-empty for every task. Encoders read the Field descriptor (axes, value_type, n_classes, value_range, mask); mask is True where a value is padding/ignored. meta.class_names, when… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/Icarus-dataset.textimage-classification10K<n<100K2 likes1.2k downloads3mo agoHugging Face04mlech26l /liquidrandom-data liquidrandom-data Diverse seed data for ML/LLM training data generation pipelines. Used by the liquidrandom Python package. Dataset Summary This dataset contains 520,080 seed data samples across 24 categories, generated using a hierarchical taxonomy tree approach with LLM-based quality validation and fuzzy deduplication. Data is stored as Parquet with zstd compression. Categories Category Samples File Coding Tasks 30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.tabulartext-generation100K<n<1M0 likes1k downloads2mo agoHugging Face05paulpacaud /rlbenchfail_train_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.tabularvisual-question-answering10K<n<100K0 likes968 downloads7mo agoHugging Face06anonymousllbench /llbench-dataset LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute. LL-Bench is a large-scale, human-preference benchmark for evaluating low-level vision restoration in the era of large generative models (LGMs). It compares 10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.imageimage-to-image100K<n<1M0 likes825 downloads4mo agoHugging Face07nesteo-datasets /nesteo-prototype NestEO: Modular and Hierarchical EO Dataset Framework NestEO is a hierarchical, resolution-aligned, UTM-based nested grid dataset framework supporting general-purpose, multi-scale multimodal Earth Observation workflows. Built from diverse EO sources and enriched with metadata for landcover, climate zones, and population, it enables scalable, representative and progressive sampling for AI4EO. Grid Levels: 120000m, 12000m, 2400m, 1200m, 600m, 300m, 150mGrid Metadata: ESA WorldCover… See the full description on the dataset page: https://huggingface.co/datasets/nesteo-datasets/nesteo-prototype.tabularimage-segmentation10K<n<100K1 likes539 downloads1y agoHugging Face08datapointai /text-2-image-human-preferences-2mgated Text-to-image human preferences: 2M votes across 30 models This dataset contains the complete voting record behind the Datapoint Image Bench leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of 216,116 image pairs. The votes compare 30 text-to-image models in a complete round-robin on 500 prompts, judged by annotators from over 200 countries. Every vote includes the annotator's trust score at the time the vote was cast. Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.imagetext-to-image1M<n<10M21 likes438 downloads1mo agoHugging Face09zentardev /handwritten-digit-dataset Handwritten Digit Dataset This dataset contains a collection of handwritten digits (0-9) contributed by users through an interactive web-based drawing application. The dataset is continuously updated, reflecting real-world human handwriting variability. Dataset Details The images are pre-processed to match the standard machine learning format for digit recognition: Dimensions: 28x28 pixels. Format: Grayscale (single channel). Processing: Each digit is cropped to… See the full description on the dataset page: https://huggingface.co/datasets/zentardev/handwritten-digit-dataset.tabularimage-classificationn<1K0 likes393 downloads16h agoHugging Face10paulpacaud /rlbenchfail_test_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes391 downloads7mo agoHugging Face11paulpacaud /rlbenchfail_val_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_val_dataset.tabularvisual-question-answering1K<n<10K0 likes350 downloads7mo agoHugging Face12aiacademy-kg /house_kg_full_dataset house.kg — Kyrgyzstan Real Estate (multimodal) A complete snapshot of house.kg, the largest real-estate board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller identities, agency ratings, reviews — and 227,294 photographs. Field names are English; values are kept in the original language (Russian/Kyrgyz), exactly as the site renders them. 💻 Scraper source code on GitHub → The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.imagetabular-regression100K<n<1M0 likes344 downloads2mo agoHugging Face13Davidup1 /GeoChrono-Data ChronoBench & ChronoInstruct ChronoBench is a comprehensive, multi-dimensional, and multi-granularity benchmark for high-resolution long-temporal remote sensing understanding. It decomposes long-term remote sensing understanding into a four-level cognitive hierarchy — from Land Cover Perception through Temporal Recognition and Long-Term Memory to Spatio-Temporal Reasoning — comprising 12 sub-tasks and 17,689 rigorously validated QA pairs derived from 3,469 high-resolution… See the full description on the dataset page: https://huggingface.co/datasets/Davidup1/GeoChrono-Data.tabularvisual-question-answering10K<n<100K0 likes276 downloads1mo agoHugging Face14mrdbourke /Recap-DataComp-1B-FoodOrDrink Recap-DataComp-1B: Food or Drink A filtered subset of Recap-DataComp-1B containing 106,230,157 rows classified as food/drink content, enriched with structured food/drink extraction from FoodExtract-v2. Overview Count Percentage Total rows 106,230,157 100% Food/drink (Stage 5 label) 96,618,895 91.0% Not food/drink (Stage 5 label) 9,611,262 9.0% FoodExtract (re_caption): food/drink 79,519,489 74.9% FoodExtract (re_caption): not food/drink 26,710,156… See the full description on the dataset page: https://huggingface.co/datasets/mrdbourke/Recap-DataComp-1B-FoodOrDrink.imagetext-classification100M<n<1B1 likes247 downloads6mo agoHugging Face15cycloevan /Ransomware_PE_Header_Feature_Dataset Dataset Card for Ransomware PE Header Feature Dataset Dataset Description Dataset Summary This dataset contains PE header features (first 1024 bytes) from 2,157 Windows executable samples, comprising 1,134 legitimate software (goodware) and 1,023 ransomware samples across 25 ransomware families. Each sample is represented by numerical features extracted from the raw PE header. Supported Tasks Binary Classification: Distinguish between goodware and… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/Ransomware_PE_Header_Feature_Dataset.documentimage-classification1K<n<10K0 likes231 downloads8mo agoHugging Face16drksci /trade_vision_dataset TradeVision: Hierarchical Physical Business & Multimodal Retail Provenance Dataset This dataset is continuously seeded from OpenStreetMap, matched to Google Place IDs, harvested for temporal store photos, and enriched with zero-shot computer vision using Hugging Face Hub native pipelines. Dataset Structure The dataset is partitioned into three relational subsets loadable via Hugging Face datasets: from datasets import load_dataset # 1. Load Canonical Businesses… See the full description on the dataset page: https://huggingface.co/datasets/drksci/trade_vision_dataset.tabularzero-shot-object-detectionn<1K0 likes214 downloads18h agoHugging Face17nrl-ai /anylearning-data AnyLearning datasets This repository contains reproducible sample datasets used to develop and test AnyLearning OSS. Dataset licenses are recorded in LICENSES.md. The repository's scripts and original documentation are Apache-2.0, but that license does not override the terms of any dataset. Check the dataset license before use. Licence-cleared Task Dataset Licence Image classification ZhangLabData: Chest X-Ray CC BY 4.0 Object detection Safety Helmet… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/anylearning-data.imageimage-classification1K<n<10K0 likes197 downloads24d agoHugging Face18notgoodkeeper /cnn-based-drowsiness-detection-data CNN-Based Drowsiness Detection - Dataset Preprocessed, auto-labeled face-crop images used to train the model in notgoodkeeper/cnn-based-drowsiness-detection. Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection Collection Frames were captured from a webcam, then run through: Haar Cascade face detection -> crop + pad + resize to 412x412 MediaPipe Selfie Segmentation -> background replaced with white CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.imageimage-classification1K<n<10K1 likes191 downloads20d agoHugging Face19paulpacaud /Guardian-FailCoT-OOD-datasets Guardian FailCoT — Out-of-Distribution Real-Robot Benchmarks This repository bundles the three real-world failure-detection benchmarks used to evaluate the Guardian vision-language model in the paper Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation (Pacaud et al., 2026): UR5-Fail — our newly collected three-view real-robot benchmark. RoboFail — single-view real-robot manipulation failure benchmark from Liu et al. (CoRL 2023). RoboVQA —… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/Guardian-FailCoT-OOD-datasets.tabularvisual-question-answering1K<n<10K1 likes189 downloads5mo agoHugging Face20Kimhi /hardness_data_mix Hardness Data Mix - Resolution Sufficiency Dataset A large-scale dataset of document images with labels indicating the minimum resolution required to accurately answer questions about those documents. Dataset Description This dataset contains 81,924 document image-question pairs labeled with resolution sufficiency information. Each sample is annotated with a "hardness" label indicating the minimum resolution level needed to answer questions about that document accurately.… See the full description on the dataset page: https://huggingface.co/datasets/Kimhi/hardness_data_mix.tabularimage-classification10K<n<100K2 likes156 downloads6mo agoHugging Face21avihayamor /tripmatch-ai-dataset TripMatch AI Dataset A reproducible multimodal dataset for the TripMatch AI Final Project. It contains 10,000 synthetic text trip plans with a raw idea generated for every row by the pretrained Hugging Face model google/flan-t5-small, plus 5,000 real street-view images retained as extra multimodal work. The two configurations are separate so Dataset Viewer can load each schema correctly. Dataset statistics Configuration Rows Main fields Intended task… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-dataset.imagetext-retrieval10K<n<100K0 likes148 downloads2mo agoHugging Face22paulpacaud /bdv2fail_train_dataset Guardian: BridgeDataV2-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_train_dataset.tabularvisual-question-answering10K<n<100K1 likes141 downloads7mo agoHugging Face23paulpacaud /ur5fail_test_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.tabularvisual-question-answeringn<1K1 likes125 downloads7mo agoHugging Face24paulpacaud /ur5fail_train_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_train_dataset.tabularvisual-question-answering1K<n<10K0 likes121 downloads7mo agoHugging Face25paulpacaud /bdv2fail_val_dataset Guardian: BridgeDataV2-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_val_dataset.tabularvisual-question-answering1K<n<10K0 likes96 downloads7mo agoHugging Face26minute1028 /forestllava-dataset Forest-LLaVA Multimodal Tree-Species Dataset Forest-LLaVA is a multimodal remote-sensing dataset for tree-species recognition and structured vision-language research. Each record is indexed by a numeric sample_id from the US subset of GlobalGeoTree and is linked to a 60 m × 60 m patch from NAIP Optical, Sentinel-2 MSI and Sentinel-1 SAR data, four-level taxonomic labels and geographic/environmental records. The repository contains the complete image archives for the… See the full description on the dataset page: https://huggingface.co/datasets/minute1028/forestllava-dataset.tabularimage-classification100K<n<1M1 likes94 downloads10d agoHugging Face27paulpacaud /ur5fail_val_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_val_dataset.tabularvisual-question-answeringn<1K0 likes91 downloads7mo agoHugging Face28paulpacaud /bdv2fail_test_dataset Guardian: BridgeDataV2-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes71 downloads7mo agoHugging Face29KBlueLeaf /danbooru2023-metadata-databasegated Metadata Database for Danbooru2023 Danbooru 2023 datasets: https://huggingface.co/datasets/nyanko7/danbooru2023 The latest entry of this database is id 7,866,491. Which is newer than nyanko7's dataset. This dataset contains a sqlite db file which have all the tags and posts metadata in it. The Peewee ORM config file is provided too, plz check it for more information. (Especially on how I link posts and tags together) The original data is from the official dump of the posts info.… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-metadata-database.imageimage-classification1M<n<10M83 likes68 downloads2y agoHugging Face30aiacademy-kg /house_kg_full_dataset_frames house.kg — Kyrgyzstan Real Estate, over time Sale and rental listings scraped from house.kg, the largest real-estate board in Kyrgyzstan, re-measured on a schedule. Field names are English; values are kept in the original language (Russian), exactly as the site renders them. Coverage: 2026-09-08. This is the baseline snapshot; later runs append new partitions. Subsets subset rows description listings 25,264 one row per advertisement — current state plus… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset_frames.imagetabular-regression100K<n<1M0 likes65 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.