CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Narsil /image_dummy\audion<1K0 likes147k downloads5y agoHugging Face02safwatkhokha /nawaqes-backup-v2imagen<1K13 likes39k downloads13d agoHugging Face03Goku-OpenLab /nano-banana-pro-prompts-datasets 🖼️ Nano Banana Pro Prompt Dataset 🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.imagetext-to-image10K<n<100K1 likes21k downloads2mo agoHugging Face04FrontisAI /NatureBench Dataset Card for NatureBench NatureBench is a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, spanning 6 scientific domains. It is designed to evaluate whether AI coding agents can move beyond reproduction toward discovery: each task asks an agent to solve a real scientific machine-learning problem and is scored against the source paper's reported state of the art. 📄 arXiv paper: https://arxiv.org/abs/2606.24530 💻 GitHub code… See the full description on the dataset page: https://huggingface.co/datasets/FrontisAI/NatureBench.imagen<1K14 likes17k downloads9d agoHugging Face05nader39 /audio-filesaudion<1K0 likes8.1k downloads22h agoHugging Face06naver-clova-ix /cord-v2image1K<n<10K126 likes8k downloads4y agoHugging Face07safwatkhokha /nawaqes-backupimagen<1K5 likes7.9k downloads2mo agoHugging Face08nateraw /imagenet-sketch-dataimage0 likes6.5k downloads4y agoHugging Face09nannullna /laion_subset Dataset Card for "laion_subset" More Information needed image1K<n<10K1 likes5.8k downloads3y agoHugging Face10naver-clova-ix /synthdog-en Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-en.image10K<n<100K27 likes4.7k downloads3y agoHugging Face11ibm-nasa-geospatial /Landslide4sense Landslide4Sense Dataset Description This dataset is originally introduced in GitHub repo Landslide4Sense-2022. The Landslide4Sense dataset has three splits, training/validation/test, consisting of 3799, 245, and 800 image patches, respectively. Each image patch is a composite of 14 bands that include: Multispectral data from Sentinel-2: B1, B2, B3, B4, B5, B6, B7, B8, B9, B10, B11, B12. Slope data from ALOS PALSAR: B13. Digital elevation model (DEM) from ALOS… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/Landslide4sense.image1K<n<10K7 likes4.7k downloads2y agoHugging Face12sevenc-nanashi /kiiteitte Kiiteitte history Kiiteitte が収集した、今までの選曲履歴。 1時間おきに更新されます。 型 { // 動画ID "video_id": "sm44670499", // タイトル "title": "library->w4nderers / 足立レイ、つくよみちゃん", // 投稿者 "author": "名無し。", // サムネイルのURL "thumbnail": "https://nicovideo.cdn.nimg.jp/thumbnails/44670499/44670499.91820835", // 選曲日時 "date": "2025-02-22 12:51:51", // 新しく増えたお気に入り数。不明の場合は null "new_faves": 5, // 回ったユーザーの数。不明の場合は null "spins": 13, // イチ押しリストのユーザーのURL。イチ押しリスト以外から選曲された場合は null… See the full description on the dataset page: https://huggingface.co/datasets/sevenc-nanashi/kiiteitte.image100K<n<1M2 likes4.3k downloads1h agoHugging Face13seongwon980 /Nameonly_generated Just Say the Name: Online Continual Learning with Category Names Only via Data Generation We provide the dataset used for Name-only continual learning, generated using Stable Diffusion XL, DALL.E-2, CogView2, and DeepFloyd IF models. Disclaimer This dataset is created solely for academic purposes. We minimized human intervention to ensure a fair comparison with the baseline methods discussed in our paper. Despite our efforts, the extensive size of the dataset prevented us… See the full description on the dataset page: https://huggingface.co/datasets/seongwon980/Nameonly_generated.image10K<n<100K1 likes3.5k downloads2y agoHugging Face14BangumiBase /nanarehananare Bangumi Image Base of Na Nare Hana Nare This is the image base of bangumi Na Nare Hana Nare, we detected 146 characters, 6818 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/nanarehananare.image1K<n<10K0 likes3.5k downloads2y agoHugging Face15c13752hz /NavSafe3d1K<n<10K3 likes3.4k downloads13d agoHugging Face16andyvhuynh /NatureMultiView Nature Multi-View (NMV) Dataset Datacard To encourage development of better machine learning methods for operating with diverse, unlabeled natural world imagery, we introduce Nature Multi-View (NMV), a multi-view dataset of over 3 million ground-level and aerial image pairs from over 1.75 million citizen science observations for over 6,000 native and introduced plant species across California. Characteristics and Challenges Long-Tail Distribution: The dataset… See the full description on the dataset page: https://huggingface.co/datasets/andyvhuynh/NatureMultiView.image1M<n<10M10 likes3.3k downloads2y agoHugging Face17naver-clova-ix /synthdog-ko Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-ko.image10K<n<100K18 likes3.2k downloads3y agoHugging Face18BangumiBase /narutoshippuden Bangumi Image Base of Naruto Shippuden This is the image base of bangumi Naruto Shippuden, we detected 196 characters, 36722 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/narutoshippuden.image10K<n<100K0 likes3k downloads3y agoHugging Face19Cognitive-Lab /NayanaOCR_Corpus_2025 🪷 NayanaOCR Corpus 2025 A 1M-page, 22-language fully-parallel synthetic OCR + VQA corpus for document-centric vision-language models — every page rendered in every language. NayanaOCR Corpus 2025 is one of the largest open-source multilingual, multi-task document datasets for training and evaluating OCR, layout detection, and visual question answering (VQA) in low-resource and underrepresented languages. The headline property: it's a true parallel corpus. The same ~45,700 source… See the full description on the dataset page: https://huggingface.co/datasets/Cognitive-Lab/NayanaOCR_Corpus_2025.imageimage-to-text1M<n<10M18 likes2.9k downloads4mo agoHugging Face20KotiyaSanae /nanatsunomaken Bangumi Image Base of Nanatsu No Maken This is the image base of bangumi Nanatsu no Maken, we detected 118 characters, 6989 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/KotiyaSanae/nanatsunomaken.image10K<n<100K0 likes2.4k downloads3y agoHugging Face21andropar /relaion2b-natural LAION-Natural: Naturalness Scores for ReLAION-2B (CCN 2025, Roth & Hebart) LAION-Natural is a large-scale naturalness scoring dataset covering 2.1 billion images from ReLAION-2B-en-research-safe. Each image receives a score predicting how "natural" or "photographic" it looks versus artificial/rendered content. At the recommended threshold of 0.7, the dataset identifies ~500 million natural photographs suitable for vision research, cognitive science, and model training. Also… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural.imageimage-classification1B<n<10B5 likes2.3k downloads6mo agoHugging Face22naver-clova-ix /synthdog-ja Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-ja.image10K<n<100K5 likes2.3k downloads3y agoHugging Face23Fhrozen /openimages-narratives-v2 Open Images Narratives v2 Original Source | Google Localized Narrative 📌 Introduction This dataset comprises images and annotations from the original Open Images Dataset V7. Out of the 9M images, a subset of 1.9M images has been annotated with automatic methods (Image-text-to-text models). Description This dataset comprises all 1.9M images with bounding boxes annotations from the Open Images V7 project. Captions The annotations… See the full description on the dataset page: https://huggingface.co/datasets/Fhrozen/openimages-narratives-v2.imageimage-text-to-text1M<n<10M1 likes2.3k downloads10mo agoHugging Face24naver-clova-ix /synthdog-zh Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-zh.image10K<n<100K18 likes2.2k downloads3y agoHugging Face25nateraw /rendered-sst2 Rendered SST-2 The Rendered SST-2 Dataset from Open AI. Rendered SST2 is an image classification dataset used to evaluate the models capability on optical character recognition. This dataset was generated by rendering sentences in the Standford Sentiment Treebank v2 dataset. This dataset contains two classes (positive and negative) and is divided in three splits: a train split containing 6920 images (3610 positive and 3310 negative), a validation split containing 872 images (444… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/rendered-sst2.imageimage-classification1K<n<10K0 likes2.1k downloads4y agoHugging Face26BangumiBase /nanatsunotaizaimokushirokunoyonkishi Bangumi Image Base of Nanatsu No Taizai - Mokushiroku No Yonkishi This is the image base of bangumi Nanatsu no Taizai - Mokushiroku no Yonkishi, we detected 99 characters, 9289 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/nanatsunotaizaimokushirokunoyonkishi.image1K<n<10K0 likes1.9k downloads2y agoHugging Face27ibm-nasa-geospatial /hls_burn_scarsThis dataset contains Harmonized Landsat and Sentinel-2 imagery of burn scars and the associated masks for the years 2018-2021 over the contiguous United States. There are 804 512x512 scenes. Its primary purpose is for training geospatial machine learning models.image1K<n<10K27 likes1.8k downloads3y agoHugging Face28naufalso /cityscape-adverse Cityscape‑Adverse A benchmark for evaluating semantic segmentation robustness under realistic adverse conditions. Overview Cityscape‑Adverse extends the original Cityscapes dataset by applying eight realistic environmental modifications—rainy, foggy, spring, autumn, snowy, sunny, night, and dawn—using diffusion‑based image editing. All transformations preserve the original 2048×1024 semantic labels, enabling direct evaluation of model robustness in… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cityscape-adverse.image3 likes1.6k downloads1y agoHugging Face29nakasyou /captcha-like-suicaimagen<1K1 likes1.6k downloads2mo agoHugging Face30anony-008 /offroad-global-nav Offroad-global-nav Geospatial Dataset Overview This repository contains the dataset introduced in: “Learning Traversability-Aware Global Planners for Long Horizon Off-Road Navigation” The dataset is designed to support long-range off-road navigation using multi-modal geospatial data, combining large-scale overhead sensing with real-world human driving behavior. Unlike traditional off-road datasets that focus on local perception, this dataset enables global… See the full description on the dataset page: https://huggingface.co/datasets/anony-008/offroad-global-nav.geospatialroboticsn<1K2 likes1.6k downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.