CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01timm /imagenet-1k-wdsgated Dataset Summary ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated. 💡… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-1k-wds.imageimage-classification10K<n<100K35 likes5k downloads3y agoHugging Face02imageomics /TreeOfLife-10M Dataset Card for TreeOfLife-10M Dataset Summary With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M.documentimage-classification1M<n<10M50 likes4.2k downloads8mo agoHugging Face03timm /imagenet-12k-wdsgated Dataset Summary This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm. The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-12k-wds.imageimage-classification100K<n<1M10 likes1.6k downloads3y agoHugging Face04dark-xet /imagenet-1k-wds Dataset Summary ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated. 💡… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/imagenet-1k-wds.imageimage-classification1M<n<10M0 likes614 downloads1y agoHugging Face056DammK9 /danbooru2024-latents-sdxl-1ktar Danbooru 2024 SDXL VAE latents in 1k tar Dedicated dataset to align deepghs/danbooru2024-webp-4Mpixel. "4MP-Focus" for average raw image resolution. Latents are ARB with maximum size of 1024x1024 as the recommended setting in kohyas. Major reason is to make sure I can finetune with RTX 3090. VRAM usage will raise drastically after 1024. Generated from prepare_buckets_latents_v2.py, modified from prepare_buckets_latents.py. Used for kohya-ss/sd-scripts. In theory it may replace… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/danbooru2024-latents-sdxl-1ktar.textimage-classification1M<n<10M7 likes528 downloads7mo agoHugging Face06Shu1L0n9 /CleanSTL-10 Dataset Card for STL-10 Cleaned (Deduplicated Training Set) Paper | Code Dataset Description This dataset is a modified version of the STL-10 dataset. The primary modification involves deduplicating the training set by removing any images that are exact byte-for-byte matches (based on SHA256 hash) with images present in the original STL-10 test set. The dataset comprises this cleaned training set and the original, unmodified STL-10 test set. The goal is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/Shu1L0n9/CleanSTL-10.imageimage-classification100K<n<1M2 likes342 downloads1y agoHugging Face07ek826 /imagenet-gen-sd1.5 ImageNet Generated using Stable Diffusion v1.5 The following repository mimics the size and class structure of the original ImageNet database. The classes can be found in the classes.txt file. This dataset contains approximately 1300 images per class over 1000 classes for a total of 1.3 million images. Here is an excerpt from classes.txt: 0 tench, Tinca tinca 1 goldfish, Carassius auratus 2 great white shark, white shark, man-eater, man-eating shark, Carcharodon caharias 3 tiger… See the full description on the dataset page: https://huggingface.co/datasets/ek826/imagenet-gen-sd1.5.imageimage-classification1M<n<10M5 likes223 downloads3y agoHugging Face08ScarlettChan /TreeOfLife-10M Dataset Card for TreeOfLife-10M Dataset Summary With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/ScarlettChan/TreeOfLife-10M.documentimage-classification1M<n<10M0 likes210 downloads8mo agoHugging Face09Ethicalpirate91 /TreeOfLife-10M Dataset Card for TreeOfLife-10M Dataset Summary With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/Ethicalpirate91/TreeOfLife-10M.documentimage-classification1M<n<10M0 likes51 downloads7mo agoHugging Face10asdfeWRF /TreeOfLife-10M Dataset Card for TreeOfLife-10M Dataset Summary With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/asdfeWRF/TreeOfLife-10M.documentimage-classification1M<n<10M0 likes50 downloads4mo agoHugging Face11Metavolve-Labs /alexandria-aeternum-1k Alexandria Aeternum — Genesis Your Entry Point to Cognitive Nutrition 1,000 curated paintings · Masters only · 4,000+ tokens each · Free sample Not scraped. Not auto-captioned. Translated from human knowledge. Monet, Van Gogh, Rembrandt, Degas, Hokusai, Cezanne, and 100+ master artists. Full 10K Dataset · MCP Access (2M+ Artworks) · Research Paper · Explore Full Archive · Scale With Us MCP Access — AI Agent Marketplace The complete high-resolution… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/alexandria-aeternum-1k.imagetext-to-image1K<n<10K1 likes38 downloads7mo agoHugging Face12lingcarzy /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M0 likes36 downloads6mo agoHugging Face13LYC-0291 /TreeOfLife-10M Dataset Card for TreeOfLife-10M Dataset Summary With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/LYC-0291/TreeOfLife-10M.documentimage-classification1M<n<10M0 likes28 downloads5mo agoHugging Face146DammK9 /e621_2024-latents-sdxl-1ktar E621 2024 SDXL VAE latents in 1k tar Dedicated dataset to align both NebulaeWis/e621-2024-webp-4Mpixel and deepghs/e621_newest-webp-4Mpixel. "4MP-Focus" for average raw image resolution. Latents are ARB with maximum size of 1024x1024 as the recommended setting in kohyas. Major reason is to make sure I can finetune with RTX 3090. VRAM usage will raise drastically after 1024. Generated from prepare_buckets_latents_v2.py, modified from prepare_buckets_latents.py. Used for… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/e621_2024-latents-sdxl-1ktar.textimage-classification1M<n<10M1 likes27 downloads2y agoHugging Face15alexey-zhavoronkin /CINIC10CINIC10 dataset with interface of CIFAR10. It is faster than the common CINIC10 due to the fact that all images are loaded into RAM while initing dataset instance. You should save cinic10.py from this repo in local directory. And then import the CINIC10 class from it: import torchvision import torch from torchvision import transforms from cinic10 import CINIC10 data_mean = [0.47889522, 0.47227842, 0.43047404] data_std = [0.24205776, 0.23828046, 0.25874835] transform_train =… See the full description on the dataset page: https://huggingface.co/datasets/alexey-zhavoronkin/CINIC10.textimage-classificationn<1K0 likes25 downloads2y agoHugging Face16mpatrick1991 /srtm-3-arc-second-global SRTM 3 Arc-Second Global Raw ASCII heightmaps of the Earth's surface labelled according to latitude and longitude. Mission Description The Shuttle Radar Topography Mission (SRTM) was flown aboard the space shuttle Endeavour February 11-22, 2000. The National Aeronautics and Space Administration (NASA) and the National Geospatial-Intelligence Agency (NGA) participated in an international project to acquire radar data which were used to create the first near-global set of… See the full description on the dataset page: https://huggingface.co/datasets/mpatrick1991/srtm-3-arc-second-global.textunconditional-image-generation10K<n<100K0 likes18 downloads9mo agoHugging Face176DammK9 /e621_2024-captions-1ktar E621 2024 captions only in 1k tar Raw captions jointed from lodestones/e621-captions It doesn't align to any dataset yet. meta_cap.json has been provided in compressed format if you want to train with kohyas triner. Currently I'm trying to merge this with my 2024 version. Core logic The script building this 1ktar There is not much choice, I don't have GPU to run for 1M captions with VLM so I just "take it or leave it". rearranged_tags = [row.regular_summary… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/e621_2024-captions-1ktar.textimage-classification1M<n<10M0 likes17 downloads2y agoHugging Face18telecomadm1145 /danbooru-convnext-embeddings2gated Dataset Card for Danbooru ConvNeXt Embeddings 2 Danbooru ConvNeXt 向量数据集 2 Dataset Details / 数据集详情 Dataset Description / 数据集描述 English: This dataset contains approximately 5,312,000 image embeddings (vectors). It was generated by extracting features from the massive Danbooru anime image dataset using the convnext_large.dinov3_lvd1689m computer vision model. These embeddings represent the visual features of the images in a high-dimensional space… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/danbooru-convnext-embeddings2.textimage-classification1M<n<10M1 likes8 downloads9mo agoHugging Face19just-a-try /anime-classification-v1.5gated Anime Image Classification Dataset (v1.5) This is the webdataset dataset, containing 323060 images in total. Images here are resized to min(width, height) <= 640. How to Use It from datasets import load_dataset dataset = load_dataset('just-a-try/anime-classification-v1.5') print(dataset["train"][0]) Images 323060 images in total. Split Image Count Total Size train 257996 14.5 GB test 32506 1.83 GB val 32558 1.84 GB Class Image Count… See the full description on the dataset page: https://huggingface.co/datasets/just-a-try/anime-classification-v1.5.imageimage-classification100K<n<1M1 likes4 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.