CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /dummy_image_text_data Dataset Card for "dummy_image_text_data" More Information needed imagen<1K1 likes11k downloads4y agoHugging Face02data-is-better-together /open-image-preferences-v1 Open Image Preferences Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K. Image 1 Image 2 Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed. Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.imagetext-to-image1K<n<10K31 likes5k downloads2y agoHugging Face03Carzit /SukaSuka-image-dataset 该数据集包含了《末日时在做什么?有没有空?可以来拯救吗?》大部分主要角色角色的图像数据,来源为动漫截图与同人二创。 为方便LoRA模型训练,所有图片尺寸均截为512x640尺寸,相应打标主要由Waifu Diffusion 1.4 Tagger V2自动完成,部分手工调整。 欢迎提交PR补充或修正本数据集! Alpha:8.21号之后的clone都是放大了两倍的图片,这是为了sdxl做准备,如果你还需要512*640尺寸的数据集,你可以在clone之后,执行下面的命令 git checkout 183e253c4c304fc6c5ef5046f1940712c349c94e 相关数据集的更正作业正在火热的进行中,请期待继续的更新吧~ imagen<1K5 likes3.1k downloads1y agoHugging Face04Rajarshi-Roy-research /Defactify_Image_Dataset Defactify_Image_Dataset This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection. 📝 Dataset Description Dataset Summary The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.imageimage-classification10K<n<100K22 likes2.4k downloads4mo agoHugging Face05HabibaAbderrahim /Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset Description This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations. It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.imagetranslationn<1K0 likes1.9k downloads1y agoHugging Face06irene93 /malfunction-image-datasetimage10K<n<100K0 likes1.5k downloads1y agoHugging Face07svjack /Chinese_Children_Image_Captioning_Dataset_Split0 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.image1K<n<10K0 likes862 downloads2y agoHugging Face08nojiyoon /pagoda-text-and-image-dataset Dataset Card for "pagoda-text-and-image-dataset" More Information needed imagen<1K1 likes527 downloads3y agoHugging Face09vargr /yt_full_image_dataset Dataset Card for "yt_full_image_dataset" More Information needed image100K<n<1M1 likes519 downloads3y agoHugging Face10svjack /Chinese_Children_Image_Captioning_Dataset_Split1 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split1.image1K<n<10K0 likes367 downloads2y agoHugging Face11Shanmuk4622 /ai-image-detection-datasetgated AI-Image Detection Dataset Paired real / AI images, with shared image-grounded captions, for training and evaluating AI-generated-image detectors. Each of 10,000 real photos is captioned once (BLIP-2) and paired with one synthetic partner per generator (6 generators → 60,000 AI images). A real image and all of its AI partners share the same prompt, so the only systematic difference between the classes is the generative process itself. A detector trained here is pushed toward the… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-image-detection-dataset.imageimage-classification100K<n<1M4 likes355 downloads3mo agoHugging Face12Aayush672 /fashion-image-datasetimage10K<n<100K0 likes352 downloads1y agoHugging Face13datapointai /text-2-image-human-preferences-2mgated Text-to-image human preferences: 2M votes across 30 models This dataset contains the complete voting record behind the Datapoint Image Bench leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of 216,116 image pairs. The votes compare 30 text-to-image models in a complete round-robin on 500 prompts, judged by annotators from over 200 countries. Every vote includes the annotator's trust score at the time the vote was cast. Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.imagetext-to-image1M<n<10M21 likes319 downloads1mo agoHugging Face14data-is-better-together /open-image-preferences-v1-binarized Open Image Preferences Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K. Image 1 Image 2 Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed. Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1-binarized.image1K<n<10K59 likes309 downloads2y agoHugging Face15vargr /yt_main_image_dataset Dataset Card for "yt_main_image_dataset" More Information needed image100K<n<1M0 likes303 downloads3y agoHugging Face16zwa73 /SoulTide-ImageData-Dataset 目录结构 character/ └── [char]/ ├── resource/ | 原始资源 │ ├── rotate/ | 旋转变体资源 │ ├── alpha_bg/ | 透明背景的资源 │ ├── white_bg/ | 从透明背景添加黑色背景的资源 │ ├── black_bg/ | 从透明背景添加白色背景的资源 │ ├── unused/ | 不使用的资源 │ └── ready/ | 准备使用但还未分类的资源 ├── categorized/ | 分类完成的资源 ├── processed/ | 完成预处理的资源 └── training_set/ | 以 [训练次数_概念] 命名的训练集 工具 搭配此管理器来生成所需的训练集:https://github.com/Sosarciel/SoulTide-ImageData-Manager image1K<n<10K0 likes249 downloads1mo agoHugging Face17Kanrawee /my-image-caption-datasetimage100K<n<1M1 likes236 downloads2y agoHugging Face18theojiang /image-text-dataset-subset-300k-captions_onlyimage100K<n<1M1 likes234 downloads3y agoHugging Face19jiuhai /image_dataimage10M<n<100M0 likes234 downloads1y agoHugging Face20Agents-X /PyVision-Image-SFT-Data PyVision-Image-RL-Data Project Page | Paper | GitHub This repository contains the reinforcement learning (RL) data used to train PyVision-Image-RL, as presented in the paper "PyVision-RL: Forging Open Agentic Vision Models via RL". PyVision-RL is a reinforcement learning framework for open-weight multimodal models designed to stabilize training and sustain interaction, preventing interaction collapse and encouraging multi-turn tool use in agentic tasks. Citation… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/PyVision-Image-SFT-Data.textimage-text-to-text1K<n<10K3 likes224 downloads7mo agoHugging Face21tarn59 /character_turnaround_sheet_qwen_image_edit_2509_datasetBase images were generated by Qwen Image and I used Wan to do 360 degree rotation. I then took frames from the rotation and concatenated them together using imagemagick. imagen<1K5 likes211 downloads10mo agoHugging Face22AhmedSSabir /Textual-Image-Caption-Dataset Update: OCT-2023 Add v2 with recent SoTA model swinV2 classifier for both soft/hard-label visual_caption_cosine_score_v2 with person label (0.2, 0.3 and 0.4) Introduction Modern image captaining relies heavily on extracting knowledge, from images such as objects, to capture the concept of static story in the image. In this paper, we propose a textual visual context dataset for captioning, where the publicly available dataset COCO caption (Lin et al., 2014) has been… See the full description on the dataset page: https://huggingface.co/datasets/AhmedSSabir/Textual-Image-Caption-Dataset.textimage-to-text7 likes207 downloads1y agoHugging Face23Wauplin /example-space-to-dataset-imageDemo to save data from a Space to a Dataset. Goal is to provide reusable snippets of code. Documentation: https://huggingface.co/docs/huggingface_hub/main/en/guides/upload#scheduled-uploads Space: https://huggingface.co/spaces/Wauplin/space_to_dataset_saver/ JSON dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-json Image dataset: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image Image (zipped) dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Wauplin/example-space-to-dataset-image.imagen<1K1 likes204 downloads1y agoHugging Face24BEE-spoke-data /govdocs1-image BEE-spoke-data/govdocs1-image This contains .jpg files from govdocs1. Light deduplication was applied (i.e. jdupes on all files) which removed ~500 duplicate images. DatasetDict({ train: Dataset({ features: ['image'], num_rows: 108895 }) }) source Source info/page: https://digitalcorpora.org/corpora/file-corpora/files/ @inproceedings{garfinkel2009bringing, title={Bringing Science to Digital Forensics with Standardized Forensic Corpora}… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-image.image100K<n<1M0 likes180 downloads9mo agoHugging Face25ductai199x /image-manipulation-dataset-compilationimage10K<n<100K0 likes179 downloads2y agoHugging Face26nojiyoon /pagoda-text-and-image-dataset-small Dataset Card for "pagoda-text-and-image-dataset-small" More Information needed imagen<1K0 likes171 downloads3y agoHugging Face27dwb2023 /brain-tumor-image-dataset-semantic-segmentation Dataset Card for "brain-tumor-image-dataset-semantic-segmentation" Dataset Description The Brain Tumor Image Dataset (BTID) for Semantic Segmentation contains MRI images and annotations aimed at training and evaluating segmentation models. This dataset was sourced from Kaggle and includes detailed segmentation masks indicating the presence and boundaries of brain tumors. This dataset can be used for developing and benchmarking algorithms for medical image segmentation… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/brain-tumor-image-dataset-semantic-segmentation.image1K<n<10K6 likes156 downloads2y agoHugging Face28ozertuu /turkish-image-description-dataset-shard-42 Turkish Image Description Dataset - Shard 42 This dataset contains translated image descriptions from English to Turkish. Contents Images with their Turkish and original English descriptions How to use from datasets import load_dataset # Load the dataset dataset = load_dataset("ozertuu/turkish-image-description-dataset-shard-42") # Access data for item in dataset["train"]: image = item["image"] # PIL.Image object turkish_description =… See the full description on the dataset page: https://huggingface.co/datasets/ozertuu/turkish-image-description-dataset-shard-42.image10K<n<100K0 likes155 downloads1y agoHugging Face29ozertuu /turkish-image-description-dataset-shard-17 Turkish Image Description Dataset - Shard 17 This dataset contains translated image descriptions from English to Turkish. Contents Images with their Turkish and original English descriptions How to use from datasets import load_dataset # Load the dataset dataset = load_dataset("ozertuu/turkish-image-description-dataset-shard-17") # Access data for item in dataset["train"]: image = item["image"] # PIL.Image object turkish_description =… See the full description on the dataset page: https://huggingface.co/datasets/ozertuu/turkish-image-description-dataset-shard-17.image10K<n<100K0 likes154 downloads1y agoHugging Face30crossingminds /shopping-queries-image-dataset Shopping Queries Image Dataset (SQID 🦑): An Image-Enriched ESCI Dataset for Exploring Multimodal Learning in Product Search Introduction The Shopping Queries Image Dataset (SQID) is a dataset that includes image information for over 190,000 products. This dataset is an augmented version of the Amazon Shopping Queries Dataset, which includes a large number of product search queries from real Amazon users, along with a list of up to 40 potentially relevant results and… See the full description on the dataset page: https://huggingface.co/datasets/crossingminds/shopping-queries-image-dataset.image100K<n<1M7 likes153 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.