CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AbstractPhil /diffusion-pretrain-set-ft1 diffusion-pretrain-set-ft1 A multi-source image-caption pretraining dataset assembled from ten upstream sources via a uniform ingest pipeline. Designed for a full pretrain or finetune pipeline meant to curate for any major diffusion model preliminary, with the sole intent to create a more powerful baseline preliminary train and a baseline for synthesizing images to train the next generation of the VLM model. This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.image1M<n<10M2 likes1.8k downloads3mo agoHugging Face02AbstractPhil /diffusion-pretrain-set-ft1-1024 diffusion-pretrain-set-ft1-1024 1024px (2x) upscale of AbstractPhil/diffusion-pretrain-set-ft1. WARNING MUCH OF THIS DATA WAS MODEL UPSCALED USING RAPID UPSCALERS. THIS IS NOT CONSISTENTLY HIGH FIDELITY NOR IS IT EVEN CLOSE TO FAIR FIDELITY AT TIMES. PLEASE use this ONLY for pretraining, new concepts, and simple design purposes ONLY. HEAVILY PRUNE FOR FINETUNING. Thank you, good luck my friends. Details Model: realesr-general-x4v3 (SRVGG Compact… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1-1024.image1M<n<10M0 likes882 downloads3mo agoHugging Face03AbstractPhil /qwen-deepfashion-fused qwen-deepfashion-fused Every SFW row of AbstractPhil/qwen-deepfashion processed by the qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs (tasks_json) → FusedScene (fused_json: entities with mask-containment-owned stratified attributes, relations with continuous offsets, counts, shared basin) → deterministic fused prompt (prompt_fused). Shards are strictly under… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-deepfashion-fused.image100K<n<1M1 likes732 downloads2mo agoHugging Face04AbstractPhil /qwen-synth-characters-fused qwen-synth-characters-fused Every SFW row of AbstractPhil/qwen-synth-characters processed by the qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs (tasks_json) → FusedScene (fused_json: entities with mask-containment-owned stratified attributes, relations with continuous offsets, counts, shared basin) → deterministic fused prompt (prompt_fused). Shards are… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-fused.image10K<n<100K0 likes696 downloads2mo agoHugging Face05AbstractPhil /sdxl-qwen-phase0 SDXL–Qwen Phase-0 dataset Purpose-built training set for AbstractPhil/geolip-sdxl-aleph. Each row pairs a Qwen-Image-Lightning render with the caption that produced it and an encoder-invariant geometric "aleph" address derived from the caption's bytes. It exists to retrain SDXL (which stays the base model) around a new text encoder (Qwen in place of CLIP-G) under a rectified-flow objective: the render is the flow-matching target, and the student learns to reproduce it from the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase0.imagetext-to-image10K<n<100K3 likes642 downloads4mo agoHugging Face06AbstractPhil /qwen-synth-characters Qwen Synthetic Characters A dataset of 60,847 fully synthetic (AI-generated) human portrait/character images produced with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA, with a prompt-augmentation policy designed to give balanced demographics, diverse facial expressions, and varied attributes — and to counter the base model's tendency to default to a narrow set of faces. [!IMPORTANT] These are not real people. Every image is generated by a diffusion model from a text… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters.imagetext-to-image10K<n<100K0 likes637 downloads3mo agoHugging Face07AbstractPhil /flux-schnell-teacher-latents Flux Schnell Teacher Latents Pre-computed latents, decoded images, and text embeddings from FLUX.1-schnell for distillation and research. Usage from datasets import load_dataset # Load specific subset ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_512") ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_2_512") ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_3_512") Subsets Config… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/flux-schnell-teacher-latents.imageimage-to-image100K<n<1M0 likes477 downloads8mo agoHugging Face08AbstractPhil /qwen-deepfashion Qwen DeepFashion (Real + Synthetic Full-Body Outfits) A dataset of 160,015 fully synthetic (AI-generated) full-body fashion images produced with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA. Outfit descriptions come from two sources — the real DeepFashion caption set and a synthetic outfit generator — and a shared prompt-augmentation policy (fashion-v1) renders them as head-to-toe outfit photographs with diversified wearers, backgrounds, poses, and framing while preserving… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-deepfashion.imagetext-to-image100K<n<1M0 likes444 downloads3mo agoHugging Face09AbstractPhil /CN_pose3D_V10_512 CN_pose3D_V10 Processed Processed version of tori29umai/CN_pose3D_V10 Progress Processed: 66500/66500 images (100.0%) Shards uploaded: 67 Processing: Resized to 512x512 (LANCZOS) Binary masks for white background removal GPU-accelerated batch processing Columns: image: RGB (512x512) conditioning_image: RGB pose (512x512) mask: Binary (512x512) - 0=ignore white bg, 255=keep text: Text prompt Attribution Original dataset:… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/CN_pose3D_V10_512.image10K<n<100K0 likes380 downloads4mo agoHugging Face10superchthonic /simple_abstract_circular_logosimagen<1K0 likes350 downloads4y agoHugging Face11AbstractPhil /synthetic-object-relations Synthetic Object Relations Dataset A synthetic image dataset generated with Flux Schnell featuring clean object-relation prompts designed for training spatial reasoning in vision and diffusion models. Dataset Description This dataset contains images generated from structured prompts describing spatial relationships between objects. Unlike typical caption datasets that use free-form text, our prompts follow consistent patterns that explicitly encode: Object identities… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-object-relations.imagetext-to-image100K<n<1M2 likes232 downloads8mo agoHugging Face12ebykAI /AbstractImageNet100 AbstractImageNet100 v2 (Markdown) Summary AbstractImageNet100 v2 is a synthetic, validation-only image dataset derived from the ImageNet100 validation split. It is a structured-Markdown prompt iteration of the AbstractImageNet100 release. Dataset scope Split: validation only Size: 5,000 images Classes: 100 ImageNet classes Images per class: 50 Pairing: each image is paired to an ImageNet100 validation source image Contents and layout… See the full description on the dataset page: https://huggingface.co/datasets/ebykAI/AbstractImageNet100.image1K<n<10K1 likes232 downloads1mo agoHugging Face13AbstractPhil /synthetic-characters Synthetic Characters Dataset A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models. Recommended Filters: Age People Count Hair color Camera angle Anime/Realistic Nudity/Clothed I'll likely recaption everything with a list of classifications attached to them for easy filtering later. Dataset Description This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.imagetext-to-image100K<n<1M0 likes186 downloads4mo agoHugging Face14DucktorV /AbstractEdit Abstract Image Editing Benchmark A benchmark for evaluating instruction-following image editing models on abstract (open-ended) vs explicit (fully specified) editing instructions. Context images are drawn from Open Images V7 (validation split). Each item pairs an abstract edit instruction (e.g. "Pack these houses into boxes for shipping") with an explicit counterpart that lists every atomic change required. The dataset exposes three configurations: benchmark — context images… See the full description on the dataset page: https://huggingface.co/datasets/DucktorV/AbstractEdit.imageimage-to-image10K<n<100K0 likes154 downloads5mo agoHugging Face15AbstractPhil /CN_pose3D_V7_512 CN_pose3D_V7 Processed Processed version of tori29umai/CN_pose3D_V10 The tags are actually pretty terrible. This is essentially a human body doll set that can be utilized in various forms of pretraining. To make use of such a dataset you require many subsequent images, which the sd15 pretrain dataset has in plenty. I advise training this unscaled rather than using 0.181 and ensuring everything is tagged with ai-generated and masked. My first train will be testing it with timesteps… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/CN_pose3D_V7_512.image100K<n<1M0 likes131 downloads4mo agoHugging Face16AbstractPhil /kkcosmos_instagram-images-with-bad-captions-sd-scriptsimage10K<n<100K1 likes115 downloads1y agoHugging Face17AbstractPhil /irodkin_ffhq_with_llava_shorter_captionsimage10K<n<100K0 likes80 downloads1y agoHugging Face18AbstractPhil /imagenet-syntheticimage10K<n<100K0 likes74 downloads8mo agoHugging Face19AbstractPhil /geolip-sdxl-fid-scoringimage1K<n<10K0 likes70 downloads4mo agoHugging Face20adrianrm /wikiart-genre-abstract-paintingimage1K<n<10K0 likes55 downloads5mo agoHugging Face21AbstractPhil /tower-probes-resultsimagen<1K0 likes53 downloads2mo agoHugging Face22AbstractPhil /ffhq_flux_latents_repairedimage10K<n<100K0 likes46 downloads4mo agoHugging Face23AbstractPhil /qwen-synth-characters-100-json-test qwen-synth-characters-100-json-test A 100-row bench-test slice of AbstractPhil/qwen-synth-characters processed end-to-end by the qwen-test-runner 12-process extraction system: the 11 deterministic specialist vision tasks (tasks_json) plus caption→JSON-schema structuring of all three prepared captions (struct_*). Built to measure wall-clock, schema conformance, and grounding before scaling to the full 60,847-row set. [!IMPORTANT] These are not real people — every image is… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-100-json-test.image1K<n<10K0 likes45 downloads2mo agoHugging Face24AbstractPhil /tomytjandra_h-and-m-fashion-caption-12k-sd-scriptsimage10K<n<100K0 likes44 downloads1y agoHugging Face25AbstractPhil /flux-1s-generated-human-associations-10kGenerated With: Flux 1S Size: 1024x1024 Steps: 4 Guidance: 0 Generated using random captions from a tokenization conjunction generation system - meant to introduce subject based linkers from one diffusion concept to another. This needs additional AI passes to filter out things like floating hands, bad anatomy and so on. This dataset is full of diffusion poison, but it's all generated from Flux 1S so it can be used for classification, object identification, association identification, and so on… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/flux-1s-generated-human-associations-10k.imagetext-to-image0 likes39 downloads1y agoHugging Face26nirmalendu01 /abstract-concepts-diffusiondb-imagesimage1K<n<10K0 likes39 downloads7mo agoHugging Face27AbstractPhil /limingcv_LAION_Aesthetics_1024-sd-scripts-5000Each zip contains; image.png/jpg/etc -> is the image image.txt -> contains the caption The captions are what I extracted from the json files at runtime. Hindsight says I should have kept the json but it is what it is for now. I'll run a better one later. This one took quite a few hours as it was. image1M<n<10M1 likes37 downloads1y agoHugging Face28AbstractPhil /synthetic-object-relations-jsonimage1K<n<10K0 likes35 downloads4mo agoHugging Face29histin116 /abstractgarmentimagen<1K1 likes32 downloads2y agoHugging Face30AbstractPhil /MiXaiLL76_TextOCR_OCR-sd-scriptsimage100K<n<1M0 likes31 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.