datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
diffusion-pretrain-set-ft1
diffusion-pretrain-set-ft1
A multi-source image-caption pretraining dataset assembled from ten upstream
sources via a uniform ingest pipeline. Designed for a full pretrain or finetune
pipeline meant to curate for any major diffusion model preliminary, with the sole
intent to create a more powerful baseline preliminary train and a baseline
for synthesizing images to train the next generation of the VLM model.
This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.diffusion-pretrain-set-ft1-1024
diffusion-pretrain-set-ft1-1024
1024px (2x) upscale of AbstractPhil/diffusion-pretrain-set-ft1.
WARNING
MUCH OF THIS DATA WAS MODEL UPSCALED USING RAPID UPSCALERS.
THIS IS NOT CONSISTENTLY HIGH FIDELITY NOR IS IT EVEN CLOSE TO FAIR FIDELITY AT TIMES.
PLEASE use this ONLY for pretraining, new concepts, and simple design purposes ONLY. HEAVILY PRUNE FOR FINETUNING.
Thank you, good luck my friends.
Details
Model: realesr-general-x4v3 (SRVGG Compact… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1-1024.qwen-deepfashion-fused
qwen-deepfashion-fused
Every SFW row of AbstractPhil/qwen-deepfashion processed by the
qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption
structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs
(tasks_json) → FusedScene (fused_json: entities with mask-containment-owned
stratified attributes, relations with continuous offsets, counts, shared basin) →
deterministic fused prompt (prompt_fused).
Shards are strictly under… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-deepfashion-fused.qwen-synth-characters-fused
qwen-synth-characters-fused
Every SFW row of AbstractPhil/qwen-synth-characters processed by the
qwen-test-runner 12-process fused extraction system: age gate (strict) → 3×caption
structuring (Qwen3.5-9B, slot-registry schema) → 12 deterministic vision task JSONs
(tasks_json) → FusedScene (fused_json: entities with mask-containment-owned
stratified attributes, relations with continuous offsets, counts, shared basin) →
deterministic fused prompt (prompt_fused).
Shards are… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-fused.sdxl-qwen-phase0
SDXL–Qwen Phase-0 dataset
Purpose-built training set for AbstractPhil/geolip-sdxl-aleph.
Each row pairs a Qwen-Image-Lightning render with the caption that produced it and an
encoder-invariant geometric "aleph" address derived from the caption's bytes. It exists to
retrain SDXL (which stays the base model) around a new text encoder (Qwen in place of
CLIP-G) under a rectified-flow objective: the render is the flow-matching target, and the
student learns to reproduce it from the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase0.qwen-synth-characters
Qwen Synthetic Characters
A dataset of 60,847 fully synthetic (AI-generated) human portrait/character images produced
with Qwen-Image + the Qwen-Image-Lightning 4-step LoRA, with a prompt-augmentation policy
designed to give balanced demographics, diverse facial expressions, and varied attributes — and
to counter the base model's tendency to default to a narrow set of faces.
[!IMPORTANT]
These are not real people. Every image is generated by a diffusion model from a text… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters.flux-schnell-teacher-latents
Flux Schnell Teacher Latents
Pre-computed latents, decoded images, and text embeddings from FLUX.1-schnell for distillation and research.
Usage
from datasets import load_dataset
# Load specific subset
ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_512")
ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_2_512")
ds = load_dataset("AbstractPhil/flux-schnell-teacher-latents", "train_3_512")
Subsets
Config… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/flux-schnell-teacher-latents.qwen-deepfashion
Qwen DeepFashion (Real + Synthetic Full-Body Outfits)
A dataset of 160,015 fully synthetic (AI-generated) full-body fashion images produced with
Qwen-Image + the Qwen-Image-Lightning 4-step LoRA. Outfit descriptions come from two
sources — the real DeepFashion caption set and a synthetic outfit generator — and a shared
prompt-augmentation policy (fashion-v1) renders them as head-to-toe outfit photographs with
diversified wearers, backgrounds, poses, and framing while preserving… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-deepfashion.CN_pose3D_V10_512
CN_pose3D_V10 Processed
Processed version of tori29umai/CN_pose3D_V10
Progress
Processed: 66500/66500 images (100.0%)
Shards uploaded: 67
Processing:
Resized to 512x512 (LANCZOS)
Binary masks for white background removal
GPU-accelerated batch processing
Columns:
image: RGB (512x512)
conditioning_image: RGB pose (512x512)
mask: Binary (512x512) - 0=ignore white bg, 255=keep
text: Text prompt
Attribution
Original dataset:… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/CN_pose3D_V10_512.simple_abstract_circular_logossynthetic-object-relations
Synthetic Object Relations Dataset
A synthetic image dataset generated with Flux Schnell featuring clean object-relation prompts designed for training spatial reasoning in vision and diffusion models.
Dataset Description
This dataset contains images generated from structured prompts describing spatial relationships between objects. Unlike typical caption datasets that use free-form text, our prompts follow consistent patterns that explicitly encode:
Object identities… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-object-relations.AbstractImageNet100
AbstractImageNet100 v2 (Markdown)
Summary
AbstractImageNet100 v2 is a synthetic, validation-only image dataset derived
from the ImageNet100 validation split. It is a structured-Markdown prompt
iteration of the AbstractImageNet100 release.
Dataset scope
Split: validation only
Size: 5,000 images
Classes: 100 ImageNet classes
Images per class: 50
Pairing: each image is paired to an ImageNet100 validation source image
Contents and layout… See the full description on the dataset page: https://huggingface.co/datasets/ebykAI/AbstractImageNet100.synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.AbstractEdit
Abstract Image Editing Benchmark
A benchmark for evaluating instruction-following image editing models on abstract
(open-ended) vs explicit (fully specified) editing instructions. Context images are
drawn from Open Images V7
(validation split). Each item pairs an abstract edit instruction
(e.g. "Pack these houses into boxes for shipping") with an explicit counterpart that
lists every atomic change required.
The dataset exposes three configurations:
benchmark — context images… See the full description on the dataset page: https://huggingface.co/datasets/DucktorV/AbstractEdit.CN_pose3D_V7_512
CN_pose3D_V7 Processed
Processed version of tori29umai/CN_pose3D_V10
The tags are actually pretty terrible.
This is essentially a human body doll set that can be utilized in various forms of pretraining.
To make use of such a dataset you require many subsequent images, which the sd15 pretrain dataset has in plenty.
I advise training this unscaled rather than using 0.181 and ensuring everything is tagged with ai-generated and masked.
My first train will be testing it with timesteps… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/CN_pose3D_V7_512.kkcosmos_instagram-images-with-bad-captions-sd-scriptsirodkin_ffhq_with_llava_shorter_captionsimagenet-syntheticgeolip-sdxl-fid-scoringwikiart-genre-abstract-paintingtower-probes-resultsffhq_flux_latents_repairedqwen-synth-characters-100-json-test
qwen-synth-characters-100-json-test
A 100-row bench-test slice of AbstractPhil/qwen-synth-characters
processed end-to-end by the qwen-test-runner 12-process extraction system:
the 11 deterministic specialist vision tasks (tasks_json) plus caption→JSON-schema
structuring of all three prepared captions (struct_*). Built to measure wall-clock,
schema conformance, and grounding before scaling to the full 60,847-row set.
[!IMPORTANT]
These are not real people — every image is… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/qwen-synth-characters-100-json-test.tomytjandra_h-and-m-fashion-caption-12k-sd-scriptsflux-1s-generated-human-associations-10kGenerated With: Flux 1S
Size: 1024x1024
Steps: 4
Guidance: 0
Generated using random captions from a tokenization conjunction generation system - meant to introduce subject based linkers from one diffusion concept to another.
This needs additional AI passes to filter out things like floating hands, bad anatomy and so on.
This dataset is full of diffusion poison, but it's all generated from Flux 1S so it can be used for classification, object identification, association identification, and so on… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/flux-1s-generated-human-associations-10k.abstract-concepts-diffusiondb-imageslimingcv_LAION_Aesthetics_1024-sd-scripts-5000Each zip contains;
image.png/jpg/etc -> is the image
image.txt -> contains the caption
The captions are what I extracted from the json files at runtime. Hindsight says I should have kept the json but it is what it is for now.
I'll run a better one later. This one took quite a few hours as it was.
synthetic-object-relations-jsonabstractgarmentMiXaiLL76_TextOCR_OCR-sd-scripts
