CoolFace
Datasetpublic

agentlans/stock-photos-asian-people

Stock Photos (Asian People, Stable Diffusion 1.5) Collage of randomly selected images from the dataset Collection of synthetic stock photographs created with Stable Diffusion 1.5, emphasizing Asian people (including East Asians, Southeast Asians, South Asians, and Central Asians). Images were generated using a diverse set of prompts and filtered for quality, realism, and safety. Supported Tasks Image generation Image captioning Visual representation… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/stock-photos-asian-people.

sourceHugging Facecc0-1.0updated 1y agoView on Hugging Face
1likes91downloads
Dataset Card

Stock Photos (Asian People, Stable Diffusion 1.5)

<figure> <img src="Examples.jpg" alt="Collage of randomly selected images from the dataset" width="600"/> <figcaption>Collage of randomly selected images from the dataset</figcaption> </figure>

Collection of synthetic stock photographs created with Stable Diffusion 1.5, emphasizing Asian people (including East Asians, Southeast Asians, South Asians, and Central Asians). Images were generated using a diverse set of prompts and filtered for quality, realism, and safety.

Supported Tasks

  • —Image generation
  • —Image captioning
  • —Visual representation learning
  • —Fairness and diversity analysis in generative models

Dataset Structure

Each entry in the dataset includes:

  • —The generated image (512 &times; 512 RGB)
  • —The automatically generated caption

The metadata of each image file contains its generation parameters. Note that the input prompts may not resemble the generated images and their captions.

Dataset Creation

Generation Process

  • —Software: AUTOMATIC1111/stable-diffusion-webui v1.10.1 (Python 3.10.12, Torch 2.1.2+cu121, xformers 0.0.23.post1)
  • —Models used:
  • —leosamsHelloworldXL_filmGrain20
  • —redcraftCADSUpdatedApr28_relustion15FP32
  • —evalisenniaSD15Ultra_v10
  • —asianRealismByStable_v4
  • —Prompts: Each image was generated using a random sample of 10 phrases from the first 100&thinsp;000 rows of agentlans/noun-phrases.
  • —Negative prompt: CyberRealistic_Negative-neg
  • —Sampling: DPM++ 2M, 20 steps, CFG Scale 7
  • —Additional processing: ADetailer with person, face, and hand detection models

Filtering

  • —Graphics and NSFW images removed using:
  • —timm/tinyvit21m224.distin22kftin1k (for graphics: drawing, cartoon, infographic, webpage, illustration, comic)
  • —Marqo/nsfw-image-detection-384 (for NSFW)
  • —Outlier removal with isolation forest (contamination=0.1) using:
  • —resnet18 embeddings
  • —CLIP-based detection (openai/clip-vit-large-patch14) of problematic body parts (repeated 3 times)

Captioning

  • —Captions generated with Qwen/Qwen2.5-VL-3B-Instruct
  • —System prompt: "You are an expert image captioner."
  • —User prompt: "Write a concise, clear caption describing the picture."

Languages

Captions are in English. Prompts are in English and consist of diverse noun phrases.

Limitations

  • —Some images exhibit artifacts, unrealistic features, or visual distortions.
  • —The dataset includes individuals of various ethnic backgrounds, not exclusively Asians.
  • —Contains identifiable individuals, including celebrities and politicians.
  • —Due to prompt limitations, the dataset does not fully capture the diversity of Asian cultures, art forms, or traditional contexts.
  • —Captions are automatically generated and may not be appropriate for all use cases.

Licence

Creative Commons Zero 1.0