datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-to-image-diffusiondb-2M
DiffusionDB text-to-image subset
A cleaned, safety-filtered image-prompt dataset for training a text-to-image
model, built from DiffusionDB.
Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers
part_id 1-20 (20,000 source images) before filtering. The same content is
also kept on the 20k-subset branch.
Load it with:
load_dataset("whosouravsharma/text-to-image-diffusiondb-2M")
Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.diffusion-generated-text
Diffusion-Generated Text
This dataset contains 16,791 question-response pairs generated by a diffusion language model. It is released to support research on diffusion-generated language and machine-generated text detection.
Dataset schema
Column
Type
Description
question
string
Input question or prompt.
dLLM_response
string
Response produced by the diffusion language model.
Loading
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/paoche11/diffusion-generated-text.diffusion.4.text_to_image
Dataset Card for "diffusion.4.text_to_image"
More Information needed
diffusion.4.text_to_image.book
Dataset Card for "diffusion.4.text_to_image.book"
More Information needed
5508nailset_diffusion.4.text_to_image
Dataset Card for "5508nailset_diffusion.4.text_to_image"
More Information needed
