poloclub/diffusiondb
DiffusionDB is the first large-scale text-to-image prompt dataset. It contains 2 million images generated by Stable Diffusion using prompts and hyperparameters specified by real users. The unprecedented scale and diversity of this human-actuated dataset provide exciting research opportunities in understanding the interplay between prompts and generative models, detecting deepfakes, and designing human-AI interaction tools to help users more easily use these models.
Add a new subset for the data viewer
Fix NaT issue when reading timestamp
Fix 2m subset part-000989.zip security issue
Change hf_hub_url() call
Rezip 2m-989 to test safetly scan
Add nsfw score distribution
Remove space in split names
Change pyarrow to pandas for dataset preview
Update loading script
Update reademe
Fix 7 error images and update DB 2M metadata
Add metadata for DiffusionDB Large
Add DiffusionDB Large
Update README
Fix a broken image in part-000689 (#1)
Document data loading method
Fix text_only key error
Force rebuild dataset preview
Add a new text_only config option
Merge two config options
Add a loading script
Update README
Add paper link
Add a teaser
Update readme
Update readme
Update readme
Update metadata
Add metadata table
Add up to 2000
Add up to 1740
Add up to 1664
Add up to parts 1497
Add first 100 parts
Init commit
initial commit
