CoolFace
Datasetpublic

Mitsua/art-museums-pd-440k

Art Museums PD 440K Summary This is a dataset to train text-to-image or any text and image multimodal models with minimized copyright/licensing concerns. All images and texts in this dataset are orignally shared under CC0 or public domain, and no pretrained models or any AI models are used to build this dataset except for our ElanMT model to translate English captions to Japanese. ElanMT model is trained solely on licensed corpus. Data sources… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/art-museums-pd-440k.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
17likes385downloads
Dataset Card

Art Museums PD 440K

[image]

Summary

This is a dataset to train text-to-image or any text and image multimodal models with minimized copyright/licensing concerns. All images and texts in this dataset are orignally shared under CC0 or public domain, and no pretrained models or any AI models are used to build this dataset except for our ElanMT model to translate English captions to Japanese. ElanMT model is trained solely on licensed corpus.

Data sources

Images and metadata collected from these museums open access. All images and metadata are shared under CC0 or Public Domain. We created image caption only from these metadata.

Filtering

  • —Word based pre-filtering is conducted for mitigating harmful or NSFW content.

Reporting Issues

We have taken great care to curate these datasets to ensure that no infringing content is included. However, if you review the contents of these datasets and identify any issues, please report them using this form. Upon receiving a report, our team will review the content, and if a violation is confirmed, we will remove the image from the dataset. Please note that we will not be able to reply to you regarding the status of the problem.

License

  • —The images and original metadata are licensed under CC0 by these museums.
  • —The dataset itself and compiled/translated captions are licensed under CC BY 4.0 by ELAN MITSUA Project / Abstract Engine.
  • —This means you can use, adapt and redistribute this dataset as long as you give appropriate credit.

Revisions

  • —Dec 15 2024 : Initial Commit
  • —Feb 13 2025 : v2
  • —Removed some inappropriate content