CoolFace
Datasetpublic

speed/WAON

WAON: Large-Scale and High-Quality Japanese Image-Text Pair Dataset for Vision-Language Models | 🤗 HuggingFace  | 📄 Paper  | 🧑‍💻 Code  | Introduction WAON is a Japanese (image, text) pair dataset containing approximately 155M examples, crawled from Common Crawl. It is built from snapshots taken in 2025-18, 2025-08, 2024-51, 2024-42, 2024-33, and 2024-26. The dataset is high-quality and diverse, constructed through a sophisticated… See the full description on the dataset page: https://huggingface.co/datasets/speed/WAON.

sourceHugging Faceapache-2.0updated 11mo agoView on Hugging Face
2likes318downloads

speed/WAON · main · files are served by the source, never re-hosted here