CoolFace
Datasetpublic

stzhao/AnyWord-3M

Dataset from AnyText: Multilingual Visual Text Generation And Editing. Dataset description from Anytext Team: Currently, there is a relative scarcity of public datasets for text generation tasks, especially those involving non-Latin script languages. To address this, we introduce a large-scale multilingual dataset called AnyWord-3M. The images in this dataset are sourced from Noah-Wukong, LAION-400M, and OCR recognition datasets such as ArT, COCO-Text, RCTW, LSVT, MLT, MTWI, ReCTS, etc. These… See the full description on the dataset page: https://huggingface.co/datasets/stzhao/AnyWord-3M.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
17likes10kdownloads
../
filetrain_1.parquet518.6 MBdownload
filetrain_10.parquet516.0 MBdownload
filetrain_11.parquet512.6 MBdownload
filetrain_12.parquet517.1 MBdownload
filetrain_13.parquet516.5 MBdownload
filetrain_14.parquet518.8 MBdownload
filetrain_15.parquet514.5 MBdownload
filetrain_16.parquet517.9 MBdownload
filetrain_2.parquet516.4 MBdownload
filetrain_3.parquet515.2 MBdownload
filetrain_4.parquet518.0 MBdownload
filetrain_5.parquet522.2 MBdownload
filetrain_6.parquet514.9 MBdownload
filetrain_7.parquet517.4 MBdownload
filetrain_8.parquet517.2 MBdownload
filetrain_9.parquet520.0 MBdownload

stzhao/AnyWord-3M · main · files are served by the source, never re-hosted here