stzhao/AnyWord-3M
Dataset from AnyText: Multilingual Visual Text Generation And Editing. Dataset description from Anytext Team: Currently, there is a relative scarcity of public datasets for text generation tasks, especially those involving non-Latin script languages. To address this, we introduce a large-scale multilingual dataset called AnyWord-3M. The images in this dataset are sourced from Noah-Wukong, LAION-400M, and OCR recognition datasets such as ArT, COCO-Text, RCTW, LSVT, MLT, MTWI, ReCTS, etc. These… See the full description on the dataset page: https://huggingface.co/datasets/stzhao/AnyWord-3M.
1710k
../
train_1.parquetdownload
train_10.parquetdownload
train_11.parquetdownload
train_12.parquetdownload
train_13.parquetdownload
train_14.parquetdownload
train_15.parquetdownload
train_16.parquetdownload
train_2.parquetdownload
train_3.parquetdownload
train_4.parquetdownload
train_5.parquetdownload
train_6.parquetdownload
train_7.parquetdownload
train_8.parquetdownload
train_9.parquetdownload
