CoolFace
14 results

synthdog

naver-clova-ix /synthdog-en Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-en.image10K<n<100K27 likes4.7k downloads3y agoHugging Facenaver-clova-ix /synthdog-ko Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-ko.image10K<n<100K18 likes3.2k downloads3y agoHugging Facenaver-clova-ix /synthdog-ja Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-ja.image10K<n<100K5 likes2.3k downloads3y agoHugging Facenaver-clova-ix /synthdog-zh Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets For more information, please visit https://github.com/clovaai/donut The links to the SynthDoG-generated datasets are here: synthdog-en: English, 0.5M. synthdog-zh: Chinese, 0.5M. synthdog-ja: Japanese, 0.5M. synthdog-ko: Korean, 0.5M. To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details. How to Cite If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-zh.image10K<n<100K18 likes2.2k downloads3y agoHugging FaceAkajackson /donut_synthdog_rus Dataset Card for "donut_rus" More Information needed image100K<n<1M4 likes437 downloads3y agoHugging FaceWueNLP /Synthdog-Multilingual-100 Synthdog Multilingual The Synthdog dataset created for training in Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model. Using the official Synthdog code, we created >1 million training samples for improving OCR capabilities in Large Vision-Language Models. Dataset Details We provide the images for download in two .tar.gz files. Download and extract them in folders of the same name (so cat images.tar.gz.* | tar xvzf -C images; tar xvzf… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/Synthdog-Multilingual-100.imageimage-to-text1M<n<10M4 likes394 downloads2y agoHugging Face