datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthdog-en
Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets
For more information, please visit https://github.com/clovaai/donut
The links to the SynthDoG-generated datasets are here:
synthdog-en: English, 0.5M.
synthdog-zh: Chinese, 0.5M.
synthdog-ja: Japanese, 0.5M.
synthdog-ko: Korean, 0.5M.
To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details.
How to Cite
If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-en.synthdog-ko
Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets
For more information, please visit https://github.com/clovaai/donut
The links to the SynthDoG-generated datasets are here:
synthdog-en: English, 0.5M.
synthdog-zh: Chinese, 0.5M.
synthdog-ja: Japanese, 0.5M.
synthdog-ko: Korean, 0.5M.
To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details.
How to Cite
If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-ko.synthdog-ja
Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets
For more information, please visit https://github.com/clovaai/donut
The links to the SynthDoG-generated datasets are here:
synthdog-en: English, 0.5M.
synthdog-zh: Chinese, 0.5M.
synthdog-ja: Japanese, 0.5M.
synthdog-ko: Korean, 0.5M.
To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details.
How to Cite
If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-ja.synthdog-zh
Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets
For more information, please visit https://github.com/clovaai/donut
The links to the SynthDoG-generated datasets are here:
synthdog-en: English, 0.5M.
synthdog-zh: Chinese, 0.5M.
synthdog-ja: Japanese, 0.5M.
synthdog-ko: Korean, 0.5M.
To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details.
How to Cite
If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/naver-clova-ix/synthdog-zh.donut_synthdog_rus
Dataset Card for "donut_rus"
More Information needed
Synthdog-Multilingual-100
Synthdog Multilingual
The Synthdog dataset created for training in Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model.
Using the official Synthdog code, we created >1 million training samples for improving OCR capabilities in Large Vision-Language Models.
Dataset Details
We provide the images for download in two .tar.gz files. Download and extract them in folders of the same name (so cat images.tar.gz.* | tar xvzf -C images; tar xvzf… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/Synthdog-Multilingual-100.SynthDoG-rusynthdog-podbi-enSynthDoG-enSynthDoG-kksynthdog_cleaned
synthdog_cleaned
The synthdog__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
445,694
QA turns
1,613,204
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
547
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/synthdog_cleaned.synthdog-en-detection
SynthDoG detection 🐕
OCR annotations with bounding boxes for synthdog-en generated with PaddleOcr.
This dataset contains annotations where the bounding boxes have the following formats:
2_coord: [(xmin, ymin), (xmax,ymax)]
2_coord: [(xmin/w, ymin/h), (xmax/w,ymax/h)] normalized version of 2_coord where (h, w) are the image height and width
4_coord: [(x1, y1), (x2,y2), (x3,y3), (x4, y4)] all corners of the rectangle enclosing the text span
4_coord_norm: [(x1/w, y1/h), (x2/w,y2/h)… See the full description on the dataset page: https://huggingface.co/datasets/nnethercott/synthdog-en-detection.SynthDoG-kk-500synthdog_en_barcode
Dataset Card for "synthdog_en_barcode"
More Information needed
SynthDog-RU_EN-prepared
Dataset Card for "SynthDog-RU_EN-prepared"
More Information needed
synthdog_cleaned
synthdog_cleaned
The synthdog__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
445,694
QA turns
1,613,204
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
547
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/synthdog_cleaned.SynthDog-RU_EN
Dataset Card for "SynthDog-RU_EN"
More Information needed
synthdog-podbi-plsynthdog-not-rusSynthDog_hu2synthdog-huSynthDog_husynthdog-endeepseek ocr labeled
SynthDoGhi hi
synthdog-en
Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets
For more information, please visit https://github.com/clovaai/donut
The links to the SynthDoG-generated datasets are here:
synthdog-en: English, 0.5M.
synthdog-zh: Chinese, 0.5M.
synthdog-ja: Japanese, 0.5M.
synthdog-ko: Korean, 0.5M.
To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details.
How to Cite
If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/ewqeqw312/synthdog-en.synthdog-zh-twsynthdog-en
Donut 🍩 : OCR-Free Document Understanding Transformer (ECCV 2022) -- SynthDoG datasets
For more information, please visit https://github.com/clovaai/donut
The links to the SynthDoG-generated datasets are here:
synthdog-en: English, 0.5M.
synthdog-zh: Chinese, 0.5M.
synthdog-ja: Japanese, 0.5M.
synthdog-ko: Korean, 0.5M.
To generate synthetic datasets with our SynthDoG, please see ./synthdog/README.md and our paper for details.
How to Cite
If you find this work useful… See the full description on the dataset page: https://huggingface.co/datasets/sdcfsdsfsdds/synthdog-en.clovaai-donut-synthdog-vi
