datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OCR-Data
OCR Text Detection and Recognition Dataset
Dataset Description
A large-scale, multi-source OCR dataset aggregating 14 public benchmarks for text detection and recognition in both scene images and handwritten documents. Each image is paired with:
Transcribed text for each text region
Bounding boxes (axis-aligned rectangles) for each text region
Polygon coordinates (precise boundary points) for each text region
The dataset is stored in HuggingFace Parquet format with… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/OCR-Data.nsfw1024openocr-web-images
OpenOCR Web Images
从网络爬取的多语言图片数据集,配套多个 OCR 模型的打标结果,用于 OCR 模型
训练/评估数据构建。
数据集结构
data_download/
en/ # 英文候选图片
data/train-part*.parquet # 图片数据,字段见下方"图片数据字段"
labels/
ppocrv6/part*/label-*.parquet # PP-OCRv6 打标结果
hunyuanocr/part*/label-*.parquet # HunyuanOCR (VLM) 打标结果
ch/ # 中文候选图片,目前只有图片,还没跑打标
data/train-part*.parquet
mlt/ # 多语言候选图片,目前只有图片,还没跑打标… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/openocr-web-images.SEEDBench-en-yes-noyesbut
YesBut Dataset (https://yesbut-dataset.github.io/)
Understanding satire and humor is a challenging task for even current Vision-Language models. In this paper, we propose the challenging tasks of Satirical Image Detection (detecting whether an image is satirical), Understanding (generating the reason behind the image being satirical), and Completion (given one half of the image, selecting the other half from 2 given options, such that the complete image is satirical) and release a… See the full description on the dataset page: https://huggingface.co/datasets/bansalaman18/yesbut.SEEDBench-en-yes-no-originyesbut_cleaned
yesbut_cleaned
The yesbut__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
1,077
QA turns
8,287
answers rewritten by the cleaning pass
1,244
QA created by the cleaning pass (new_qa)
3,997 (48.2%)
shards
2
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds wrong but salvageable… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/yesbut_cleaned.scicap-caption-no-more-than-100-tokens-yes-subfigstripes_yesocr_image_urlssampleMMBench-en-yes-noMMMU-en-yes-no-binaryMMMU-en-yes-no-originYes-No-Brain-Tumorimplicitscicap-first-sentence-yes-subfig3x5x500_Datasets_YesAnoMMBench-en-yes-no-binaryTextRIRO_ScenePairvqa-rad-tr-yesno-2025bhutanese-textileMMBench-en-yes-no-originwhatsup-coco-yesnotwitter-yes00988388-2025.09.04-1963446029971591413-qlK3vk_UYLoYN5g-part1uncertain-models-pipeline-artifactsweb_images_store
