datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RenderedTextThis dataset has been created by Stability AI and LAION.
This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions.
Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.olmocr-pre-rendered
olmOCR-bench Pre-Rendered
Pre-rendered PNG images of the olmOCR-bench benchmark dataset, ready for zero-setup evaluation of any OCR / vision model.
What This Is
The official olmOCR benchmark requires downloading 1,403 PDFs locally and rendering each page to a PNG image before sending it to a model. Every benchmark runner in the official repo does this same rendering step internally — see olmocr/data/renderpdf.py::render_pdf_to_base64png().
This dataset eliminates that… See the full description on the dataset page: https://huggingface.co/datasets/shhdwi/olmocr-pre-rendered.rendered-wikipedia-english
Dataset Card for Team-PIXEL/rendered-wikipedia-english
Dataset Summary
This dataset contains the full English Wikipedia from February 1, 2018, rendered into images of 16x8464 resolution.
The original text dataset was built from a Wikipedia dump. Each example in the original text dataset contained the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). Each rendered example contains a subset of one full article.… See the full description on the dataset page: https://huggingface.co/datasets/Team-PIXEL/rendered-wikipedia-english.IconStack-48M-Rendered-Trainrendered-bookcorpus
Dataset Card for Team-PIXEL/rendered-bookcorpus
Dataset Summary
This dataset is a version of the BookCorpus available at https://huggingface.co/datasets/bookcorpusopen with examples rendered as images with resolution 16x8464 pixels.
The original BookCorpus was introduced by Zhu et al. (2015) in Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books and contains 17868 books of various genres. The rendered BookCorpus was used… See the full description on the dataset page: https://huggingface.co/datasets/Team-PIXEL/rendered-bookcorpus.transcoda-rendered-row-343k-full-pipeline-v1
Transcoda Rendered Row 343k Full Pipeline v1
Full-page rendered Transcoda row dataset generated from synthetic and random-notation transcriptions.
Target contents: 343113
Accepted contents: 343027
Failed/dropped contents: 86
Renderings per accepted content: 4
Accepted images: 1372108
Source counts: {'random': 99990, 'synth': 243037}
Staging repo: cminst/transcoda-rendered-row-343k-full-pipeline-v1-shards
Each row contains one transcription and four independently rendered page… See the full description on the dataset page: https://huggingface.co/datasets/cminst/transcoda-rendered-row-343k-full-pipeline-v1.rendered-wikipedia-8x8-withTextdaruma-SFT-renderedrendered-bookcorpus-16x16rendered-wikipedia-en-8x8mm_rendered_textnayana-renderedrendered-bookcorpus-8x8-withTextNayanaBench-rendered-splitsvg-rendered
SVG to PNG Rendered Dataset
Dataset Summary
This dataset is a processed version of the svgen-500k-instruct dataset, where SVG images have been converted to PNG format for easier consumption in computer vision and machine learning pipelines. Each successfully converted image maintains the original SVG's visual representation while providing a standardized raster format.
Data Fields
png_processed: Boolean flag indicating whether the conversion was successful… See the full description on the dataset page: https://huggingface.co/datasets/thesantatitan/svg-rendered.wds_renderedsst2IconStack-48M-Rendered-DevRendered_512_32_Testgsm8k-rendered-vlm-v2
GSM8K Rendered-VL v2
1319 rendered GSM8K test problems for the VLM modality study (Phase 1).
Contributors
Rodela Ghosh — study design, pilot (v1), dataset packaging and Hugging Face release (scripts/prepare_hf_v2_release.py)
Aviral Gupta — benchmark infrastructure (src/), v2 rendering protocol (src/rendering.py), Phase 1 model runs
Code: https://github.com/Ro-netizen004/vlm-modality-research
Not interchangeable with v1: RodelaG/gsm8k-rendered-vlm
v1
v2… See the full description on the dataset page: https://huggingface.co/datasets/vlm-modality-research/gsm8k-rendered-vlm-v2.Rendered_Room_Dataset_Samplesvg-rendered-blip_captioned
SVG to PNG Rendered Dataset
Dataset Summary
This dataset is a processed version of the svgen-500k-instruct dataset, where SVG images have been converted to PNG format for easier consumption in computer vision and machine learning pipelines. Each successfully converted image maintains the original SVG's visual representation while providing a standardized raster format.
Data Fields
caption: Image captions generated using Salesforce's BLIP model
png_processed:… See the full description on the dataset page: https://huggingface.co/datasets/thesantatitan/svg-rendered-blip_captioned.rendered-kh-table-suryaocr2pa-warm-start-sft-25b-rendered-review
pa-warm-start-sft-25b rendered review sample (n=200)
200 uniformly-sampled conversations from geodesic-research/pa-warm-start-sft-heavy-25b-mix
(default/train, the control-pretraining 30B baseline SFT corpus), rendered EXACTLY as the
training pack renders them: the library's _chat_preprocess (tool-call normalization +
think-HISTORY chat template + assistant-only loss mask).
Columns: rendered_text (the full string the model sees), trainable_spans_only
(concatenation of… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-25b-rendered-review.mentis-cad-recode-renderedrendered_kh_tables_testlong_surya
Surya OCR 2 Table Recognition on sopheakvoatei/rendered_khmer_tables
Surya OCR 2 table recognition using offline vLLM inference.
The table pipeline automatically detects tall table images and applies
vertical overlapping tiling before reconstructing the final HTML table.
Processing Details
Source Dataset: sopheakvoatei/rendered_khmer_tables
Model: datalab-to/surya-ocr-2
Task: table
Table mode: full
Input column: image
Output column: markdown
Structured column:… See the full description on the dataset page: https://huggingface.co/datasets/sopheakvoatei/rendered_kh_tables_testlong_surya.wds_renderedsst2_test
Rendered SST2 (Test set only)
Original paper: The Visual Task Adaptation Benchmark
Homepage: https://github.com/openai/CLIP/blob/main/data/rendered-sst2.md
Derived from SST2: https://nlp.stanford.edu/sentiment/treebank.html
Bibtex:
@article{zhai2019visual,
title={The Visual Task Adaptation Benchmark},
author={Xiaohua Zhai and Joan Puigcerver and Alexander Kolesnikov and
Pierre Ruyssen and Carlos Riquelme and Mario Lucic and
Josip… See the full description on the dataset page: https://huggingface.co/datasets/djghosh/wds_renderedsst2_test.wds_renderedsst2text2svg-stack-renderedwds_renderedsst2rendered_khmer_tables
