datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RenderedTextThis dataset has been created by Stability AI and LAION.
This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions.
Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.olmocr-pre-rendered
olmOCR-bench Pre-Rendered
Pre-rendered PNG images of the olmOCR-bench benchmark dataset, ready for zero-setup evaluation of any OCR / vision model.
What This Is
The official olmOCR benchmark requires downloading 1,403 PDFs locally and rendering each page to a PNG image before sending it to a model. Every benchmark runner in the official repo does this same rendering step internally — see olmocr/data/renderpdf.py::render_pdf_to_base64png().
This dataset eliminates that… See the full description on the dataset page: https://huggingface.co/datasets/shhdwi/olmocr-pre-rendered.rendered-wikipedia-english
Dataset Card for Team-PIXEL/rendered-wikipedia-english
Dataset Summary
This dataset contains the full English Wikipedia from February 1, 2018, rendered into images of 16x8464 resolution.
The original text dataset was built from a Wikipedia dump. Each example in the original text dataset contained the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). Each rendered example contains a subset of one full article.… See the full description on the dataset page: https://huggingface.co/datasets/Team-PIXEL/rendered-wikipedia-english.IconStack-48M-Rendered-Traintranscoda-rendered-row-343k-full-pipeline-v1-shardsrendered-sst2
Rendered SST-2
The Rendered SST-2 Dataset from Open AI.
Rendered SST2 is an image classification dataset used to evaluate the models capability on optical character recognition. This dataset was generated by rendering sentences in the Standford Sentiment Treebank v2 dataset.
This dataset contains two classes (positive and negative) and is divided in three splits: a train split containing 6920 images (3610 positive and 3310 negative), a validation split containing 872 images (444… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/rendered-sst2.i1-rendered_text-tfrecordi1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
Princeton University
[arXiv][code][model][project page]
Overview
To prepare the dataset for training, we store the image-caption pairs as TFRecords.
This HuggingFace dataset contains the TFRecords corresponding to the rendered_text dataset at 256×256 resolution.
It also serves as an example of what a dataset processed using… See the full description on the dataset page: https://huggingface.co/datasets/i1-datasets/i1-rendered_text-tfrecord.i1-rendered_text-512-resolution-1m-tfrecordi1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu
Princeton University
[arXiv][code][model][project page]
Overview
To prepare the dataset for training, we store the image-caption pairs as TFRecords.
This HuggingFace dataset contains the TFRecords corresponding to the rendered_text dataset at 512×512 resolution. Concretely, we only retain raw images with a shorter edge of… See the full description on the dataset page: https://huggingface.co/datasets/i1-datasets/i1-rendered_text-512-resolution-1m-tfrecord.rendered-bookcorpus-8x8rendered-bookcorpus
Dataset Card for Team-PIXEL/rendered-bookcorpus
Dataset Summary
This dataset is a version of the BookCorpus available at https://huggingface.co/datasets/bookcorpusopen with examples rendered as images with resolution 16x8464 pixels.
The original BookCorpus was introduced by Zhu et al. (2015) in Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books and contains 17868 books of various genres. The rendered BookCorpus was used… See the full description on the dataset page: https://huggingface.co/datasets/Team-PIXEL/rendered-bookcorpus.rendered-bookcorpus-bigramsrendered-sts17
Dataset Summary
This dataset is rendered to images from STS-17. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load Arabic to Arabic dataset:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts17", name="ar-ar", split="test")
Load French to English dataset:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-sts17.transcoda-rendered-row-343k-full-pipeline-v1
Transcoda Rendered Row 343k Full Pipeline v1
Full-page rendered Transcoda row dataset generated from synthetic and random-notation transcriptions.
Target contents: 343113
Accepted contents: 343027
Failed/dropped contents: 86
Renderings per accepted content: 4
Accepted images: 1372108
Source counts: {'random': 99990, 'synth': 243037}
Staging repo: cminst/transcoda-rendered-row-343k-full-pipeline-v1-shards
Each row contains one transcription and four independently rendered page… See the full description on the dataset page: https://huggingface.co/datasets/cminst/transcoda-rendered-row-343k-full-pipeline-v1.rendered-wikipedia-8x8-withTextdaruma-SFT-renderedrendered-stsb
Dataset Summary
This dataset is rendered to images from STS-benchmark. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load English train Dataset:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-stsb", name="en", split="train")
Load Chinese dev Dataset:
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-Linguist/rendered-stsb.Nano3D_rendered_v2rendered-bookcorpus-16x16rendered-wikipedia-en-8x8rendered-wiki_en-bigramsmm_rendered_textBEAT_Rendered_Videosnano3d-rendered-v1-100k
nano3d-rendered-v1-100k
Mirror of the exact bmcore v24 local holdout subset: 2000 video files.
Benchmark label: real. Classical rendering maps to real under the current video taxonomy.
Source reference: https://huggingface.co/datasets/yanlinli/Nano3D_rendered.
No new license or ownership claim is asserted by this mirror. Original source rights and restrictions remain applicable.
Source revision reviewed: 7cc2f8cebed86b6704e4fa9d60e65f1dc67bf502.
gsm8k-rendered-vlm
Rendered GSM8K-VL Dataset
Rendered GSM8K-VL is a multimodal math-reasoning dataset for vision-language model evaluation.Each example links:
a GSM8K word problem (question)
the final numeric answer (answer)
cleaned chain-of-thought style reasoning (reasoning)
a rendered image path (image)
This dataset is intended for controlled experiments comparing text-only and image-based reasoning behavior.
Canonical Dataset Artifact
The official dataset release uses:… See the full description on the dataset page: https://huggingface.co/datasets/RodelaG/gsm8k-rendered-vlm.nayana-renderedrendered-bookcorpus-8x8-withTextNayanaBench-rendered-splitrendered-sts13
Dataset Summary
This dataset is rendered to images from STS-13. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts13", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/mteb/rendered-sts13.rendered-sts15
Dataset Summary
This dataset is rendered to images from STS-15. We envision the need to assess vision encoders' abilities to understand texts. A natural way will be assessing them with the STS protocols, with texts rendered into images.
Examples of Use
Load test split:
from datasets import load_dataset
dataset = load_dataset("Pixel-Linguist/rendered-sts15", split="test")
Languages
English-only; for multilingual and cross-lingual datasets, see… See the full description on the dataset page: https://huggingface.co/datasets/mteb/rendered-sts15.svg-rendered
SVG to PNG Rendered Dataset
Dataset Summary
This dataset is a processed version of the svgen-500k-instruct dataset, where SVG images have been converted to PNG format for easier consumption in computer vision and machine learning pipelines. Each successfully converted image maintains the original SVG's visual representation while providing a standardized raster format.
Data Fields
png_processed: Boolean flag indicating whether the conversion was successful… See the full description on the dataset page: https://huggingface.co/datasets/thesantatitan/svg-rendered.
