datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
photo-typography
Photo Typography Dataset
Pulled from Pexels in 2023.
A majority of these images contain text, captioned with CogVLM.
Image filenames may be used as captions, or, the parquet table contains the same values.
This dataset contains the full images.
photo-typography
Photo Typography Dataset
Pulled from Pexels in 2023.
A majority of these images contain text, captioned with CogVLM.
Image filenames may be used as captions, or, the parquet table contains the same values.
This dataset contains the full images.
terminusresearch-photo-typography
Photo Typography — Webshart
This is the canonical, maintained location for the former terminusresearch/photo-typography dataset. The migration completed on August 19, 2026, with 13,167 captioned image samples repackaged into 209 indexed Webshart shards.
The legacy repository now contains only a relocation notice. Its original tar archives, parquet file, and prior history were removed after this copy was validated, avoiding a duplicate of roughly 222 GB. Existing users should… See the full description on the dataset page: https://huggingface.co/datasets/webshart/terminusresearch-photo-typography.hands_and_typography_1024
hands_and_typography_1024
This dataset is a recaptioned merge of two source datasets:
typography: https://huggingface.co/datasets/terminusresearch/photo-typography (13,167 source samples)
hands: https://huggingface.co/datasets/terminusresearch/photo-anatomy (15,425 source samples)
It contains 28,321 image-text samples in bucketed-shards format at 1024-family bucket resolutions. Captions are stored in each sample .txt member and are recorded in the shard metadata as… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/hands_and_typography_1024.GPT4V-captions-from-LVIS-typography
GPT4V-captions-from-LVIS-typography
by: Peter Bevan, 21 March 2023
This dataset is a typography subset of 220k-GPT4Vision-captions-from-LIVIS.
This dataset comprises a subset of 8,857 captioned images from the LVIS dataset. This subset was creating by selecting only image-caption pairs which contain typography that is accurately reflected in the caption.
The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/GPT4V-captions-from-LVIS-typography.photo-typography
Photo Typography has moved
The maintained dataset is now available at webshart/terminusresearch-photo-typography.
The replacement is a fully repackaged Webshart dataset with 13,167 captioned image samples across 209 indexed shards. Its paired JSON indexes include byte offsets, image geometry, and embedded captions for efficient random HTTP range access.
The replacement dataset card contains ready-to-use SimpleTuner and Webshart Python examples.
This legacy repository's tar… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/photo-typography.typographyORGANIC_TYPOGRAPHYarticle-13-bionic-reading-typography
You don't need the internet to run bionic reading mode. We made sure of it.
Bionic Reading Mode: Adaptive Typography and Ambient Light Adjustment for Cognitive Accessibility
The Problem
Bionic Reading is a typographic intervention that partially emboldens the initial letters of words to guide the reader's gaze across text, purportedly improving reading speed and comprehension. Despite widespread popular interest, the technique lacks rigorous empirical validation… See the full description on the dataset page: https://huggingface.co/datasets/Anticloud/article-13-bionic-reading-typography.article-13-bionic-reading-typography
You don't need the internet to run bionic reading mode. We made sure of it.
Bionic Reading Mode: Adaptive Typography and Ambient Light Adjustment for Cognitive Accessibility
The Problem
Bionic Reading is a typographic intervention that partially emboldens the initial letters of words to guide the reader's gaze across text, purportedly improving reading speed and comprehension. Despite widespread popular interest, the technique lacks rigorous empirical validation… See the full description on the dataset page: https://huggingface.co/datasets/kleinnner/article-13-bionic-reading-typography.
