datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
assistive-ocr-data-acquisition
Assistive OCR Benchmark Data
Multilingual OCR benchmark for visually impaired assistance — Indian medicine labels, packaged goods, and signage in Bengali, Hindi, and English.
Dataset Summary
Property
Value
Total images
7,004 rows in manifest
Image sources
images/hf_medicines/, images/openfoodfacts/, images/synthetic/
Domains
medicine_packaging (6,588), packaged_goods (386), signage (30)
Languages
bn+en (5,337), hi+en (868), en (799)
Splits
dev… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/assistive-ocr-data-acquisition.gova-3case-word-acquisition
GOVA Word Acquisition — 3-Case Split by Language-Bias Regime
OctoBERT's grounded word-acquisition successes on GOVA (Flickr30k), partitioned
into three regimes by how much the answer depends on language vs. vision.
The two clean poles (Case 1 / Case 3) are defined by the intersection of two
models so they are model-agnostic, not an OctoBERT artifact:
split
definition
n
case1_language_dominant
both OctoBERT-blind AND Qwen2.5-32B solve it text-only
8213… See the full description on the dataset page: https://huggingface.co/datasets/guangliangliu/gova-3case-word-acquisition.octobert-gova-word-acquisition
OctoBERT GOVA Word-Acquisition — per-example predictions (works vs. fails)
Per-example predictions of OctoBERT (World-to-Words, Ma et al., ACL 2023) on the
Grounded Open-Vocabulary Acquisition (GOVA) cloze test over Flickr30k. Each row masks one
word in the caption; the frozen released checkpoint (sled-umich/OctoBERT-flickr, pure
inference, no training) must predict it. This is an analysis / shortcut-diagnosis artifact,
not an official release by the paper's authors.… See the full description on the dataset page: https://huggingface.co/datasets/ZYXue/octobert-gova-word-acquisition.
