datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VisualWebInstruct-Recall
Introduction
This is the dataset recalled from Google Search from the seed images.
Links
Github|
Paper|
Website
Citation
@article{visualwebinstruct,
title={VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search},
author = {Jia, Yiming and Li, Jiachen and Yue, Xiang and Li, Bo and Nie, Ping and Zou, Kai and Chen, Wenhu},
journal={arXiv preprint arXiv:2503.10582},
year={2025}
}
LLaVA-ReCap-676KThis is an integrated version of LLaVA-ReCap, sourced from lmms-lab/LLaVA-ReCap-558K and lmms-lab/LLaVA-ReCap-118K.
In this version, the conversations field has been split into two separate fields: prompt and response. Additionally, the <image> special token has been removed to facilitate customization.
Inspired by the original paper, the prompt field has been further expanded with human-crafted variations. Specifically, each prompt is sampled from one of the following 30 instructions:… See the full description on the dataset page: https://huggingface.co/datasets/LimeryJorge/LLaVA-ReCap-676K.ReceiptQA
ReceiptQA: A Comprehensive Dataset for Receipt Understanding and Question Answering
ReceiptQA is a large-scale dataset specifically designed to support and advance research in receipt understanding through question-answering (QA) tasks. This dataset offers a wide range of questions derived from real-world receipt images, addressing diverse challenges such as text extraction, layout understanding, and numerical reasoning. ReceiptQA provides a benchmark for evaluating and improving… See the full description on the dataset page: https://huggingface.co/datasets/mahmoud2019/ReceiptQA.cookpad-scrape-recipes
Cookpad India Recipe Archive
Request More ScrapesOrder Private Scrapes
Overview
This repository contains a dataset scraped from cookpad.com/in, a popular community-driven recipe sharing platform. The dataset serves as an extensive archive of diverse, human-created culinary data, capturing home-cooked recipes, ingredient lists, step-by-step instructions, and related web metadata.
Purpose and Usage
This dataset is published publicly and strictly for… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/cookpad-scrape-recipes.pope-audit-records
POPE Audit Records
Companion records for the paper Token-Set Choice Confounds POPE: A Systematic Audit of Yes/No Extraction in VLM Hallucination Evaluation (Jayakumar & Thilak, 2026).
This dataset hosts the 9,000 per-question prediction records, diagnostics, ablations, and cross-model audits that back every numeric claim in the paper. Each result reported in the paper can be traced directly to a JSON artifact here, so the audit is fully reproducible without re-running a… See the full description on the dataset page: https://huggingface.co/datasets/kesav2k04/pope-audit-records.OKReddit-Visionary
Dataset Summary
OKReddit Visionary is a collection of 50 GiB (~74K pairs) of image Question & Answers. This dataset has been prepared for research or archival purposes.
Curated by: KaraKaraWitch
Funded by: Recursal.ai
Shared by: KaraKaraWitch
Special Thanks: harrison (Suggestion)
Language(s) (NLP): Mainly English.
License: Refer to Licensing Information for data license.
Dataset Sources
Source Data: Academic Torrents by (stuck_in_the_matrix, Watchful1… See the full description on the dataset page: https://huggingface.co/datasets/recursal/OKReddit-Visionary.
