datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scaffold
SCAFFOLD
SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces is a large-scale multimodal reasoning dataset designed for training and evaluating Vision-Language Models (VLMs) on scientific figure understanding and visual reasoning.
The dataset is constructed from figures extracted from publicly available arXiv research papers and contains 157,387 question-answer pairs covering diverse scientific… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/scaffold.laion70m-no-peopleScaleCap-450k
[Paper] https://arxiv.org/abs/2506.19848
[GitHub] https://github.com/Cooperx521/ScaleCap
ScaleCap450k-Hyper detailed and high quality image caption
Dataset details
This dataset contains 450k image-caption pairs, where the captions are annotated using the ScaleCap pipeline.
For more details, please refer to the paper.
In collecting images for our dataset, we primarily focus on two
aspects: diversity and richness of image content. Given that the ShareGPT4V-100k already… See the full description on the dataset page: https://huggingface.co/datasets/long-xing1/ScaleCap-450k.SCARLET
SCARLET
SCARLET (Scientific Chart and Figure Reasoning Dataset) is a large-scale multimodal reasoning dataset constructed from figures extracted from publicly available arXiv research papers.
The dataset is designed for training and evaluating Vision-Language Models (VLMs) on scientific figure understanding and visual reasoning.
Each figure is paired with multiple question-answer instances spanning several reasoning categories. Every QA pair additionally contains an… See the full description on the dataset page: https://huggingface.co/datasets/AarIsJay/SCARLET.Small-Scale_VSA_Datasetsals-scanscannetpp_mini_val_set_suite
