datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
S1-MMAlignS1-MMAlign
A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset
S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers.
Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. S1-MMAlign aims to… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.Innovator-VL-Instruct-Sciencemmcl-scienceqa
scienceqa
Repo: Chengxiang1122/mmcl-scienceqa
Visibility: public
Included archives
archives/train.tar.gz
source: /g/data/cp23/ch3329/data/scienceqa/data
archives/val.tar.gz
source: /g/data/cp23/ch3329/data/scienceqa/data
Each archive preserves the original source path contents for the corresponding split.
