datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WeThink_Multimodal_Reasoning_120K
Dataset Card for WeThink
Repository: https://github.com/yangjie-cv/WeThink
Paper: https://arxiv.org/abs/2506.07905
Dataset Structure
Question-Answer Pairs
The WeThink_Multimodal_Reasoning_120K.jsonl file contains the question-answering data in the following format:
{
"problem": "QUESTION",
"answer": "ANSWER",
"category": "QUESTION TYPE",
"abilities": "QUESTION REQUIRED ABILITIES",
"refined_cot": "THINK PROCESS",
"image_path": "IMAGE PATH"… See the full description on the dataset page: https://huggingface.co/datasets/yangjie-cv/WeThink_Multimodal_Reasoning_120K.WeThink-Multimodal-Reasoning-120K
WeThink-Multimodal-Reasoning-120K
Image Type
Images data can be access from https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k
Image Type
Source Dataset
Images
General Images
COCO
25,344
SAM-1B
18,091
Visual Genome
4,441
GQA
3,251
PISC
835
LLaVA
134
Text-Intensive Images
TextVQA
25,483
ShareTextVQA
538
DocVQA
4,709
OCR-VQA5,142
ChartQA
21,781
Scientific & Technical
GeoQA+
4,813
ScienceQA
4,990
AI2D
1,812
CLEVR-Math
677… See the full description on the dataset page: https://huggingface.co/datasets/WeThink/WeThink-Multimodal-Reasoning-120K.
