visual-cot
visual_cot_llava_qwq_32b_proj_only_siglip2tmp_visual_cot_qwq_32b_proj_only_siglip2_proj_onlyproj_only_llava_deepseek_r1_distill_llama3_8b_reasoning_visual_cot_1000_samples_10_epochsone_vision-visual_cot-38k_samplesQwen2.5-VL-7B-Grounded-Visual-CoTqwen3vl-4b-visual_cot_4k-visual_cot_gqa_2k-sft-ep10proj_only_llava_deepseek_r1_distill_llama3_8b_reasoning_visual_cot_1000_samplesvisual_cot_qwq_32b_proj_only_siglip2_redo
Visual-CoT
VisCoT Dataset Card
There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance.
To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/deepcs233/Visual-CoT.Joint-VisualCoT
Joint VisualCoT
Joint evidence SFT on Visual-CoT document pages. One assistant target:
{"bboxes_2d": [[x1,y1,x2,y2], ...], "selected_sentences": ["..."], "score_img": 0.0, "score_text": 0.0}
Boxes are integer xyxy in [0, 1000]. Images are not in this repo; resolve image under Visual-CoT cot_image_data/{image}
(deepcs233/Visual-CoT).
Code: Chenfei-Liao/MMProvenceChenfei.
Paper protocol
Image-level no-leak: Stage2 test images never enter Stage1 train (splits/image_splits.json).… See the full description on the dataset page: https://huggingface.co/datasets/Chenfei-Liao/Joint-VisualCoT.Visual-CoT
VisCoT Dataset Card
There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance.
To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/ham18053178427/Visual-CoT.Visual-CoT-Sampledpusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot
PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.Visual-CoT-46k-Distill-Sharegpt-v1
