datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Joint-VisualCoT
Joint VisualCoT
Joint evidence SFT on Visual-CoT document pages. One assistant target:
{"bboxes_2d": [[x1,y1,x2,y2], ...], "selected_sentences": ["..."], "score_img": 0.0, "score_text": 0.0}
Boxes are integer xyxy in [0, 1000]. Images are not in this repo; resolve image under Visual-CoT cot_image_data/{image}
(deepcs233/Visual-CoT).
Code: Chenfei-Liao/MMProvenceChenfei.
Paper protocol
Image-level no-leak: Stage2 test images never enter Stage1 train (splits/image_splits.json).… See the full description on the dataset page: https://huggingface.co/datasets/Chenfei-Liao/Joint-VisualCoT.Visual-CoT-Sampledpusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot
PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.Visual-CoT-46k-Distill-Sharegpt-v1Visual-CoT-27k-SFT-Sharegpt-v1Visual-CoT-46k-SFT-Sharegpt-v1pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot
PushT int1 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_int1_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
Message format:
Each user turn is the PushT prompt text plus one… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot.Visual-CoT-60kVisual-CoT-4k-Sharegpt-ImagesVisual-CoT-4k-Sharegptvisualcot-step-900visual-cotVisual-CoT-40k-SFT-Sharegpt-v2Visual-CoT-4kVisual-CoT-4k-Distill-SharegptVisual-CoT-4k-Distill-Sharegpt-Matched1kVisual-CoT-27k-SFT-SharegptVisual-CoTVisual-CoT-GQA-2kVisual-CoT-GQA-2k-SharegptVisual-CoT-GQA-2k-Distill-SharegptVisual-CoT-60k-SFTvisualcot_1k_latent_400steps_grpo_1.0temp_64bsz_20260128_182125
