datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Visual-CoT
VisCoT Dataset Card
There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance.
To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/deepcs233/Visual-CoT.Joint-VisualCoT
Joint VisualCoT
Joint evidence SFT on Visual-CoT document pages. One assistant target:
{"bboxes_2d": [[x1,y1,x2,y2], ...], "selected_sentences": ["..."], "score_img": 0.0, "score_text": 0.0}
Boxes are integer xyxy in [0, 1000]. Images are not in this repo; resolve image under Visual-CoT cot_image_data/{image}
(deepcs233/Visual-CoT).
Code: Chenfei-Liao/MMProvenceChenfei.
Paper protocol
Image-level no-leak: Stage2 test images never enter Stage1 train (splits/image_splits.json).… See the full description on the dataset page: https://huggingface.co/datasets/Chenfei-Liao/Joint-VisualCoT.Visual-CoT
VisCoT Dataset Card
There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance.
To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/ham18053178427/Visual-CoT.Visual-CoT-Sampledpusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot
PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.Visual-CoT-46k-Distill-Sharegpt-v1visual_cot_sample_200
Visual CoT Sample Dataset
This is a sampled subset from the Visual-CoT dataset.
Dataset Description
This dataset contains a random sample of data points from the original Visual CoT dataset,
which focuses on Chain-of-Thought reasoning for multi-modal language models.
Files
sample_200.json: Annotation file containing sampled data
sample_200_images/: Directory containing corresponding images
Usage
from datasets import load_dataset
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/ubowang/visual_cot_sample_200.Visual-CoT-27k-SFT-Sharegpt-v1Visual-CoT-46k-SFT-Sharegpt-v1pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot
PushT int1 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_int1_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
Message format:
Each user turn is the PushT prompt text plus one… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot.Visual-CoT-60kVisual-CoT-4k-Sharegpt-ImagesVisual-CoT-4k-Sharegptvisualcot-step-900pusht_96_norm4_visual_nomarker_stopreq_candidate_shuffle_cot_aligned100k
PushT 96 Norm4 Visual Nomarker Stop-Required Candidate-Shuffle CoT Aligned 100k
This repository contains the exact PushT CoT dataset used for the 2026-06-04 ordered SFT CoT run.
Archive:
pusht_96_norm4_visual_nomarker_stopreq_candidate_shuffle_cot_aligned100k_20260604_ordered.tar.gz
Contents after extraction:
data/train/: 100000 CoT JSONL records in 8 gzip shards.
data/test/: 100 CoT JSONL records in 1 gzip file.
metadata/train_alignment_manifest.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/novastar113/pusht_96_norm4_visual_nomarker_stopreq_candidate_shuffle_cot_aligned100k.visual-cotVisual-CoT-40k-SFT-Sharegpt-v2Visual-CoT-4kVisual-CoT-4k-Distill-SharegptVisual-CoT-4k-Distill-Sharegpt-Matched1kVisual-CoT-27k-SFT-SharegptVisual-CoTVisual-CoT-GQA-2kVisual-CoT-GQA-2k-SharegptVisual-CoT-GQA-2k-Distill-Sharegptqwen-visual-cot-evaluation-logsVisual-CoT-60k-SFTvisualcot_1k_latent_400steps_grpo_1.0temp_64bsz_20260128_182125Visual-CoT-PTvg_visual_cot
