CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01deepcs233 /Visual-CoT VisCoT Dataset Card There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance. To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/deepcs233/Visual-CoT.image-text-to-text63 likes6.8k downloads2y agoHugging Face02Chenfei-Liao /Joint-VisualCoT Joint VisualCoT Joint evidence SFT on Visual-CoT document pages. One assistant target: {"bboxes_2d": [[x1,y1,x2,y2], ...], "selected_sentences": ["..."], "score_img": 0.0, "score_text": 0.0} Boxes are integer xyxy in [0, 1000]. Images are not in this repo; resolve image under Visual-CoT cot_image_data/{image} (deepcs233/Visual-CoT). Code: Chenfei-Liao/MMProvenceChenfei. Paper protocol Image-level no-leak: Stage2 test images never enter Stage1 train (splits/image_splits.json).… See the full description on the dataset page: https://huggingface.co/datasets/Chenfei-Liao/Joint-VisualCoT.imagevisual-question-answering1M<n<10M0 likes358 downloads21h agoHugging Face03ham18053178427 /Visual-CoT VisCoT Dataset Card There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance. To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/ham18053178427/Visual-CoT.image-text-to-text0 likes256 downloads8mo agoHugging Face04ohjoonhee /Visual-CoT-Sampledimage10K<n<100K0 likes85 downloads10mo agoHugging Face05novastar112 /pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker. Each row contains one full successful trajectory from the first move through the final stop action. Main files: training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows. testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows. metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.imageimage-to-text100K<n<1M0 likes79 downloads4mo agoHugging Face06ohjoonhee /Visual-CoT-46k-Distill-Sharegpt-v1image10K<n<100K0 likes67 downloads8mo agoHugging Face07ubowang /visual_cot_sample_200 Visual CoT Sample Dataset This is a sampled subset from the Visual-CoT dataset. Dataset Description This dataset contains a random sample of data points from the original Visual CoT dataset, which focuses on Chain-of-Thought reasoning for multi-modal language models. Files sample_200.json: Annotation file containing sampled data sample_200_images/: Directory containing corresponding images Usage from datasets import load_dataset # Load the… See the full description on the dataset page: https://huggingface.co/datasets/ubowang/visual_cot_sample_200.imageimage-text-to-textn<1K0 likes52 downloads9mo agoHugging Face08ohjoonhee /Visual-CoT-27k-SFT-Sharegpt-v1image10K<n<100K0 likes42 downloads8mo agoHugging Face09ohjoonhee /Visual-CoT-46k-SFT-Sharegpt-v1image10K<n<100K0 likes30 downloads8mo agoHugging Face10novastar112 /pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot PushT int1 Visual Nomarker All-Step Thinking Trickiness COT This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_int1_visual_nomarker. Each row contains one full successful trajectory from the first move through the final stop action. Main files: training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows. testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows. Message format: Each user turn is the PushT prompt text plus one… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_int1_visual_nomarker_allstep_thinking_trickiness_cot.imageimage-to-text100K<n<1M0 likes24 downloads4mo agoHugging Face11ohjoonhee /Visual-CoT-60kimage10K<n<100K0 likes19 downloads8mo agoHugging Face12ohjoonhee /Visual-CoT-4k-Sharegpt-Imagesimage1K<n<10K0 likes17 downloads9mo agoHugging Face13ohjoonhee /Visual-CoT-4k-Sharegptimage1K<n<10K0 likes13 downloads9mo agoHugging Face14luckychao /visualcot-step-900imagen<1K0 likes12 downloads9mo agoHugging Face15novastar113 /pusht_96_norm4_visual_nomarker_stopreq_candidate_shuffle_cot_aligned100k PushT 96 Norm4 Visual Nomarker Stop-Required Candidate-Shuffle CoT Aligned 100k This repository contains the exact PushT CoT dataset used for the 2026-06-04 ordered SFT CoT run. Archive: pusht_96_norm4_visual_nomarker_stopreq_candidate_shuffle_cot_aligned100k_20260604_ordered.tar.gz Contents after extraction: data/train/: 100000 CoT JSONL records in 8 gzip shards. data/test/: 100 CoT JSONL records in 1 gzip file. metadata/train_alignment_manifest.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/novastar113/pusht_96_norm4_visual_nomarker_stopreq_candidate_shuffle_cot_aligned100k.visual-question-answeringn<1K0 likes12 downloads4mo agoHugging Face16ZachSun /visual-cottext100K<n<1M0 likes10 downloads1y agoHugging Face17ohjoonhee /Visual-CoT-40k-SFT-Sharegpt-v2image10K<n<100K0 likes10 downloads8mo agoHugging Face18ohjoonhee /Visual-CoT-4kimage1K<n<10K0 likes8 downloads10mo agoHugging Face19ohjoonhee /Visual-CoT-4k-Distill-Sharegptimage1K<n<10K0 likes8 downloads8mo agoHugging Face20ohjoonhee /Visual-CoT-4k-Distill-Sharegpt-Matched1kimage1K<n<10K0 likes7 downloads8mo agoHugging Face21ohjoonhee /Visual-CoT-27k-SFT-Sharegptimage10K<n<100K0 likes7 downloads8mo agoHugging Face22ohjoonhee /Visual-CoTimagen<1K0 likes6 downloads10mo agoHugging Face23ohjoonhee /Visual-CoT-GQA-2kimage1K<n<10K0 likes6 downloads9mo agoHugging Face24ohjoonhee /Visual-CoT-GQA-2k-Sharegptimage1K<n<10K0 likes6 downloads9mo agoHugging Face25ohjoonhee /Visual-CoT-GQA-2k-Distill-Sharegptimage1K<n<10K0 likes6 downloads8mo agoHugging Face26caes0r /qwen-visual-cot-evaluation-logs0 likes6 downloads7mo agoHugging Face27ohjoonhee /Visual-CoT-60k-SFTimagen<1K0 likes5 downloads8mo agoHugging Face28luckychao /visualcot_1k_latent_400steps_grpo_1.0temp_64bsz_20260128_182125imagen<1K0 likes4 downloads8mo agoHugging Face29candido-ai /Visual-CoT-PT0 likes2 downloads1y agoHugging Face30Jimyeong1532 /vg_visual_cot0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.