CoolFace
14 results

visual-cot

deepcs233 /Visual-CoT VisCoT Dataset Card There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance. To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/deepcs233/Visual-CoT.image-text-to-text63 likes6.8k downloads2y agoHugging FaceChenfei-Liao /Joint-VisualCoT Joint VisualCoT Joint evidence SFT on Visual-CoT document pages. One assistant target: {"bboxes_2d": [[x1,y1,x2,y2], ...], "selected_sentences": ["..."], "score_img": 0.0, "score_text": 0.0} Boxes are integer xyxy in [0, 1000]. Images are not in this repo; resolve image under Visual-CoT cot_image_data/{image} (deepcs233/Visual-CoT). Code: Chenfei-Liao/MMProvenceChenfei. Paper protocol Image-level no-leak: Stage2 test images never enter Stage1 train (splits/image_splits.json).… See the full description on the dataset page: https://huggingface.co/datasets/Chenfei-Liao/Joint-VisualCoT.imagevisual-question-answering1M<n<10M0 likes358 downloads21h agoHugging Faceham18053178427 /Visual-CoT VisCoT Dataset Card There is a shortage of multimodal datasets for training multi-modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance. To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying… See the full description on the dataset page: https://huggingface.co/datasets/ham18053178427/Visual-CoT.image-text-to-text0 likes256 downloads8mo agoHugging Faceohjoonhee /Visual-CoT-Sampledimage10K<n<100K0 likes85 downloads10mo agoHugging Facenovastar112 /pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker. Each row contains one full successful trajectory from the first move through the final stop action. Main files: training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows. testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows. metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.imageimage-to-text100K<n<1M0 likes79 downloads4mo agoHugging Faceohjoonhee /Visual-CoT-46k-Distill-Sharegpt-v1image10K<n<100K0 likes67 downloads8mo agoHugging Face