CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01depth2world /VLADBenchimage1K<n<10K3 likes51k downloads8mo agoHugging Face02UCSC-VLAA /gpt-edit-simplerimage1M<n<10M13 likes13k downloads1y agoHugging Face03UCSC-VLAA /Recap-DataComp-1B Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions. Dataset Details Dataset Description Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.imagezero-shot-classification1B<n<10B205 likes9.9k downloads2y agoHugging Face04ricl-vla /collected_demos_trainingimage10K<n<100K0 likes8.2k downloads1y agoHugging Face05UCSC-VLAA /GPT-Image-Edit-1.5M GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset 📃Arxiv | 🌐 Project Page | 💻Github GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1. 📣 News [2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download. [2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.imageimage-to-image1M<n<10M90 likes7.1k downloads1y agoHugging Face06UCSC-VLAA /HQ-Edit Dataset Card for HQ-EDIT HQ-Edit, a high-quality instruction-based image editing dataset with total 197,350 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3. HQ-Edit’s high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/HQ-Edit.image1K<n<10K42 likes4k downloads2y agoHugging Face07VLABench /vlm_evaluation_v1.0 Datacard This dataset is the evaluation VLM dataset used in VLABench. It is designed to evaluate the planning capabilities of Vision-Language Models (VLMs) in embodied scenarios. Source Project Page: https://vlabench.github.io/ Arxiv Paper: https://arxiv.org/abs/2412.18194 Code: https://github.com/OpenMOSS/VLABench Uses The dataset structure is as follows: vlm_evaluation_v1.0/ ├── CommenSence/ ├── add_condiment_common_sense/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlm_evaluation_v1.0.image1K<n<10K0 likes3.4k downloads1y agoHugging Face08VladPyatov /CADFS CADFS Dataset CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models 🚀 Project Page 📃 Paper 💻 Code 🤗 Model A large-scale dataset for parametric CAD model generation from text descriptions and multi-view images. Models are represented as FeatureScript programs, enabling direct import into Onshape environment. This dataset was used to train and evaluate CADFS-2B, a fine-tuned Qwen2-VL-2B multimodal language model… See the full description on the dataset page: https://huggingface.co/datasets/VladPyatov/CADFS.3dtext-to-3d100K<n<1M6 likes3.1k downloads2mo agoHugging Face09VLABench /vlabench_primitive_ft_lerobotimage100K<n<1M2 likes2.4k downloads1y agoHugging Face10VLA-Arena /VLA_Arena_L0_L_lerobot_openpi VLA-Arena Dataset (L0 - Large Variant) About VLA-Arena VLA-Arena is an open-source benchmark designed for the systematic evaluation of Vision-Language-Action (VLA) models. It provides a complete and unified toolchain covering scene modeling, demonstration collection, model training, and evaluation. Featuring 150+ tasks across 11 specialized suites, VLA-Arena assesses models through hierarchical difficulty levels (L0-L2) to ensure comprehensive metrics for safety… See the full description on the dataset page: https://huggingface.co/datasets/VLA-Arena/VLA_Arena_L0_L_lerobot_openpi.imagerobotics100K<n<1M0 likes2k downloads7mo agoHugging Face11UCSC-VLAA /Complex-Edit Complex-Edit: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark 📃Arxiv | 🌐Project Page | 💻Github | 📚Dataset | 📄HF Paper We introduce Complex-Edit, a comprehensive benchmark designed to systematically evaluate instruction-based image editing models across instructions of varying complexity. To develop this benchmark, we harness GPT-4o to automatically collect a diverse set of editing instructions at scale. Our approach follows a well-structured… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Complex-Edit.imageimage-to-image1K<n<10K4 likes1.8k downloads1y agoHugging Face12ricl-vla /collected_demosimagen<1K0 likes1.5k downloads1y agoHugging Face13Kanden1112 /surg-vla-datasetimage100K<n<1M0 likes1k downloads3mo agoHugging Face14Vladx /imagenet-w21-wds-dinov2image10M<n<100M0 likes1k downloads1y agoHugging Face15UCSC-VLAA /VLAA-Thinking SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄 Arxiv • 💻 Code 🤗 VLAA-Thinker Family • 🤔 VLAA-Thinking Dataset 🤗 VLAA-Thinker-Qwen2.5-3B • 🤗 VLAA-Thinker-Qwen2.5-7B Both VLAA-Thinker-Qwen2.5-3B and VLAA-Thinker-Qwen2.5-7Bachieve SOTA performance on OpenCompass Multimodal Reasoning Leaderboard as of April 7th, 2025. Contents Quick Start 🚀… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLAA-Thinking.documentvisual-question-answeringn<1K20 likes819 downloads1y agoHugging Face16sxx1012 /vladataimage10K<n<100K0 likes741 downloads3mo agoHugging Face17dadm022 /determ_lerobot_rgb_depth_vlaimage100K<n<1M0 likes701 downloads2mo agoHugging Face18mim-chess-vlas /rs-all-objectsAssets from https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Objects-Kitchen-MJCF Please visit https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Objects-Kitchen-MJCF for more information 3d1K<n<10K1 likes622 downloads5mo agoHugging Face19UCSC-VLAA /MedTrinity-25Mgated Tutorial of using Medtrinity-25M MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities, with multigranular annotations for more than 65 diseases. These enriched annotations encompass both global textual information, such as disease/lesion type, modality, region-specific descriptions, and inter-regional relationships, as well as detailed local annotations for regions of interest (ROIs), including bounding… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedTrinity-25M.imagequestion-answering10M<n<100M214 likes568 downloads2y agoHugging Face20jiaruiguan /omnid_vlabench_dataset_v_0image100K<n<1M0 likes552 downloads1y agoHugging Face21HsiaVyse87 /wm-latentcorr-vla-recoveryimage1K<n<10K0 likes549 downloads4mo agoHugging Face22VLA-Arena /VLA_Arena_L1_L_lerobot_openpi VLA-Arena Dataset (L1 - Large Variant) About VLA-Arena VLA-Arena is an open-source benchmark designed for the systematic evaluation of Vision-Language-Action (VLA) models. It provides a complete and unified toolchain covering scene modeling, demonstration collection, model training, and evaluation. Featuring 150+ tasks across 11 specialized suites, VLA-Arena assesses models through hierarchical difficulty levels (L0-L2) to ensure comprehensive metrics for safety… See the full description on the dataset page: https://huggingface.co/datasets/VLA-Arena/VLA_Arena_L1_L_lerobot_openpi.imagerobotics100K<n<1M0 likes530 downloads7mo agoHugging Face23UCSC-VLAA /MedVLThinker-pmc_vqaCode: https://github.com/UCSC-VLAA/MedVLThinker Project Page: https://ucsc-vlaa.github.io/MedVLThinker/ 📊 Datasets Available Datasets Our project provides several curated datasets for medical vision-language understanding and training: Dataset Modality Description Download MedVLThinker-m23k-tokenized Text-only Tokenized version of the m23k dataset 🤗 HF MedVLThinker-pmc_vqa-gpt_4o_reasoning-tokenized Image-Text Tokenized PMC-VQA dataset with GPT-4o generated… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedVLThinker-pmc_vqa.image100K<n<1M2 likes508 downloads1y agoHugging Face24mattpidden /vla0-context-trace-full-demo-datasetimagen<1K0 likes498 downloads1mo agoHugging Face25UCSC-VLAA /MedVLThinker-EvalCode: https://github.com/UCSC-VLAA/MedVLThinker Project Page: https://ucsc-vlaa.github.io/MedVLThinker/ 📊 Datasets Available Datasets Our project provides several curated datasets for medical vision-language understanding and training: Dataset Modality Description Download MedVLThinker-m23k-tokenized Text-only Tokenized version of the m23k dataset 🤗 HF MedVLThinker-pmc_vqa-gpt_4o_reasoning-tokenized Image-Text Tokenized PMC-VQA dataset with GPT-4o generated… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedVLThinker-Eval.image1K<n<10K3 likes481 downloads1y agoHugging Face26VLAIResearchLab /lerobot_libero_viThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "panda", "total_episodes": 1693, "total_frames": 273465, "total_tasks": 40, "chunks_size": 1000, "fps": 10.0, "splits": { "train": "0:1693" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/VLAIResearchLab/lerobot_libero_vi.imagerobotics100K<n<1M0 likes470 downloads2mo agoHugging Face27vladislavbro /imagesimagen<1K0 likes442 downloads1y agoHugging Face28VLA-Arena /VLA_Arena_L0_L_lerobot_smolvla VLA-Arena Dataset (L0 - Large Variant) About VLA-Arena VLA-Arena is an open-source benchmark designed for the systematic evaluation of Vision-Language-Action (VLA) models. It provides a complete and unified toolchain covering scene modeling, demonstration collection, model training, and evaluation. Featuring 150+ tasks across 11 specialized suites, VLA-Arena assesses models through hierarchical difficulty levels (L0-L2) to ensure comprehensive metrics for safety… See the full description on the dataset page: https://huggingface.co/datasets/VLA-Arena/VLA_Arena_L0_L_lerobot_smolvla.imagerobotics100K<n<1M0 likes330 downloads9mo agoHugging Face29danielkha /VLA_pick_and_place_processedimage100K<n<1M1 likes318 downloads9d agoHugging Face30Ni1111 /giga-vlaimage10K<n<100K0 likes317 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.