CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tencent /Hy-Embodied-0.5-VLA-Data Hy-Embodied-0.5-VLA From Vision-Language-Action Models to a Real-World Robot Learning Stack Tencent Robotics X × Tencent Hy Team 📖 Abstract We introduce Hy-Embodied-0.5-VLA (Hy-VLA) — an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, pre-training, supervised fine-tuning, RL post-training, and real-world deployment. Built on the Hy-Embodied-0.5 MoT backbone, Hy-VLA integrates a flow-matching… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data.tabularroboticsn<1K23 likes84k downloads3mo agoHugging Face02InnovatorLab /Innovator-VL-Instruct-46M Innovator-VL-Instruct-46M Paper | Code 🤗🤗 The data is being uploaded continuously Introduction To further enhance the model’s ability to handle a broad range of visual tasks with accurate, grounded, and instruction-aligned responses, we perform full-parameter visual instruction supervised fine-tuning (SFT).This SFT stage serves as a critical bridge between multimodal pretraining and subsequent reinforcement learning, providing both general capability coverage and a… See the full description on the dataset page: https://huggingface.co/datasets/InnovatorLab/Innovator-VL-Instruct-46M.imageimage-text-to-text10M<n<100M9 likes67k downloads8mo agoHugging Face03picbreeder-vlm /picbreeder-vlm-archive Picbreeder-VLM Archive Every image evolved by the swarm of vision-language-model "breeders" in In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models (GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the lineage graphs, and the analysis artifacts behind the paper and blog. The original Picbreeder (Secretan et al., 2008) let crowds of humans collaboratively evolve images from CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.imageimage-to-text100K<n<1M14 likes66k downloads3mo agoHugging Face04Eyz /VLNVerse_scene0 likes62k downloads2mo agoHugging Face05depth2world /VLADBenchimage1K<n<10K3 likes60k downloads8mo agoHugging Face06vLAR /LavalObjaverseDataset Laval Objaverse Dataset vLAR Group | SIGGRAPH Asia 2026 A large-scale, high-quality dataset for multi-view relighting. 📖 Dataset Summary The Laval Objaverse Dataset is a comprehensive dataset designed for multi-view relighting and novel view synthesis tasks. It combines high-quality 3D assets from Objaverse with realistic, diverse illumination conditions… See the full description on the dataset page: https://huggingface.co/datasets/vLAR/LavalObjaverseDataset.3dimage-to-image10M<n<100M11 likes46k downloads1d agoHugging Face07vLAR /PhysInOnePhysInOne: Visual Physics Learning and Reasoning in One Suite vLAR Group | The Hong Kong Polytechnic University | Syai Singapore | Meta CVPR 2026 🧭 Navigation 📌 Summary 🚀 Release Timetable 📦 Repositories & Downloads 📊 Data Splits 🧱 3D Assets 🛠️ Data Processing 🏆 Leaderboard Evaluation Data 🎞️ Rendered Data &nbsp;&nbsp;&nbsp;1. Download Scripts &nbsp;&nbsp;&nbsp;2. Install Dependencies… See the full description on the dataset page: https://huggingface.co/datasets/vLAR/PhysInOne.videotext-to-video1M<n<10M20 likes34k downloads2d agoHugging Face08VLABench /vlabench_primitive_pretrain_lerobot Datacard This is the official VLABench primitive pretraining dataset converted to the LeRobot format. The dataset contains language-conditioned manipulation trajectories collected with a Franka Panda robot in VLABench simulation. This LeRobot version is hosted at: https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot Source Project Page: https://vlabench.github.io/ Arxiv Paper: https://arxiv.org/abs/2412.18194 Code:… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot.0 likes33k downloads3mo agoHugging Face09Eyz /VLNVerse_data0 likes23k downloads6mo agoHugging Face10RoganInglis /vllm-control-arena vLLM Main Tasks Dataset AI coding tasks generated from vLLM git commits Dataset Description This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work. Dataset Structure The dataset contains the following columns: commit_hash: The git commit hash parent_hash: The parent commit hash commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.tabulartext-generation1K<n<10K0 likes21k downloads1y agoHugging Face11VoCuc /vlm-teacher-embedding0 likes19k downloads3d agoHugging Face12mikasa-robo /mikasa-robo-vla-rlds0 likes19k downloads3mo agoHugging Face13princeton-vl /InFlux-Synth InFlux++ Synth InFlux++ Synth is a large-scale synthetic training dataset for the InFlux project, providing per-frame ground truth camera intrinsics and camera pose for videos with dynamic intrinsics. The dataset contains 441,840 annotated frames from 1,841 procedurally generated high-resolution videos. Every video contains 240 frames at a resolution of 1280 × 720. The dataset spans indoor and nature scenes and features changing zoom and focus, dynamic objects, and realistic… See the full description on the dataset page: https://huggingface.co/datasets/princeton-vl/InFlux-Synth.1 likes18k downloads9d agoHugging Face14princeton-vl /LayeredFlow-Syn LayeredFlow-Syn Extracted Ground Truth This repository contains extracted LayeredFlow ground-truth annotations in Parquet format. Layout data/<scene>/<sample>.parquet For example: data/0/0_0.parquet Each Parquet file stores one extracted sample. Rows correspond to files from the extracted sample directory and include a leftmost visualization image preview when available, relative_path, num_bytes, sha256, and binary content. Rows are ordered by frame0/left… See the full description on the dataset page: https://huggingface.co/datasets/princeton-vl/LayeredFlow-Syn.image1M<n<10M0 likes17k downloads3mo agoHugging Face15ethz-vlg /mv3dpt-datasets Multi-View 3D Point Tracking Datasets This repository hosts the training and evaluation datasets associated with the paper Multi-View 3D Point Tracking. Project Page: https://ethz-vlg.github.io/mvtracker/ Code/Github Repository: https://github.com/ethz-vlg/mvtracker Abstract We introduce the first data-driven multi-view 3D point tracker, designed to track arbitrary points in dynamic scenes using multiple camera views. Unlike existing monocular trackers, which struggle… See the full description on the dataset page: https://huggingface.co/datasets/ethz-vlg/mv3dpt-datasets.keypoint-detection3 likes16k downloads6mo agoHugging Face16mim-chess-vlas /eval-resultsvideo100K<n<1M0 likes16k downloads1h agoHugging Face17MonsterDie /VLN_Dataset_2This repository contains encrypted visual features for an ongoing academic research project. Decryption keys are managed internally for reproducibility. 3 likes15k downloads1mo agoHugging Face18PaddlePaddle /PaddleOCR-VL_demoimagen<1K2 likes12k downloads10mo agoHugging Face19UCSC-VLAA /gpt-edit-simplerimage1M<n<10M13 likes11k downloads1y agoHugging Face20lerobot /vlabench-assets1 likes10k downloads5mo agoHugging Face21Telkwevr /Bench2Drive-VL-base Bench2Drive-VL: Full-Stack Software for Closed-Loop Autonomous Driving with Vision Language Models Project Page | GitHub | Paper Bench2Drive-VL is a comprehensive closed-loop benchmark for Vision-Language Models in Autonomous Driving (VLM4AD). It extends the Bench2Drive benchmark by introducing closed-loop evaluation and the DriveCommenter expert model for automated annotation. This repository contains the natural language annotations for the Bench2Drive-Base1000 dataset. These… See the full description on the dataset page: https://huggingface.co/datasets/Telkwevr/Bench2Drive-VL-base.image-text-to-text10M<n<100M1 likes9.9k downloads6mo agoHugging Face22UCSC-VLAA /Recap-DataComp-1B Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced LLaVA-1.5-LLaMA3-8B model to enhance the alignment and detail of textual descriptions. Dataset Details Dataset Description Our paper aims to bridge this community effort, leveraging the powerful and open-sourced LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/Recap-DataComp-1B.imagezero-shot-classification1B<n<10B205 likes9.5k downloads2y agoHugging Face23merve /vlm_test_imagesBunch of random test cases for vision language in the wild. imagen<1K9 likes8.8k downloads6mo agoHugging Face24nvidia /Nemotron-VLM-Dataset-v2 Nemotron-VLM-Dataset v2 Versions Date Commit Changes 2025-11-05 head Fix nights_cot dataset. Fix/filter broken <think> entries. Update fintabnet instructions. Update indexes. 2025-10-28 214051e Initial Release Dataset Description Following up on Llama Nemotron VLM Dataset V1 with 3 million samples, we are releasing the Nemotron VLM Dataset V2 with almost three times as many high-quality samples. This time, our focus was on three… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-VLM-Dataset-v2.textvisual-question-answering1M<n<10M98 likes8.8k downloads9mo agoHugging Face25VLMEval /OpenVLMRecords OpenVLM Records Here we maintain all the evaluation records generated by VLMEvalKit, which also reflects on the OpenVLM Leaderboard. Before using the scripts to browse and utilize those record files, you should first have VLMEvalKit installed (use pip install -e . --no-deps when you encounter some dependency errors). Naming System & Record Browsing In this repo, records are organized with the following naming system: The record file of evaluating MLLM VLM-A on the… See the full description on the dataset page: https://huggingface.co/datasets/VLMEval/OpenVLMRecords.visual-question-answering1M<n<10M15 likes8.5k downloads1y agoHugging Face26ricl-vla /collected_demos_trainingimage10K<n<100K0 likes8.3k downloads1y agoHugging Face27LejuRobotics /LET-KUAVO-VLA-1.0-Datasetgated LET-KUAVO-VLA-1.0-Dataset videon<1K3 likes7.4k downloads18d agoHugging Face28InternRobotics /VLAC-Cut-FullData VLAC-Cut-FullData VLAC-Cut-FullData is the full-data release for VLAC-Cut. It provides the complete raw-data archive set, benchmark-style JSON files, and a lightweight frame-extraction workflow for reproducing evaluation on the released benchmark protocol. Contents benchmark_style_all/ train/video_progress_benchmark_file.json test_expert_seen/video_progress_benchmark_file.json test_expert_unseen/video_progress_benchmark_file.json… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/VLAC-Cut-FullData.3 likes7.2k downloads2mo agoHugging Face29open-cn-llm-leaderboard /vlm_results0 likes7.2k downloads1y agoHugging Face30UCSC-VLAA /GPT-Image-Edit-1.5M GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset 📃Arxiv | 🌐 Project Page | 💻Github GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1. 📣 News [2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download. [2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.imageimage-to-image1M<n<10M90 likes6.8k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.