CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Video-R1 /Video-R1-dataThis repository contains the data presented in Video-R1: Reinforcing Video Reasoning in MLLMs. Code: https://github.com/tulerfeng/Video-R1 Video data folder: CLEVRER, LLaVA-Video-178K, NeXT-QA, PerceptionTest, STAR Image data folder: Chart, General, Knowledge, Math, OCR, Spatial Video-R1-COT-165k.json is for SFT cold start, and Video-R1-260k.json is for RL training. Data Format in Video-R1-COT-165k: { "problem_id": 2, "problem": "What appears on the screen in Russian during the… See the full description on the dataset page: https://huggingface.co/datasets/Video-R1/Video-R1-data.imagevideo-text-to-text10K<n<100K24 likes4.6k downloads1y agoHugging Face02IffYuan /Embodied-R1.5-SFT-Dataset Embodied-R1.5-SFT-Dataset 🌐 Project Page &nbsp;|&nbsp; 📄 arXiv &nbsp;|&nbsp; 💻 Code &nbsp;|&nbsp; 🧰 EmbodiedEvalKit &nbsp;|&nbsp; 🤗 Models & Datasets 🗓️ Update — 2026-08-20 (20260820). All 34 Stage 1 SFT JSON annotation files have been uploaded to sft_datasets_json/. The complete JSON ↔ image/video data mapping is documented in the Dataset composition table below. ⚠️ Partial release. This repository currently contains only a subset of the full Stage 1 SFT… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-SFT-Dataset.imageimage-text-to-text10 likes3.5k downloads1mo agoHugging Face03Mobile-R1 /Mobile-R1 Dataset Card for Mobile-R1 Dataset Structure images/: All the screenshots data.jsonl: The trajectory data Data Fields All screenshots are stored in the images/ directory. We describe the structure of a single trajectory entry from the file data.jsonl, which contains the full interaction trajectories and action history. app_name: String. The name of the mobile application (e.g., "闲鱼" / Xianyu) where the task is performed. trajectory_length: Integer. Number of… See the full description on the dataset page: https://huggingface.co/datasets/Mobile-R1/Mobile-R1.imageimage-text-to-text1K<n<10K0 likes2.8k downloads5mo agoHugging Face04PG23 /Mobile-R1 Dataset Card for Mobile-R1 Dataset Structure images/: All the screenshots data.jsonl: The trajectory data Data Instances Data Fields All screenshots are stored in the images/ directory. We describe the structure of a single trajectory entry from the file data.jsonl, which contains the full interaction trajectories and action history. app_name: String. The name of the mobile application (e.g., "闲鱼" / Xianyu) where the task is performed.… See the full description on the dataset page: https://huggingface.co/datasets/PG23/Mobile-R1.imageimage-text-to-text1K<n<10K8 likes1.1k downloads5mo agoHugging Face05SassyRong /book2skill-qc-r11image1K<n<10K0 likes952 downloads14d agoHugging Face06IffYuan /Embodied-R1.5-RFT-Dataset Embodied-R1.5-RFT-Dataset 🌐 Project Page &nbsp;|&nbsp; 📄 arXiv &nbsp;|&nbsp; 💻 Code &nbsp;|&nbsp; 🧰 EmbodiedEvalKit &nbsp;|&nbsp; 🤗 Models & Datasets 🗓️ Update — 2026-08-20 (20260820). All 28 Stage 2 RFT JSON annotation files have been uploaded to rft_datasets_json/. The complete JSON ↔ media archive mapping is documented in the Dataset composition table below. ⚠️ Partial release. This repository currently contains only a subset of the full Stage 2 RFT data… See the full description on the dataset page: https://huggingface.co/datasets/IffYuan/Embodied-R1.5-RFT-Dataset.imageimage-text-to-text2 likes754 downloads1mo agoHugging Face07R1l3y-w /im2vec-svg-stack-sample im2vec-svg-stack-sample Rendered (SVG, PNG) pairs for training Im2Vec, a raster-logo-to-SVG model. This is a pre-rendered sample of starvector/svg-stack (2.17M rows total, scraped from permissively-licensed GitHub repos): each .svg is paired with a same-stem .png rendered at 256x256 via cairosvg, matching this project's training pipeline (im2vec/data/render.py). Why this dataset Added alongside im2vec-svg-emoji to give the model more data and more shape diversity… See the full description on the dataset page: https://huggingface.co/datasets/R1l3y-w/im2vec-svg-stack-sample.image10K<n<100K0 likes554 downloads10d agoHugging Face08lmms-lab /multimodal-open-r1-8k-verifiedimage1K<n<10K76 likes534 downloads2y agoHugging Face09huCamus /Mobile-R1 Dataset Card for Mobile-R1 Dataset Structure images/: All the screenshots data.jsonl: The trajectory data Data Fields All screenshots are stored in the images/ directory. We describe the structure of a single trajectory entry from the file data.jsonl, which contains the full interaction trajectories and action history. app_name: String. The name of the mobile application (e.g., "闲鱼" / Xianyu) where the task is performed. trajectory_length:… See the full description on the dataset page: https://huggingface.co/datasets/huCamus/Mobile-R1.imageimage-text-to-text1K<n<10K0 likes512 downloads3mo agoHugging Face10Jiaqi-hkust /Robust-R1Robust-R1:Degradation-Aware Reasoning for Robust Visual Understanding ✅ Copyright Statement Our dataset is built upon a subset of A-OKVQA (Schwenket al. 2022), comprising 10K samples for training and 1K for validation. This statement serves to clarify that the copyright for original visual content is held by the provider. If you wish to use original content itself, you must obtain permission from the respective copyright holders. ⭐️ Citation If you find Robust-R1 useful… See the full description on the dataset page: https://huggingface.co/datasets/Jiaqi-hkust/Robust-R1.imagevisual-question-answering10K<n<100K5 likes462 downloads9mo agoHugging Face11leonardPKU /GEOQA_R1V_Train_8KProcessed from Geo170K, we filtered out the questions with the same (image, answer) pair. image1K<n<10K14 likes417 downloads2y agoHugging Face12minglingfeng /Ocean_R1_collected_visual_dataimage100K<n<1M3 likes398 downloads2y agoHugging Face13omlab /VLM-R1This repository contains the dataset used in the paper VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model. Code: https://github.com/om-ai-lab/VLM-R1 image18 likes332 downloads1y agoHugging Face14yeyimilk /CrowdVLM-R1-BiggerDSPaper: https://huggingface.co/papers/2504.03724 Git repo: https://github.com/yeyimilk/CrowdVLM-R1 Citation @misc{wang2025crowdvlmr1expandingr1ability, title={CrowdVLM-R1: Expanding R1 Ability to Vision Language Model for Crowd Counting using Fuzzy Group Relative Policy Reward}, author={Zhiqiang Wang and Pengbin Feng and Yanbin Lin and Shuzhang Cai and Zongao Bian and Jinghua Yan and Xingquan Zhu}, year={2025}, eprint={2504.03724}, archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/yeyimilk/CrowdVLM-R1-BiggerDS.image10K<n<100K0 likes306 downloads1y agoHugging Face15Osilly /Vision-R1-rlimage10K<n<100K3 likes303 downloads1y agoHugging Face16PetrusCrus /Nav-R1image10K<n<100K0 likes273 downloads9mo agoHugging Face17JefferyZhan /Vision-R1-DataThis repository contains the dataset used in the paper Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning. Code: https://github.com/jefferyZhan/Griffon imageimage-text-to-text10K<n<100K3 likes256 downloads1y agoHugging Face18yuanqiuye /qrcode_image_100k_r16_textdrop0.5image100K<n<1M0 likes234 downloads2y agoHugging Face19JM-Chen /Remember-R1 [ACM MM 2026] Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning This repository provides the training dataset for Remember-R1. 🔗 Quick Links 📄 Paper 💻 GitHub 🤗 Remember-R1-3B 🤗 Remember-R1-7B 🧠 Remember-R1 Dataset 📌 Citation If you find Remember-R1 useful in your research, please cite our paper: @inproceedings{chen2026rememberr1, title={Remember-R1: Mitigating Long-Context Visual Forgetting through… See the full description on the dataset page: https://huggingface.co/datasets/JM-Chen/Remember-R1.image0 likes231 downloads2mo agoHugging Face20mm-vl /mcot_r1_mcq_66kimage10K<n<100K0 likes197 downloads1y agoHugging Face21theaiinstitute /sole-r1-oxe-annot-1026image100K<n<1M0 likes196 downloads4mo agoHugging Face22R1l3y-w /im2vec-svg-emoji im2vec-svg-emoji Rendered (SVG, PNG) pairs for training Im2Vec, a raster-logo-to-SVG model. This is a pre-rendered version of starvector/svg-emoji: each .svg file is paired with a same-stem .png rendered at 256x256 via cairosvg, matching this project's training pipeline (im2vec/data/render.py). Why this dataset The project's original training data, FIGR-8, turned out to contain zero color information — every sampled training SVG relies on the SVG default fill… See the full description on the dataset page: https://huggingface.co/datasets/R1l3y-w/im2vec-svg-emoji.image10K<n<100K0 likes177 downloads13d agoHugging Face23tonia123 /R1-ShareVL-Synimage10K<n<100K0 likes173 downloads8mo agoHugging Face24leonardPKU /GEOQA_8K_R1Vimage1K<n<10K2 likes163 downloads2y agoHugging Face25MMInstruction /Clevr_CoGenT_TrainA_R1image10K<n<100K48 likes140 downloads2y agoHugging Face26LZXzju /UI-R1-3B-TrainThis repository contains the training data presented in UI-R1: Enhancing Action Prediction of GUI Agents by Reinforcement Learning. Project page: https://github.com/lll6gg/UI-R1 imagen<1K7 likes131 downloads1y agoHugging Face27svjack /Koni_R15_Anime_Girl_Images_ZH_Captioned imagen<1K1 likes123 downloads10mo agoHugging Face28jacob-valdez /synthux-economy-r10-w184 SynthUX Computer-Use Dataset — economy sim r10 Grounded visual computer-use trajectories generated by SynthUX. Each record is one worker's device session inside a simulated company: a goal expands into a node tree, terminal nodes drive real desktop-simulator apps (Terminal, Notes, VS Code, Browser, Slack/Teams, Mail, Sheets, Slides, …) through low-level mouse/keyboard input, and the observed trajectory is recorded as a screen video plus per-node frames. App-native… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-economy-r10-w184.imageothern<1K0 likes122 downloads4mo agoHugging Face29adamgavora /SK-IBAN-synthetic-r1 Slovak IBAN OCR Synthetic Dataset Hard samples — obrázky, ktoré Qwen3-VL-4B neprečítal správne, ale GLM-OCR (zai-org/GLM-OCR) ich prečítal korektne (čitateľné). Pipeline filtrovania Generovanie syntetických IBAN obrázkov s degradáciami Qwen3-VL-4B filter — ponechané iba vzorky, ktoré Qwen neprečítal správne (hard) GLM-OCR filter — z hard vzoriek ponechané iba tie, ktoré GLM-OCR prečítal správne (overenie čitateľnosti — nečitateľné obrázky sú zahodené) Štatistiky… See the full description on the dataset page: https://huggingface.co/datasets/adamgavora/SK-IBAN-synthetic-r1.imageimage-to-text1K<n<10K0 likes119 downloads6mo agoHugging Face30yeyimilk /CrowdVLM-R1-dataPaper: https://huggingface.co/papers/2504.03724 Git repo: https://github.com/yeyimilk/CrowdVLM-R1 Citation @misc{wang2025crowdvlmr1expandingr1ability, title={CrowdVLM-R1: Expanding R1 Ability to Vision Language Model for Crowd Counting using Fuzzy Group Relative Policy Reward}, author={Zhiqiang Wang and Pengbin Feng and Yanbin Lin and Shuzhang Cai and Zongao Bian and Jinghua Yan and Xingquan Zhu}, year={2025}, eprint={2504.03724}, archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/yeyimilk/CrowdVLM-R1-data.image1K<n<10K0 likes118 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.