CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01surogate /ro_sft_finepdfs Dataset Description FinePDFs is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Here we provide the Romanian split of FinePDFs training set, prepared for OCR: pairs of images (pages) and extracted text. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al.… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_finepdfs.image100K<n<1M0 likes520 downloads28d agoHugging Face02surogate /ro_sft_cosyn Dataset Description CoSyn is a collection of synthetic question-answer pairs about very diverse range of computer-generated images. Here we provide the Romanian translation of the CoSyn dataset (matplotlib-chart, plotly-chart and plotly-table), translated (code + data) with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_cosyn.image100K<n<1M0 likes488 downloads28d agoHugging Face03Ultimatech /rosaryimagen<1K0 likes363 downloads1y agoHugging Face04liangyou03 /SPInfer-ROSMAP SPInfer Brain IF Data This repository contains the compact public SPInfer brain immunofluorescence annotation dataset. Layout dapi/ DAPI channel images: <stem>.tiff marker/ matched marker channel images: <stem>_marker.tiff cellbodies/ manual cell-body instance annotations: <stem>_cellbodies.npy dapimultimask/ manual DAPI/nucleus instance annotations for the annotated subset Root files: README.md manifest.tsv dataset_metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/liangyou03/SPInfer-ROSMAP.imageimage-segmentationn<1K0 likes316 downloads4mo agoHugging Face05surogate /ro_sft_pixmo_cap Dataset Description PixmoCap is a dataset of very long (roughly 200 words on average), detailed captions. Here we provide the Romanian translation of the PixmoCap dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation @inproceedings{deitke2025molmo, title={Molmo and pixmo: Open weights and… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_cap.image100K<n<1M0 likes285 downloads28d agoHugging Face06surogate /ro_sft_pixmo_points Dataset Description PixmoPoints is a dataset of images paired with referring expressions and points marking the locations the referring expression refers to in the image. Here we provide the Romanian translation of the PixmoPoints dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_points.image100K<n<1M0 likes278 downloads28d agoHugging Face07roshbeed /ai-residency-blog-data AI Residency — blog toy data Small slices of the datasets used by the nine systems in RoshBeed/ai-residency, cut down so the toy models in the posts on roshbeed.com train in seconds on a GitHub Actions runner. Every post pins a commit revision of this dataset rather than tracking main, so a figure on the site cannot change because something here did. path what it is source text8/text8-2m.txt first 2,000,000 characters of text8 roshbeed/ai-residency-text8… See the full description on the dataset page: https://huggingface.co/datasets/roshbeed/ai-residency-blog-data.image1K<n<10K0 likes215 downloads10h agoHugging Face08surogate /ro_sft_llava_mix Dataset Description LlavaMix is a dataset constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability. Here we provide the Romanian translation of the LlavaMix dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation @article{liu2023visual… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_llava_mix.image100K<n<1M0 likes214 downloads28d agoHugging Face09Charlotte-Nao /blue-rose-musicaudion<1K0 likes197 downloads7mo agoHugging Face10surogate /ro_sft_laion Dataset Description Laion is a subset of LAION/CC/SBU dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning. Here we provide the Romanian translation of the Laion dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_laion.image100K<n<1M0 likes196 downloads28d agoHugging Face11Neperl /skin-disease-acne-rosacea-normalimage1K<n<10K1 likes188 downloads3mo agoHugging Face12surogate /ro_seedbench2 Dataset Description SEED-Bench-2 is a comprehensive large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs), featuring 24K multiple-choice questions with precise human annotations. It spans 27 evaluation dimensions, assessing both text and image generation. Here we provide the Romanian translation of SEED-Bench-2, translated with gpt-4.1-mini. This dataset is used as a benchmark and is part of the evaluation protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_seedbench2.image10K<n<100K0 likes173 downloads28d agoHugging Face13BreCaspian /ROBOMASTER-2025-LiDAR-ROSBAG ROBOMASTER-2025 · 华北理工大学HORIZON战队 · LiDAR ROSBAG 📖 概述 数据来源: 华北理工大学 HORIZON 战队 — 雷达组依托平台: 华北理工 RM 创新实验室录制时间地点: ROBOMASTER 2025 超级对抗赛,北京理工大学(珠海)南部赛区现场实录数据用途: ROBOMASTER 场景下的点云识别、目标检测、三维建图等任务 🗂️ 数据概览 文件名 时长 大小 消息数 点云话题 RM-LiDAR-ROSBAG_01.bag 13分22秒 11.2 GB 8037 /cloudpoints RM-LiDAR-ROSBAG_02.bag 13分59秒 12.9 GB 8399 /cloudpoints 数据格式为标准 ROS 1 .bag 文件,未压缩,采样频率约为 10 Hz。 ​ 🎥… See the full description on the dataset page: https://huggingface.co/datasets/BreCaspian/ROBOMASTER-2025-LiDAR-ROSBAG.imageroboticsn<1K0 likes164 downloads1y agoHugging Face14OpenLLM-Ro /ro_sft_finepdfs Dataset Description FinePDFs is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Here we provide the Romanian split of FinePDFs training set, prepared for OCR: pairs of images (pages) and extracted text. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al.… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_finepdfs.image100K<n<1M1 likes161 downloads4mo agoHugging Face15surogate /ro_sft_pixmo_aa Dataset Description PixmoAA is an instruction-tuning dataset for vision-language models. It contains human-authored question-answer pairs about diverse images with long-form answers. Here we provide the Romanian translation of the PixmoAA dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_aa.image100K<n<1M0 likes160 downloads28d agoHugging Face16Junlinp /rosbag_disparity_701image1K<n<10K0 likes140 downloads6mo agoHugging Face17WaterMelon2333 /Rosetum3Dimage10K<n<100K1 likes138 downloads4mo agoHugging Face18ROSCOSMOS /Movie-Poster-WebURL-Dataset-1874-2025 Movie Poster WebURL Dataset 1874–2025 A TMDB-derived metadata index of movie poster WebURLs covering 1874–2025. The dataset contains metadata and external TMDB poster URLs. Poster image binaries are not redistributed in this repository. Data Split: train Rows: 804,304 Format: Parquet Columns: 15 The publication artifact was produced from a larger local TMDB harvest and passed a conservative metadata-based content filtering and post-filter verification process… See the full description on the dataset page: https://huggingface.co/datasets/ROSCOSMOS/Movie-Poster-WebURL-Dataset-1874-2025.image100K<n<1M2 likes125 downloads17d agoHugging Face19TheKarna /skin-disease-acne-rosacea-normalimage1K<n<10K0 likes118 downloads1mo agoHugging Face20hexasix /rosesimage1K<n<10K0 likes117 downloads2y agoHugging Face21xrkong /nuway_rosbag nUWAy2 ROS Data Repository Overview This repository contains a collection of ROS bag files (in MCAP format) and video streams from an autonomous shuttle bus. It includes data from single run, featuring a variety of sensors such as VLP-16 LiDAR, safety lidar, CAN bus signals, GPS, a 9-axis IMU, and two camera streams. Dataset Description The data is organized into multiple runs, each containing synchronized streams from the following sensors: VLP-16 LiDAR: 3D… See the full description on the dataset page: https://huggingface.co/datasets/xrkong/nuway_rosbag.imageobject-detectionn<1K1 likes115 downloads2y agoHugging Face22OpenLLM-Ro /ro_sft_llava_mix Dataset Description LlavaMix is a dataset constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability. Here we provide the Romanian translation of the LlavaMix dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation @article{liu2023visual… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_llava_mix.image100K<n<1M1 likes98 downloads4mo agoHugging Face23surogate /ro_sft_pixmo_count Dataset Description PixmoCount is a dataset of images paired with number of objects in the image. Here we provide the Romanian translation of the PixmoCount dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation @inproceedings{deitke2025molmo, title={Molmo and pixmo: Open weights and open data… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_count.image10K<n<100K0 likes94 downloads28d agoHugging Face24roshbeed /ai-residency-multimodal-captioning-dataimagen<1K0 likes90 downloads1mo agoHugging Face25Roshnig /Traffic_Sign_Dataset_Parquetimage10K<n<100K1 likes81 downloads3y agoHugging Face26rose-lab-mines /rose-lunar-vtb-assetsimage1K<n<10K0 likes66 downloads8d agoHugging Face27sysuwyh357 /ROSE-v0.1 ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models ROSE (Reference-conditioned Oddity and Symbolic Execution) is a controlled benchmark for evaluating whether multimodal large language models can turn fine-grained visual evidence into the symbolic action required by the current task context. Paper: arXiv:2606.19965 PDF: arXiv PDF Project page: https://xbdxwyh.github.io/ROSE-v0.1/ Evaluation code: https://github.com/xbdxwyh/ROSE-v0.1 Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/sysuwyh357/ROSE-v0.1.image1K<n<10K0 likes61 downloads3mo agoHugging Face28Goamhobala /instaroad-rosa-tilesimage1K<n<10K0 likes61 downloads2d agoHugging Face29OpenLLM-Ro /ro_seedbench2 Dataset Description SEED-Bench-2 is a comprehensive large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs), featuring 24K multiple-choice questions with precise human annotations. It spans 27 evaluation dimensions, assessing both text and image generation. Here we provide the Romanian translation of SEED-Bench-2, translated with gpt-4.1-mini. This dataset is used as a benchmark and is part of the evaluation protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_seedbench2.image10K<n<100K0 likes60 downloads4mo agoHugging Face30OpenLLM-Ro /ro_sft_laion Dataset Description Laion is a subset of LAION/CC/SBU dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning. Here we provide the Romanian translation of the Laion dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_laion.image100K<n<1M0 likes60 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.