datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ro_sft_finepdfs
Dataset Description
FinePDFs is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages.
Here we provide the Romanian split of FinePDFs training set, prepared for OCR: pairs of images (pages) and extracted text. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al.… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_finepdfs.ro_sft_cosyn
Dataset Description
CoSyn is a collection of synthetic question-answer pairs about very diverse range of computer-generated images.
Here we provide the Romanian translation of the CoSyn dataset (matplotlib-chart, plotly-chart and plotly-table), translated (code + data) with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_cosyn.rosarySPInfer-ROSMAP
SPInfer Brain IF Data
This repository contains the compact public SPInfer brain immunofluorescence annotation dataset.
Layout
dapi/ DAPI channel images: <stem>.tiff
marker/ matched marker channel images: <stem>_marker.tiff
cellbodies/ manual cell-body instance annotations: <stem>_cellbodies.npy
dapimultimask/ manual DAPI/nucleus instance annotations for the annotated subset
Root files:
README.md
manifest.tsv
dataset_metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/liangyou03/SPInfer-ROSMAP.ro_sft_pixmo_cap
Dataset Description
PixmoCap is a dataset of very long (roughly 200 words on average), detailed captions.
Here we provide the Romanian translation of the PixmoCap dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@inproceedings{deitke2025molmo,
title={Molmo and pixmo: Open weights and… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_cap.ro_sft_pixmo_points
Dataset Description
PixmoPoints is a dataset of images paired with referring expressions and points marking the locations the referring expression refers to in the image.
Here we provide the Romanian translation of the PixmoPoints dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_points.ai-residency-blog-data
AI Residency — blog toy data
Small slices of the datasets used by the nine systems in
RoshBeed/ai-residency, cut down so the
toy models in the posts on roshbeed.com train in seconds on a
GitHub Actions runner.
Every post pins a commit revision of this dataset rather than tracking main, so a
figure on the site cannot change because something here did.
path
what it is
source
text8/text8-2m.txt
first 2,000,000 characters of text8
roshbeed/ai-residency-text8… See the full description on the dataset page: https://huggingface.co/datasets/roshbeed/ai-residency-blog-data.ro_sft_llava_mix
Dataset Description
LlavaMix is a dataset constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability.
Here we provide the Romanian translation of the LlavaMix dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@article{liu2023visual… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_llava_mix.blue-rose-musicro_sft_laion
Dataset Description
Laion is a subset of LAION/CC/SBU dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning.
Here we provide the Romanian translation of the Laion dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_laion.skin-disease-acne-rosacea-normalro_seedbench2
Dataset Description
SEED-Bench-2 is a comprehensive large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs), featuring 24K multiple-choice questions with precise human annotations. It spans 27 evaluation dimensions, assessing both text and image generation.
Here we provide the Romanian translation of SEED-Bench-2, translated with gpt-4.1-mini. This dataset is used as a benchmark and is part of the evaluation protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_seedbench2.ROBOMASTER-2025-LiDAR-ROSBAG
ROBOMASTER-2025 · 华北理工大学HORIZON战队 · LiDAR ROSBAG
📖 概述
数据来源: 华北理工大学 HORIZON 战队 — 雷达组依托平台: 华北理工 RM 创新实验室录制时间地点: ROBOMASTER 2025 超级对抗赛,北京理工大学(珠海)南部赛区现场实录数据用途: ROBOMASTER 场景下的点云识别、目标检测、三维建图等任务
🗂️ 数据概览
文件名
时长
大小
消息数
点云话题
RM-LiDAR-ROSBAG_01.bag
13分22秒
11.2 GB
8037
/cloudpoints
RM-LiDAR-ROSBAG_02.bag
13分59秒
12.9 GB
8399
/cloudpoints
数据格式为标准 ROS 1 .bag 文件,未压缩,采样频率约为 10 Hz。
🎥… See the full description on the dataset page: https://huggingface.co/datasets/BreCaspian/ROBOMASTER-2025-LiDAR-ROSBAG.ro_sft_finepdfs
Dataset Description
FinePDFs is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages.
Here we provide the Romanian split of FinePDFs training set, prepared for OCR: pairs of images (pages) and extracted text. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al.… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_finepdfs.ro_sft_pixmo_aa
Dataset Description
PixmoAA is an instruction-tuning dataset for vision-language models. It contains human-authored question-answer pairs about diverse images with long-form answers.
Here we provide the Romanian translation of the PixmoAA dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_aa.rosbag_disparity_701Rosetum3DMovie-Poster-WebURL-Dataset-1874-2025
Movie Poster WebURL Dataset 1874–2025
A TMDB-derived metadata index of movie poster WebURLs covering 1874–2025.
The dataset contains metadata and external TMDB poster URLs. Poster image binaries are not redistributed in this repository.
Data
Split: train
Rows: 804,304
Format: Parquet
Columns: 15
The publication artifact was produced from a larger local TMDB harvest and passed a conservative metadata-based content filtering and post-filter verification process… See the full description on the dataset page: https://huggingface.co/datasets/ROSCOSMOS/Movie-Poster-WebURL-Dataset-1874-2025.skin-disease-acne-rosacea-normalrosesnuway_rosbag
nUWAy2 ROS Data Repository
Overview
This repository contains a collection of ROS bag files (in MCAP format) and video streams from an autonomous shuttle bus. It includes data from single run, featuring a variety of sensors such as VLP-16 LiDAR, safety lidar, CAN bus signals, GPS, a 9-axis IMU, and two camera streams.
Dataset Description
The data is organized into multiple runs, each containing synchronized streams from the following sensors:
VLP-16 LiDAR: 3D… See the full description on the dataset page: https://huggingface.co/datasets/xrkong/nuway_rosbag.ro_sft_llava_mix
Dataset Description
LlavaMix is a dataset constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability.
Here we provide the Romanian translation of the LlavaMix dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@article{liu2023visual… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_llava_mix.ro_sft_pixmo_count
Dataset Description
PixmoCount is a dataset of images paired with number of objects in the image.
Here we provide the Romanian translation of the PixmoCount dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@inproceedings{deitke2025molmo,
title={Molmo and pixmo: Open weights and open data… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_count.ai-residency-multimodal-captioning-dataTraffic_Sign_Dataset_Parquetrose-lunar-vtb-assetsROSE-v0.1
ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models
ROSE (Reference-conditioned Oddity and Symbolic Execution) is a controlled benchmark for evaluating whether multimodal large language models can turn fine-grained visual evidence into the symbolic action required by the current task context.
Paper: arXiv:2606.19965
PDF: arXiv PDF
Project page: https://xbdxwyh.github.io/ROSE-v0.1/
Evaluation code: https://github.com/xbdxwyh/ROSE-v0.1
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/sysuwyh357/ROSE-v0.1.instaroad-rosa-tilesro_seedbench2
Dataset Description
SEED-Bench-2 is a comprehensive large-scale benchmark for evaluating Multimodal Large Language Models (MLLMs), featuring 24K multiple-choice questions with precise human annotations. It spans 27 evaluation dimensions, assessing both text and image generation.
Here we provide the Romanian translation of SEED-Bench-2, translated with gpt-4.1-mini. This dataset is used as a benchmark and is part of the evaluation protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_seedbench2.ro_sft_laion
Dataset Description
Laion is a subset of LAION/CC/SBU dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning.
Here we provide the Romanian translation of the Laion dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_laion.
