CoolFace
Datasetpublic

surogate/ro_sft_pixmo_cap

Dataset Description PixmoCap is a dataset of very long (roughly 200 words on average), detailed captions. Here we provide the Romanian translation of the PixmoCap dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026). Citation @inproceedings{deitke2025molmo, title={Molmo and pixmo: Open weights and… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_cap.

sourceHugging Facecc-by-nc-4.0updated 1mo agoView on Hugging Face
0likes285downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
surogate/ro_sft_pixmo_cap · CoolFace