surogate/ro_sft_laion
Dataset Description Laion is a subset of LAION/CC/SBU dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning. Here we provide the Romanian translation of the Laion dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_laion.
Dataset Description
<!-- Provide a longer summary of what this dataset is. --> Laion is a subset of LAION/CC/SBU dataset, filtered with a more balanced concept coverage distribution. Captions are also associated with BLIP synthetic caption for reference. It is constructed for the pretraining stage for feature alignment in visual instruction tuning.
Here we provide the Romanian translation of the Laion dataset, translated with Seed-X-PPO. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. -->
@article{liu2023visual,
title={Visual instruction tuning},
author={Liu, Haotian and Li, Chunyuan and Wu, Qingyang and Lee, Yong Jae},
journal={Advances in neural information processing systems},
volume={36},
pages={34892--34916},
year={2023}
}@misc{masala2026intelegi,
title={``\^{I}n\c{t}elegi Rom\^{a}ne\c{s}te?'' A Recipe for Romanian Vision-Language Models},
author={Mihai Masala and Marius Leordeanu and Mihai Dascalu and Traian Rebedea},
year={2026},
eprint={2605.31401},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2605.31401},
}