datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VPPO_MMK12_validation
Dataset Card for VPPO_MMK12_validation
Dataset Details
Dataset Description
This dataset is the official validation split used to fine-tune the VPPO-7B and VPPO-32B models presented in our paper, "Spotlight on Token Perception for Multimodal Reinforcement Learning".
This is a direct copy of the test split of FanqingM/MMK12 dataset. We have isolated it here to ensure the exact version used in our experiments is publicly available, guaranteeing reproducibility for… See the full description on the dataset page: https://huggingface.co/datasets/chamber111/VPPO_MMK12_validation.VPPO_ViRL39K_train
Dataset Card for VPPO_ViRL39K_train
Dataset Details
Dataset Description
This dataset is the official training split used to fine-tune the VPPO-7B and VPPO-32B models presented in our paper, "Spotlight on Token Perception for Multimodal Reinforcement Learning".
This is a direct copy of the TIGER-Lab/ViRL39K dataset. We have isolated it here to ensure the exact version used in our experiments is publicly available, guaranteeing reproducibility for our research.… See the full description on the dataset page: https://huggingface.co/datasets/chamber111/VPPO_ViRL39K_train.
