bruce360568/SRPO_RL_datasets
SRPO Dataset: Reflection-Aware RL Training Data This repository provides the multimodal reasoning dataset used in the paper: SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning We release two versions of the dataset: 39K version (modified_39Krelease.jsonl + images.zip) Enhanced 47K+ version (47K_release_plus.jsonl + 47K_release_plus.zip) Both follow the same unified format, containing multimodal (image–text) reasoning data with… See the full description on the dataset page: https://huggingface.co/datasets/bruce360568/SRPO_RL_datasets.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face