CoolFace
Datasetpublic

bruce360568/SRPO_RL_datasets

SRPO Dataset: Reflection-Aware RL Training Data This repository provides the multimodal reasoning dataset used in the paper: SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning We release two versions of the dataset: 39K version (modified_39Krelease.jsonl + images.zip) Enhanced 47K+ version (47K_release_plus.jsonl + 47K_release_plus.zip) Both follow the same unified format, containing multimodal (image–text) reasoning data with… See the full description on the dataset page: https://huggingface.co/datasets/bruce360568/SRPO_RL_datasets.

sourceHugging Facemitupdated 1y agoView on Hugging Face
2likes27downloads

bruce360568/SRPO_RL_datasets · main · files are served by the source, never re-hosted here