CoolFace
Datasetpublic

bruce360568/SRPO_RL_datasets

SRPO Dataset: Reflection-Aware RL Training Data This repository provides the multimodal reasoning dataset used in the paper: SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning We release two versions of the dataset: 39K version (modified_39Krelease.jsonl + images.zip) Enhanced 47K+ version (47K_release_plus.jsonl + 47K_release_plus.zip) Both follow the same unified format, containing multimodal (image–text) reasoning data with… See the full description on the dataset page: https://huggingface.co/datasets/bruce360568/SRPO_RL_datasets.

sourceHugging Facemitupdated 1y agoView on Hugging Face
2likes27downloads
settings

This repository belongs to bruce360568 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameSRPO_RL_datasets
visibilitypublic
licencemit
gatedno
ownerbruce360568
Account settings
bruce360568/SRPO_RL_datasets · CoolFace