bruce360568/SRPO_RL_datasets
SRPO Dataset: Reflection-Aware RL Training Data This repository provides the multimodal reasoning dataset used in the paper: SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning We release two versions of the dataset: 39K version (modified_39Krelease.jsonl + images.zip) Enhanced 47K+ version (47K_release_plus.jsonl + 47K_release_plus.zip) Both follow the same unified format, containing multimodal (image–text) reasoning data with… See the full description on the dataset page: https://huggingface.co/datasets/bruce360568/SRPO_RL_datasets.
Update README.md
Update README.md
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload data/47K_release_plus.zip with huggingface_hub
Upload data/images.zip with huggingface_hub
Upload data/47K_release_plus.jsonl with huggingface_hub
Upload data/modified_39Krelease.jsonl with huggingface_hub
Upload README.md with huggingface_hub
initial commit
