CoolFace
Datasetpublic

bruce360568/SRPO_RL_datasets

SRPO Dataset: Reflection-Aware RL Training Data This repository provides the multimodal reasoning dataset used in the paper: SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning We release two versions of the dataset: 39K version (modified_39Krelease.jsonl + images.zip) Enhanced 47K+ version (47K_release_plus.jsonl + 47K_release_plus.zip) Both follow the same unified format, containing multimodal (image–text) reasoning data with… See the full description on the dataset page: https://huggingface.co/datasets/bruce360568/SRPO_RL_datasets.

sourceHugging Facemitupdated 1y agoView on Hugging Face
2likes27downloads
10 commits on main
b6d85c51y ago

Update README.md

bruce360568
9befb8e1y ago

Update README.md

bruce360568
efa34a01y ago

Upload README.md with huggingface_hub

bruce360568
e5c0e381y ago

Upload README.md with huggingface_hub

bruce360568
2d6371e1y ago

Upload data/47K_release_plus.zip with huggingface_hub

bruce360568
38611871y ago

Upload data/images.zip with huggingface_hub

bruce360568
03559801y ago

Upload data/47K_release_plus.jsonl with huggingface_hub

bruce360568
a5557ab1y ago

Upload data/modified_39Krelease.jsonl with huggingface_hub

bruce360568
a9235e11y ago

Upload README.md with huggingface_hub

bruce360568
010f7bc1y ago

initial commit

bruce360568