MauroPello/reasoning-gym-verl-datasets
reasoning-gym-verl-datasets This dataset contains procedurally generated reasoning tasks from the Reasoning Gym (r-gym) framework, structured and pre-processed in parquet format for training models with veRL. These datasets were used to train MauroPello/Qwen3-1.7B-RL-final using GRPO (Group Relative Policy Optimization). Dataset Splits & Structure Split Name Path Size (Examples) Description train train.parquet 100,000 Raw training set containing… See the full description on the dataset page: https://huggingface.co/datasets/MauroPello/reasoning-gym-verl-datasets.
085
Upload folder using huggingface_hub
initial commit
