GRPO
SearchR1-nq_hotpotqa_train-qwen2.5-7b-em-grpo-v0.3base-action-grpoOpenThinker3-7B-SFT-GRPO-DEQwen3-Next-80B-A3B-Thinking-GRPO-Uncensored-i1-GGUFQwen3-Next-80B-A3B-Thinking-GRPO-Uncensored-GGUFDeepSeek-R1-Distill-Qwen-7B-GRPO-i1-GGUFRACRO-7B-CRO-GRPO-i1-GGUFqwen2.5-coder-7b-verireason-grpo-official-full-ft
Datasets
All datasets matching “GRPO”test-grpo-vlm-log-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.EgoLoc-Contact-GRPO
EgoLoc Contact Exact-Moment Grid GRPO
This is a self-contained 3x3 image-grid dataset for GRPO training on exact
contact/start localization. The numbered cells are chronological and use 1-based
indices.
This dataset is used to improve a VLM's accuracy for the EgoLoc pipeline.
This dataset IS NOT shuffled. When undergoing GRPO, recommend shuffling the dataset.
3x3 grid dataset for VLM tuning on contact frame identification.
Splits
Training rows: 1389
Validation… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/EgoLoc-Contact-GRPO.WebShop-Qwen3-8B-GRPO-evalsurfer-grpo
Kedar84/surfer-grpo
Source run file: dataset-run-1758885513916.jsonl
Generated: 2025-09-27 (UTC)
Fields:
url: Source page URL gathered by Surfer automation.
bounding_box: Normalized viewport coordinates for the target element.
raw_ss: PNG screenshot of the raw page stored in images/raw/.
annotated_ss: PNG screenshot with bounding boxes overlayed (images/annotated/).
Loading Example
from datasets import load_dataset
ds = load_dataset("Kedar84/surfer-grpo"… See the full description on the dataset page: https://huggingface.co/datasets/Kedar84/surfer-grpo.eagle-grpo-iter19-fp4-encsimplevla-grpo-assets
SimpleVLA GRPO Grasp Assets
This dataset contains the released USD object assets used by the SimpleVLA-style GRPO grasping experiments.
Expected local layout after running scripts/download_assets.sh:
/data4/nerako/reasoning/RLinf_assets/grasp_assets/
