datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test-grpo-vlm-log-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.EgoLoc-Contact-GRPO
EgoLoc Contact Exact-Moment Grid GRPO
This is a self-contained 3x3 image-grid dataset for GRPO training on exact
contact/start localization. The numbered cells are chronological and use 1-based
indices.
This dataset is used to improve a VLM's accuracy for the EgoLoc pipeline.
This dataset IS NOT shuffled. When undergoing GRPO, recommend shuffling the dataset.
3x3 grid dataset for VLM tuning on contact frame identification.
Splits
Training rows: 1389
Validation… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/EgoLoc-Contact-GRPO.eagle-grpo-iter19-fp4-encsurfer-grpo
Kedar84/surfer-grpo
Source run file: dataset-run-1758885513916.jsonl
Generated: 2025-09-27 (UTC)
Fields:
url: Source page URL gathered by Surfer automation.
bounding_box: Normalized viewport coordinates for the target element.
raw_ss: PNG screenshot of the raw page stored in images/raw/.
annotated_ss: PNG screenshot with bounding boxes overlayed (images/annotated/).
Loading Example
from datasets import load_dataset
ds = load_dataset("Kedar84/surfer-grpo"… See the full description on the dataset page: https://huggingface.co/datasets/Kedar84/surfer-grpo.simplevla-grpo-assets
SimpleVLA GRPO Grasp Assets
This dataset contains the released USD object assets used by the SimpleVLA-style GRPO grasping experiments.
Expected local layout after running scripts/download_assets.sh:
/data4/nerako/reasoning/RLinf_assets/grasp_assets/
grpo-resultseagle-grpo-iter19-q4k-uniform-noimatrix-enc
eagle-grpo-iter19 — I-Quality pack, uniform Q4_K, NO imatrix
READ THIS BEFORE COMPARING THIS PACK TO ANY iter_267 NUMBER.
What this is
An I-Quality .iqpt pack of the Oaica V4-Flash iter19 checkpoint (SFT + GRPO final,
DeepSeek-V4-Flash 284B, 43 layers, 256 routed experts/layer), produced by
pipeline/iquality/pack_cbalanced_proposed.py --preset c-balanced-proposed.
Files are AES-256-CTR encrypted (one random IV per file). The manifest mapping
original_path ->… See the full description on the dataset page: https://huggingface.co/datasets/sprapp/eagle-grpo-iter19-q4k-uniform-noimatrix-enc.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval.
The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.grpo-gsm8k-experimentsdetails_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default.SELF_GRPO
👋 Hi, everyone!
verl is a RL training library initiated by ByteDance Seed team and maintained by the verl community.
verl: Volcano Engine Reinforcement Learning for LLMs
verl is a flexible, efficient and production-ready RL training library for large language models (LLMs).
verl is the open-source version of HybridFlow: A Flexible and Efficient RLHF Framework paper.
verl is flexible and easy to use with:
Easy extension of diverse RL algorithms: The… See the full description on the dataset page: https://huggingface.co/datasets/ZswSensetime/SELF_GRPO.GRPO-Card 强化学习环境 :verl-agent-uts.tar.gz
envace2.0-envscaler-grpo-32gpu-rollouts-20260520
EnvACE 2.0 EnvScaler GRPO 32-GPU Rollouts (2026-05-20)
This dataset is the run-artifact backup for envscaler_non_conv_rl_grpo_32gpu_aligned_20260520_055146, an aligned 32-GPU non-conversation GRPO training run.
Contents
rollouts/train/: 201 JSONL dumps, steps 0 through 200
rollouts/val/: 41 JSONL dumps, validation snapshots through step 200
logs/: 177 driver, actor, reference, reward, and environment logs
419 backed-up business files in total
The corresponding… See the full description on the dataset page: https://huggingface.co/datasets/xuzishan/envace2.0-envscaler-grpo-32gpu-rollouts-20260520.es-grpo-results
ES vs GRPO: Evaluation Results
Overview
This repository contains comprehensive evaluation results comparing Evolution Strategies (ES) and Group Relative Policy Optimization (GRPO) on mathematical reasoning tasks.
Tasks
GSM8K: Grade school math word problems
Countdown: Arithmetic equation generation puzzle
Models
Qwen2.5-3B-Instruct
Qwen2.5-3B-Base (with custom tokenizer)
Llama-3.2-3B-Instruct
Llama-3.2-3B-Base (with custom tokenizer)
File… See the full description on the dataset page: https://huggingface.co/datasets/alphaXiv/es-grpo-results.clinical-qa-grpo-qwen3finqa-grpo-inpdetails_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat.grpo_candidates_ta1grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat.NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_tau_0.3_last_timeTiCo-Bench-outputs-Qwen3-Omni-GRPO-ckpt600
TiCo (Qwen3-Omni-30B-A3B) GRPO checkpoint-600: TiCo-Bench outputs
Inference outputs of the TiCo model built on Qwen3-Omni-30B-A3B-Instruct
(SFT LoRA ga642381/TiCo-Qwen3-Omni-SFT-LoRA, then GRPO + CHORD on the
4,000 exact-duration prompts, checkpoint at step 600, LoRA merged) on the
2,000-sample TiCo-Bench speech-query benchmark (WeiChihChen/TiCo-Bench).
For each benchmark id there are two files:
<id>.wav: the generated speech (24 kHz mono).
<id>.txt: the full prompt and the… See the full description on the dataset page: https://huggingface.co/datasets/ga642381/TiCo-Bench-outputs-Qwen3-Omni-GRPO-ckpt600.medical-rl-grpo-v1grpo-completions-qwen3-0.6b
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo-completions-qwen3-0.6b.NEW_qwen2_5_MATH_1_5b_grpo_reg_grpo_bce_4grpo-dapo_shuffled-01_offline-grpo-dapo-qwen3-1.7B-Base-mbs128-n4-mbs128-n4_mathevaldetails_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine.grpo_logs
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo_logs.
