datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test-grpo-vlm-log-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.EgoLoc-Contact-GRPO
EgoLoc Contact Exact-Moment Grid GRPO
This is a self-contained 3x3 image-grid dataset for GRPO training on exact
contact/start localization. The numbered cells are chronological and use 1-based
indices.
This dataset is used to improve a VLM's accuracy for the EgoLoc pipeline.
This dataset IS NOT shuffled. When undergoing GRPO, recommend shuffling the dataset.
3x3 grid dataset for VLM tuning on contact frame identification.
Splits
Training rows: 1389
Validation… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/EgoLoc-Contact-GRPO.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval.
The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default.clinical-qa-grpo-qwen3finqa-grpo-inpdetails_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat.grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat.TiCo-Bench-outputs-Qwen3-Omni-GRPO-ckpt600
TiCo (Qwen3-Omni-30B-A3B) GRPO checkpoint-600: TiCo-Bench outputs
Inference outputs of the TiCo model built on Qwen3-Omni-30B-A3B-Instruct
(SFT LoRA ga642381/TiCo-Qwen3-Omni-SFT-LoRA, then GRPO + CHORD on the
4,000 exact-duration prompts, checkpoint at step 600, LoRA merged) on the
2,000-sample TiCo-Bench speech-query benchmark (WeiChihChen/TiCo-Bench).
For each benchmark id there are two files:
<id>.wav: the generated speech (24 kHz mono).
<id>.txt: the full prompt and the… See the full description on the dataset page: https://huggingface.co/datasets/ga642381/TiCo-Bench-outputs-Qwen3-Omni-GRPO-ckpt600.medical-rl-grpo-v1grpo-completions-qwen3-0.6b
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo-completions-qwen3-0.6b.NEW_qwen2_5_MATH_1_5b_grpo_reg_grpo_bce_4details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine.grpo_logs
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo_logs.NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_4PALACE_GRPOdetails_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted.VHM_dataset_grpoNEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_2scaf-grpo-dataset
Scaf-GRPO Training Dataset
This dataset contains all the training data used in the paper Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning. The training data is derived from the DeepScaleR-Preview-Dataset , a comprehensive collection of 40k mathematical problems sourced from AIME, AMC, MATH, Still, and Omni-MATH.
Data Filtering Strategy
To maximize training efficiency and target problems most conducive to learning, we implement a dynamic… See the full description on the dataset page: https://huggingface.co/datasets/hkuzxc/scaf-grpo-dataset.NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gpg_bce_5EgoLoc-Separation-GRPO
EgoLoc Separation GRPO Dataset
This is a self-contained 3x3 image-grid dataset for GRPO training on exact
separation/end localization. The numbered cells are chronological and use 1-based
indices.
This dataset is used to improve a VLM's accuracy for the EgoLoc pipeline.
This dataset IS NOT shuffled. When undergoing GRPO, recommend shuffling the dataset.
3x3 grid dataset for VLM tuning on separation frame identification.
Splits
Training rows: 1127
Validation rows:… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/EgoLoc-Separation-GRPO.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-cosine
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-cosine
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-cosine.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-cosine.GRPO_Responsedetails_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-v2
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-v2
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-v2.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-v2.NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_5ETCHR-GRPO-10K
ETCHR GRPO-10K
📖Paper
| 🏠Homepage
| 🤗ETCHR-FLUX.2-klein-9B Model
| 🤗ETCHR SFT-400K Dataset
| 🤗ETCHR GRPO-10K Dataset
| 🤗DL3DV-2K Benchmark
ETCHR GRPO-10K is the GRPO training data for further enhance ETCHR's editing capaibility in assisting understanding models. It contains 10000 samples of five tasks (Fine-grained Perception, Chart Understanding, Maze Solving, Jigsaw Puzzle and Spatial Understanding). Each sample contains the image to be edited, an editing… See the full description on the dataset page: https://huggingface.co/datasets/internlm/ETCHR-GRPO-10K.
