CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01qgallouedec /test-grpo-vlm-log-completions TRL Completion logs This dataset contains the completions generated during training using trl. The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument). Each file contains the following columns: step: the step of training prompt: the prompt used to generate the completion completion: the completion generated by the model <reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.tabularn<1K0 likes2k downloads6mo agoHugging Face02yuchenxie /EgoLoc-Contact-GRPO EgoLoc Contact Exact-Moment Grid GRPO This is a self-contained 3x3 image-grid dataset for GRPO training on exact contact/start localization. The numbered cells are chronological and use 1-based indices. This dataset is used to improve a VLM's accuracy for the EgoLoc pipeline. This dataset IS NOT shuffled. When undergoing GRPO, recommend shuffling the dataset. 3x3 grid dataset for VLM tuning on contact frame identification. Splits Training rows: 1389 Validation… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/EgoLoc-Contact-GRPO.imagevisual-question-answering1K<n<10K0 likes1.4k downloads2mo agoHugging Face03Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval. The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.tabular1K<n<10K0 likes687 downloads1y agoHugging Face04Lansechen /details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default.tabular1K<n<10K0 likes492 downloads1y agoHugging Face05rntc /clinical-qa-grpo-qwen3text1K<n<10K0 likes333 downloads9mo agoHugging Face06shjondhale /finqa-grpo-inptabular1K<n<10K0 likes287 downloads1y agoHugging Face07Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync.text1K<n<10K0 likes265 downloads1y agoHugging Face08Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-1epochstop-withformat.text1K<n<10K0 likes253 downloads1y agoHugging Face09bihungba1101 /grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions TRL Completion logs This dataset contains the completions generated during training using trl. Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate. The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument). Each file contains the following columns: step: the step of training prompt: the prompt used to generate the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions.tabularn<1K0 likes250 downloads4mo agoHugging Face10Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-noformat.text1K<n<10K0 likes227 downloads1y agoHugging Face11ga642381 /TiCo-Bench-outputs-Qwen3-Omni-GRPO-ckpt600 TiCo (Qwen3-Omni-30B-A3B) GRPO checkpoint-600: TiCo-Bench outputs Inference outputs of the TiCo model built on Qwen3-Omni-30B-A3B-Instruct (SFT LoRA ga642381/TiCo-Qwen3-Omni-SFT-LoRA, then GRPO + CHORD on the 4,000 exact-duration prompts, checkpoint at step 600, LoRA merged) on the 2,000-sample TiCo-Bench speech-query benchmark (WeiChihChen/TiCo-Bench). For each benchmark id there are two files: <id>.wav: the generated speech (24 kHz mono). <id>.txt: the full prompt and the… See the full description on the dataset page: https://huggingface.co/datasets/ga642381/TiCo-Bench-outputs-Qwen3-Omni-GRPO-ckpt600.audiotext-to-speech1K<n<10K0 likes208 downloads11d agoHugging Face12nhn309261 /medical-rl-grpo-v1image10K<n<100K0 likes207 downloads9mo agoHugging Face13essobi /grpo-completions-qwen3-0.6b TRL Completion logs This dataset contains the completions generated during training using trl. The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument). Each file contains the following columns: step: the step of training prompt: the prompt used to generate the completion completion: the completion generated by the model <reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo-completions-qwen3-0.6b.tabularn<1K0 likes206 downloads7mo agoHugging Face14happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_grpo_bce_4textn<1K0 likes204 downloads1y agoHugging Face15Lansechen /details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2 Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2 Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2.tabular1K<n<10K0 likes202 downloads1y agoHugging Face16Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-olympiads-aime-unique-cosine.text1K<n<10K0 likes191 downloads1y agoHugging Face17essobi /grpo_logs TRL Completion logs This dataset contains the completions generated during training using trl. The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument). Each file contains the following columns: step: the step of training prompt: the prompt used to generate the completion completion: the completion generated by the model <reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo_logs.tabularn<1K0 likes185 downloads7mo agoHugging Face18happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_4textn<1K0 likes182 downloads1y agoHugging Face19s1ghhh /PALACE_GRPOtabular100K<n<1M0 likes174 downloads7mo agoHugging Face20Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted.text1K<n<10K0 likes167 downloads1y agoHugging Face21aybora /VHM_dataset_grpoimage10K<n<100K0 likes167 downloads1y agoHugging Face22happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_2textn<1K0 likes162 downloads1y agoHugging Face23hkuzxc /scaf-grpo-dataset Scaf-GRPO Training Dataset This dataset contains all the training data used in the paper Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning. The training data is derived from the DeepScaleR-Preview-Dataset , a comprehensive collection of 40k mathematical problems sourced from AIME, AMC, MATH, Still, and Omni-MATH. Data Filtering Strategy To maximize training efficiency and target problems most conducive to learning, we implement a dynamic… See the full description on the dataset page: https://huggingface.co/datasets/hkuzxc/scaf-grpo-dataset.text10K<n<100K0 likes161 downloads11mo agoHugging Face24happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_reg_beta_0.1_gpg_bce_5textn<1K0 likes159 downloads1y agoHugging Face25yuchenxie /EgoLoc-Separation-GRPO EgoLoc Separation GRPO Dataset This is a self-contained 3x3 image-grid dataset for GRPO training on exact separation/end localization. The numbered cells are chronological and use 1-based indices. This dataset is used to improve a VLM's accuracy for the EgoLoc pipeline. This dataset IS NOT shuffled. When undergoing GRPO, recommend shuffling the dataset. 3x3 grid dataset for VLM tuning on separation frame identification. Splits Training rows: 1127 Validation rows:… See the full description on the dataset page: https://huggingface.co/datasets/yuchenxie/EgoLoc-Separation-GRPO.imagevisual-question-answering1K<n<10K0 likes159 downloads2mo agoHugging Face26Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-cosine Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-cosine Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-cosine. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-cosine.text1K<n<10K0 likes154 downloads1y agoHugging Face27JackyChunKit /GRPO_Responsetext100K<n<1M1 likes153 downloads1y agoHugging Face28Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-v2 Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-v2 Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-v2. The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-v2.text1K<n<10K0 likes149 downloads1y agoHugging Face29happynew111 /NEW_qwen2_5_MATH_1_5b_grpo_AR_Lopti_follow_kk_bce_5textn<1K0 likes148 downloads1y agoHugging Face30internlm /ETCHR-GRPO-10K ETCHR GRPO-10K 📖Paper | 🏠Homepage | 🤗ETCHR-FLUX.2-klein-9B Model | 🤗ETCHR SFT-400K Dataset | 🤗ETCHR GRPO-10K Dataset | 🤗DL3DV-2K Benchmark ETCHR GRPO-10K is the GRPO training data for further enhance ETCHR's editing capaibility in assisting understanding models. It contains 10000 samples of five tasks (Fine-grained Perception, Chart Understanding, Maze Solving, Jigsaw Puzzle and Spatial Understanding). Each sample contains the image to be edited, an editing… See the full description on the dataset page: https://huggingface.co/datasets/internlm/ETCHR-GRPO-10K.imagevisual-question-answering10K<n<100K5 likes147 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.