datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test-grpo-vlm-log-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval.
The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default.grpo-completions-qwen3-0.6b
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo-completions-qwen3-0.6b.finqa-grpo-inpPALACE_GRPOgrpo_logs
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo_logs.grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions.details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 10 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-v2.grpo-qwen1.5b-textworld-policy-logitsmultireward-grpo-gsm8k-rewards-qwen2.5-7b
Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct)
Raw rollout-level reward observations from the empirical Section of
"Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis".
This is the data that produced the headline Theorem 3 (correlation-dependent
MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on
real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on
GSM8K test prompts at temperature 0.7.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.danish-json-grpo-v1
danish-json-grpo-v1
10,015 Danish prompts for schema-directed JSON generation, built for GRPO training with a deterministic verifier (parse + key-set match + optional grounding penalty).
Task types
task_type
share
shape
extract
42%
Danish passage + schema → JSON grounded in passage
generate
26%
"Give me JSON for X with fields Y" (values open-ended)
rewrite
22%
Bullet list / semicolon-separated data → JSON with same info
fill_template
10%
JSON… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-json-grpo-v1.action-grpo-datareddit-posts-summarization-grpo
GRPO Summarization Eval Rollouts
Evaluation artifacts for all GRPO summarization checkpoints from smolcluster — a distributed GRPO training framework for Apple Silicon Mac clusters.
Two base models were fine-tuned across two training strategies and six reward configurations each, then evaluated on 200 examples from the mlabonne/smoltldr test split.
Judge: gpt-5-mini-2025-08-07 · Framework: DeepEval G-Eval · Rounds: 5 averaged · Metrics (each 0–1): Faithfulness · Coverage ·… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/reddit-posts-summarization-grpo.grpo_qwen3_1p7B_base_polaris_rolloutsdanish-ner-grpo-v1
danish-ner-grpo-v1
Danish NER rows over DANSK passages, generated by scripts/gen_ner_sft.py with --exclude-src against danish-ner-sft-v1, so no passage here appears in that dataset's train split. Built for GRPO: the intended policy already fits the v1 train split (reward ~1.0).
split
rows
train
9000
eval
400
eval_format
400
val
400
agenttune-agent-ops-GRPO-tracesdanish-if-grpo-combined-v4Code-Contests-Plus-HQ-2x-GRPOCode-Contests-Plus-HQ-1x-GRPOmrm8488__phi-4-14B-grpo-gsm8k-3e-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-gsm8k-3e
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-gsm8k-3e
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-gsm8k-3e-details.danish-if-grpo-combined-v3
danish-if-grpo-combined-v3
10,000 Danish instruction-following prompts for GRPO reward training,
built by rewriting danish-instruction-following-v4 prompts with a
sampled mix of our 46 Danish constraints + ~24 Google IFEval-schema
constraints.
What's new vs v2
Structural fixes (see scripts/constraint_compat.py on
github):
0% impossible constraint combos (v2 had 15.2%; e.g. constrained_response
paired with 20+ content-shape rules that can't co-satisfy). Enforced… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-if-grpo-combined-v3.details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-noRW
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-noRW
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-noRW.
The dataset is composed of 2 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-cosine-noRW.Dongwei__DeepSeek-R1-Distill-Qwen-7B-GRPO-details
Dataset Card for Evaluation run of Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO
Dataset automatically created during the evaluation run of model Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dongwei__DeepSeek-R1-Distill-Qwen-7B-GRPO-details.mrm8488__phi-4-14B-grpo-limo-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-limo
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-limo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-limo-details.danish-icl-grpo-v1
danish-icl-grpo-v1
In-context schema/format induction rows, generated by scripts/gen_icl_schema_format.py with --exclude-src against the danish-icl-schema-format-v3 build, so no passage here appears in that dataset's train split or in any of its published eval splits. Built for GRPO: the intended policy already fits the v3 train split (reward ~1.0), which trains recall rather than induction.
split
rows
train
5854
eval_schema
6146
caliber-extension-gemma4-e2b-grpo-rollouts
CALIBER Extension — Gemma4-E2B GRPO Rollouts
Training rollouts from matched GRPO arms on google/gemma-4-E2B-it
(new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps).
Subsets
subset
arm
τ
prior
rows
mean reward_total
accuracy
full schema
caliber
vanilla CALIBER
0.0
—
1600
2.298
0.514
0.664
mink
Min-K% prior
1.0
mink_0.2
4800
2.506
0.520
0.680
minkpp
Min-K++% prior
1.0
minkpp_0.2
4800
2.637
0.541
0.726
Load:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.yasserrmd__Coder-GRPO-3B-details
Dataset Card for Evaluation run of yasserrmd/Coder-GRPO-3B
Dataset automatically created during the evaluation run of model yasserrmd/Coder-GRPO-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yasserrmd__Coder-GRPO-3B-details.pyroDash-math-grpoarabic-rag-chat-grpo-5K
Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations)
The reinforcement-learning half of
oddadmix/arabic-rag-chat-30K:
same generator, same validator, same schema, disjoint companies. It exists
so GRPO explores fresh knowledge bases instead of taking a second pass over
material the SFT already memorised.
conversations
turns
companies
this pool
5,259
14,018
309
Company-disjointness is exact and verified: this pool shares zero
company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.
