datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grpo-qwen1.5b-textworld-policy-logitsmrm8488__phi-4-14B-grpo-gsm8k-3e-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-gsm8k-3e
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-gsm8k-3e
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-gsm8k-3e-details.caliber-extension-gemma4-e2b-grpo-rollouts
CALIBER Extension — Gemma4-E2B GRPO Rollouts
Training rollouts from matched GRPO arms on google/gemma-4-E2B-it
(new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps).
Subsets
subset
arm
τ
prior
rows
mean reward_total
accuracy
full schema
caliber
vanilla CALIBER
0.0
—
1600
2.298
0.514
0.664
mink
Min-K% prior
1.0
mink_0.2
4800
2.506
0.520
0.680
minkpp
Min-K++% prior
1.0
minkpp_0.2
4800
2.637
0.541
0.726
Load:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.mrm8488__phi-4-14B-grpo-limo-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-limo
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-limo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-limo-details.Dongwei__DeepSeek-R1-Distill-Qwen-7B-GRPO-details
Dataset Card for Evaluation run of Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO
Dataset automatically created during the evaluation run of model Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dongwei__DeepSeek-R1-Distill-Qwen-7B-GRPO-details.yasserrmd__Coder-GRPO-3B-details
Dataset Card for Evaluation run of yasserrmd/Coder-GRPO-3B
Dataset automatically created during the evaluation run of model yasserrmd/Coder-GRPO-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yasserrmd__Coder-GRPO-3B-details.arabic-rag-chat-grpo-5K
Arabic multi-turn RAG conversations — GRPO pool (5,259 conversations)
The reinforcement-learning half of
oddadmix/arabic-rag-chat-30K:
same generator, same validator, same schema, disjoint companies. It exists
so GRPO explores fresh knowledge bases instead of taking a second pass over
material the SFT already memorised.
conversations
turns
companies
this pool
5,259
14,018
309
Company-disjointness is exact and verified: this pool shares zero
company_id values… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-grpo-5K.Nitral-AI__Captain-Eris-BMO_Violent-GRPO-v0.420-details
Dataset Card for Evaluation run of Nitral-AI/Captain-Eris-BMO_Violent-GRPO-v0.420
Dataset automatically created during the evaluation run of model Nitral-AI/Captain-Eris-BMO_Violent-GRPO-v0.420
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nitral-AI__Captain-Eris-BMO_Violent-GRPO-v0.420-details.ymcki__Llama-3.1-8B-SFT-GRPO-Instruct-details
Dataset Card for Evaluation run of ymcki/Llama-3.1-8B-SFT-GRPO-Instruct
Dataset automatically created during the evaluation run of model ymcki/Llama-3.1-8B-SFT-GRPO-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ymcki__Llama-3.1-8B-SFT-GRPO-Instruct-details.pyroDash-math-grpoymcki__Llama-3.1-8B-GRPO-Instruct-details
Dataset Card for Evaluation run of ymcki/Llama-3.1-8B-GRPO-Instruct
Dataset automatically created during the evaluation run of model ymcki/Llama-3.1-8B-GRPO-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ymcki__Llama-3.1-8B-GRPO-Instruct-details.deepseek_grpo_correct_6144
DeepSeek GRPO Correct 6144
Filtered GRPO training subset generated from deepseek-reasoner math generations.
Files
train.jsonl: filtered training examples with prompt, solution, dataset_index, and DeepSeek metadata.
metadata.json: filtering metadata.
Filter
Rows are kept when the raw generation is successful, stopped, correct, deduplicated by dataset_index, and has usage_total_tokens <= 6144.
Summary
Rows: 7576
Max total tokens: 6144
Source raw… See the full description on the dataset page: https://huggingface.co/datasets/igreck/deepseek_grpo_correct_6144.deepseek_grpo_correct_8192
DeepSeek GRPO Correct 8192
Filtered GRPO training subset generated from deepseek-reasoner math generations.
Files
train.jsonl: filtered training examples with prompt, solution, dataset_index, and DeepSeek metadata.
metadata.json: filtering metadata.
Filter
Rows are kept when the raw generation is successful, stopped, correct, deduplicated by dataset_index, and has usage_total_tokens <= 8192.
Summary
Rows: 9002
Max total tokens: 8192
Source raw… See the full description on the dataset page: https://huggingface.co/datasets/igreck/deepseek_grpo_correct_8192.nemo-grpo-from083-full-edge-curation
Nemotron 0.83 Edge-Prompt Curation
This private dataset contains edge-prompt curation rollouts for the DGXChen/Tong CoT dataset.
Seed edge prompts: 134
New rollout rows after seed exclusion: 7668
New edge prompts: 1648
Full edge prompts, seed plus rollout: 1782
Full dataset rows: 7830
Edge rate over full dataset: 0.2276
The Hugging Face dataset viewer is configured to load only data/full_edge_prompts_seed_plus_rollout.jsonl.
The larger rollout and metadata files remain… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-grpo-from083-full-edge-curation.deepseek_grpo_correct_4096
DeepSeek GRPO Correct 4096
Filtered GRPO training subset generated from deepseek-reasoner math generations.
Files
train.jsonl: filtered training examples with prompt, solution, dataset_index, and DeepSeek metadata.
metadata.json: filtering metadata.
Filter
Rows are kept when the raw generation is successful, stopped, correct, deduplicated by dataset_index, and has usage_total_tokens <= 4096.
Summary
Rows: 5502
Max total tokens: 4096
Source raw… See the full description on the dataset page: https://huggingface.co/datasets/igreck/deepseek_grpo_correct_4096.deepseek_grpo_correct_2048
DeepSeek GRPO Correct 2048
Filtered GRPO training subset generated from deepseek-reasoner math generations.
Files
train.jsonl: filtered training examples with prompt, solution, dataset_index, and DeepSeek metadata.
metadata.json: filtering metadata.
Filter
Rows are kept when the raw generation is successful, stopped, correct, deduplicated by dataset_index, and has usage_total_tokens <= 2048.
Summary
Rows: 2153
Max total tokens: 2048
Source raw… See the full description on the dataset page: https://huggingface.co/datasets/igreck/deepseek_grpo_correct_2048.Nitral-AI__Captain-Eris_Violet-GRPO-v0.420-details
Dataset Card for Evaluation run of Nitral-AI/Captain-Eris_Violet-GRPO-v0.420
Dataset automatically created during the evaluation run of model Nitral-AI/Captain-Eris_Violet-GRPO-v0.420
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nitral-AI__Captain-Eris_Violet-GRPO-v0.420-details.RLVE-Test-Qwen3-1.7B-GRPO-step70-Pass8
RLVE test eval — GRPO step70 (pass@8)
Evaluation rollouts on the RLVE test split.
Model: grpo_train_Qwen3-1.7B-SFT-rlve-20K-1epoch (GRPO, step 70)
Source prompts: RLVE test split — 180 questions (RLVE-Eval Gym environments)
Sampling: 8 samples/question (pass@8) = 1440 records, temperature 0.7,
max 16384 new tokens
Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]).
Record-level accuracy (reward>0): 216 / 1440 = 15.0%, mean reward -0.677
pass@8 (>=1 of 8… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Test-Qwen3-1.7B-GRPO-step70-Pass8.verl_agent_sokoban_grpo_qwen3_1.7b_subgoalver3_rolloutNovaciano__Fusetrix-Dolphin-3.2-1B-GRPO_Creative_RP-details
Dataset Card for Evaluation run of Novaciano/Fusetrix-Dolphin-3.2-1B-GRPO_Creative_RP
Dataset automatically created during the evaluation run of model Novaciano/Fusetrix-Dolphin-3.2-1B-GRPO_Creative_RP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Novaciano__Fusetrix-Dolphin-3.2-1B-GRPO_Creative_RP-details.Danielbrdz__Barcenas-3b-GRPO-details
Dataset Card for Evaluation run of Danielbrdz/Barcenas-3b-GRPO
Dataset automatically created during the evaluation run of model Danielbrdz/Barcenas-3b-GRPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Danielbrdz__Barcenas-3b-GRPO-details.game24_GRPO_Qwen2.5_7Bnotbdq__Qwen2.5-14B-Instruct-1M-GRPO-Reasoning-details
Dataset Card for Evaluation run of notbdq/Qwen2.5-14B-Instruct-1M-GRPO-Reasoning
Dataset automatically created during the evaluation run of model notbdq/Qwen2.5-14B-Instruct-1M-GRPO-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/notbdq__Qwen2.5-14B-Instruct-1M-GRPO-Reasoning-details.hatemmahmoud__qwen2.5-1.5b-sft-raft-grpo-hra-doc-details
Dataset Card for Evaluation run of hatemmahmoud/qwen2.5-1.5b-sft-raft-grpo-hra-doc
Dataset automatically created during the evaluation run of model hatemmahmoud/qwen2.5-1.5b-sft-raft-grpo-hra-doc
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/hatemmahmoud__qwen2.5-1.5b-sft-raft-grpo-hra-doc-details.katago-grpo-datasetsdispatchr-grpo-runsfblgit__miniclaus-qw1.5B-UNAMGS-GRPO-details
Dataset Card for Evaluation run of fblgit/miniclaus-qw1.5B-UNAMGS-GRPO
Dataset automatically created during the evaluation run of model fblgit/miniclaus-qw1.5B-UNAMGS-GRPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__miniclaus-qw1.5B-UNAMGS-GRPO-details.Novaciano__FuseChat-3.2-1B-GRPO_Creative_RP-details
Dataset Card for Evaluation run of Novaciano/FuseChat-3.2-1B-GRPO_Creative_RP
Dataset automatically created during the evaluation run of model Novaciano/FuseChat-3.2-1B-GRPO_Creative_RP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Novaciano__FuseChat-3.2-1B-GRPO_Creative_RP-details.smollm2_17b_instruct_0531_inf_eli5_product_dpoed_c4_lowq_200m2b_subsample20m_grpo_promptMagusCorp__grpo_lora_enem_llama3_7b-details
Dataset Card for Evaluation run of MagusCorp/grpo_lora_enem_llama3_7b
Dataset automatically created during the evaluation run of model MagusCorp/grpo_lora_enem_llama3_7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MagusCorp__grpo_lora_enem_llama3_7b-details.
