datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fusion-pairwise-evals-finetuned
Automatic pairwise preference evaluations for: Making, not taking, the Best-of-N
Content
This data contains pairwise automatic win-rate evaluations for the m-ArenaHard-v2.0 benchmark and it compares 2 models against gemini-2.5-flash:
Fusion: is the 111B model finetuned on synthetic data generated with Fusion from 5 teachers
BoN: is the 111B model finetuned on synthetic data generated with BoN from 5 teachers
Each model’s outputs are compared in pairs with the respective… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-pairwise-evals-finetuned.GRPO-Fine-Tuned
GSM8K GRPO Dataset for Qwen3-4B Post-Training
A clean, GRPO-ready dataset derived from GSM8K for post-training the Qwen3-4B base model using Group Relative Policy Optimization (GRPO).
Dataset Purpose
This dataset is designed for the GRPO stage of post-training, where the model learns to produce correct mathematical reasoning through reward-based optimization. The key design principles are:
Verifiable answers: Every example has a single, unambiguous numeric answer
Clean… See the full description on the dataset page: https://huggingface.co/datasets/TeamClaude/GRPO-Fine-Tuned.finetune_DataSet
