datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Fast-Math-R1-GRPOThis repository contains the second-stage GRPO dataset for the paper A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning.
This dataset is crucial for the second stage of the training recipe, aiming to improve token efficiency while preserving peak mathematical reasoning performance in Large Language Models (LLMs) through Reinforcement Learning from online inference (GRPO).
We extracted the answers from the 2nd stage SFT… See the full description on the dataset page: https://huggingface.co/datasets/RabotniKuma/Fast-Math-R1-GRPO.wordle-grpogrpo-training-plotsgrpo_auto_evaleureka-rebus-grpoforensics-grpo-data
forensics-grpo-data
Generated-video dataset + annotations used to train
sdzt/forensics-grpo.
📂 Repository layout
forensics-grpo-data/
├── video/ # 5,388 .mp4 clips, packed as one .tar per generator
│ ├── 01_vidu.tar # 9.7 GB — Vidu
│ ├── 02_wan.tar # 28 GB — Wan
│ ├── 03_fcvg.tar # 27 GB — FCVG
│ ├── 04_scifi.tar # 34 GB — SciFi
│ ├── 05_ltx.tar # 6.3 GB — LTX
│… See the full description on the dataset page: https://huggingface.co/datasets/sdzt/forensics-grpo-data.qwen35-9b-grpo-unsloth-ut-tft-1000epgrpo-usageqwen35-9b-grpo-unsloth-game-tft-1000epbigmath-grpo-rolloutsGRPOtestingEvaluation_GRPOllama3.1-grpo-r128-a256_log_20250525_123129llama3.1-grpo-r256-a512_log_20250526_114710bigmath-grpo-rollouts-qwen25-3bllama3.1-grpo-r256-a512-base_log_20250526_165221
