datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
verl-aha-moment-dataset
Verl "Aha Moment" Dataset for DeepSeek-R1 Reproduction
Overview
This dataset is built for use with the Verl framework to reproduce the "aha moment" phenomenon observed in DeepSeek-R1-Zero. It combines two Hugging Face math-reasoning datasets into a single Parquet file conforming to the format specified in format.json.
Composition
Source
Rows
Description
agentica-org/DeepScaleR-Preview-Dataset
40,315
~40K competition math problems (AIME… See the full description on the dataset page: https://huggingface.co/datasets/dsa1dsa12/verl-aha-moment-dataset.verl-aha-moment-dataset
Verl Dataset for DeepSeek-R1 "Aha Moment" Reproduction
This dataset is prepared for training with the Verl framework
to reproduce the "aha moment" phenomenon observed in DeepSeek-R1-Zero.
What is the "Aha Moment"?
During RL training without any SFT, DeepSeek-R1-Zero spontaneously developed:
Self-verification: Checking answers within `` tags
Long chain-of-thought: Extended reasoning traces
Backtracking: "Wait, that seems wrong..." behavior
Metacognition:… See the full description on the dataset page: https://huggingface.co/datasets/cxzsad12e/verl-aha-moment-dataset.
