cxzsad12e/verl-aha-moment-dataset
Verl Dataset for DeepSeek-R1 "Aha Moment" Reproduction This dataset is prepared for training with the Verl framework to reproduce the "aha moment" phenomenon observed in DeepSeek-R1-Zero. What is the "Aha Moment"? During RL training without any SFT, DeepSeek-R1-Zero spontaneously developed: Self-verification: Checking answers within `` tags Long chain-of-thought: Extended reasoning traces Backtracking: "Wait, that seems wrong..." behavior Metacognition:… See the full description on the dataset page: https://huggingface.co/datasets/cxzsad12e/verl-aha-moment-dataset.
Upload verl_deepscaler_test.parquet with huggingface_hub
Upload verl_deepscaler.parquet with huggingface_hub
Upload README.md with huggingface_hub
Upload verl_deepscaler_test.parquet with huggingface_hub
Upload verl_deepscaler.parquet with huggingface_hub
Upload README.md with huggingface_hub
Upload verl_deepscaler_test.parquet with huggingface_hub
Upload verl_deepscaler.parquet with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload train_math.arrow with huggingface_hub
Upload train_gsm8k.arrow with huggingface_hub
Upload test.arrow with huggingface_hub
Upload train.arrow with huggingface_hub
Upload test_math.parquet with huggingface_hub
Upload train_math.parquet with huggingface_hub
Upload test_gsm8k.parquet with huggingface_hub
Upload test.parquet with huggingface_hub
Upload README.md with huggingface_hub
Upload train.parquet with huggingface_hub
Upload test_math.arrow with huggingface_hub
Upload train_gsm8k.parquet with huggingface_hub
Upload test_gsm8k.arrow with huggingface_hub
initial commit
