CoolFace
Datasetpublic

sastpg/CoVo_Dataset

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning The file test.jsonl is used to monitor the model performance as training proceeds. Level Source 1 GSM8K 2 MATH-500 3 AMC-23 4 AIME2024 5 AIME2025

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes101downloads
Dataset Card

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

The file test.jsonl is used to monitor the model performance as training proceeds.

LevelSource
1GSM8K
2MATH-500
3AMC-23
4AIME2024
5AIME2025