sastpg/CoVo_Dataset
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning The file test.jsonl is used to monitor the model performance as training proceeds. Level Source 1 GSM8K 2 MATH-500 3 AMC-23 4 AIME2024 5 AIME2025
075
1---2license: mit3---4 5# Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning6 7The file `test.jsonl` is used to monitor the model performance as training proceeds.8 9 10|Level |Source |11|--------|--------|12|1 |GSM8K |13|2 |MATH-500|14|3 |AMC-23 |15|4 |AIME2024|16|5 |AIME2025|17 