CoolFace
Datasetpublic

sastpg/CoVo_Dataset

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning The file test.jsonl is used to monitor the model performance as training proceeds. Level Source 1 GSM8K 2 MATH-500 3 AMC-23 4 AIME2024 5 AIME2025

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes75downloads
README.md17 linesDownload Raw Back to root
1---2license: mit3---4 5# Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning6 7The file `test.jsonl` is used to monitor the model performance as training proceeds.8 9 10|Level   |Source  |11|--------|--------|12|1       |GSM8K   |13|2       |MATH-500|14|3       |AMC-23  |15|4       |AIME2024|16|5       |AIME2025|17