sastpg/CoVo_Dataset
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning The file test.jsonl is used to monitor the model performance as training proceeds. Level Source 1 GSM8K 2 MATH-500 3 AMC-23 4 AIME2024 5 AIME2025
0100
This repository belongs to sastpg on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
CoVo_Dataset
public
mit
no
sastpg
