CoolFace
Datasetpublic

t2ance/mat-01-lr-versus-learning-signal

01 Learning rate versus the learning signal When the round-best score on a CPU-only Kaggle task stops rising during GRPO training, is the binding constraint the learning rate (too small to move the policy, or too large to keep it stable), or the learning signal itself (what the search samples and how the reward separates it)? This repository is the data root of that question: every training run's tree-search rollout archive it produced between 2026-06-29 and 2026-07-03, minus… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/mat-01-lr-versus-learning-signal.

sourceHugging Faceupdated 22d agoView on Hugging Face
0likes382downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
t2ance/mat-01-lr-versus-learning-signal · CoolFace