CoolFace
Datasetpublic

t2ance/mat-01-lr-versus-learning-signal

01 Learning rate versus the learning signal When the round-best score on a CPU-only Kaggle task stops rising during GRPO training, is the binding constraint the learning rate (too small to move the policy, or too large to keep it stable), or the learning signal itself (what the search samples and how the reward separates it)? This repository is the data root of that question: every training run's tree-search rollout archive it produced between 2026-06-29 and 2026-07-03, minus… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/mat-01-lr-versus-learning-signal.

sourceHugging Faceupdated 22d agoView on Hugging Face
0likes382downloads
8 commits on main
f1d55ad22d ago

Publish 150 files of 01-lr-versus-learning-signal (350/350) (part 3)

t2ance
5f9572d22d ago

Publish 150 files of 01-lr-versus-learning-signal (350/350) (part 2)

t2ance
026561e22d ago

Publish 150 files of 01-lr-versus-learning-signal (350/350)

t2ance
6f7939c22d ago

Publish 200 files of 01-lr-versus-learning-signal (200/350) (part 4)

t2ance
f7cd1eb22d ago

Publish 200 files of 01-lr-versus-learning-signal (200/350) (part 3)

t2ance
35f70e822d ago

Publish 200 files of 01-lr-versus-learning-signal (200/350) (part 2)

t2ance
7a4afcd22d ago

Publish 200 files of 01-lr-versus-learning-signal (200/350)

t2ance
dda8c1e22d ago

initial commit

t2ance