ulmentflam/gpt2-774m-fineweb-mojo
023
v2 re-run: trained biases restored (dbias scratch overrun fixed)
GPT-2 774M FineWeb-10B from-scratch llm.mojo run (val loss 3.0130, HellaSwag 36.34%)
initial commit
v2 re-run: trained biases restored (dbias scratch overrun fixed)
GPT-2 774M FineWeb-10B from-scratch llm.mojo run (val loss 3.0130, HellaSwag 36.34%)
initial commit