CoolFace
Datasetpublic

david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples

FIneMix Dataset: This is a 16,278,528 token, 15,897 text sample, 1094 sequence length (padded) subset of the curriculum used to train codelion/gpt-2-70m. It is designed for smoke testing, hyperparameter optimization, and baseline comparison on novel small language model architectures before scaling up and burning more resources. Data Sources and selection: 50% - FinePDFs (500M tokens): High-quality PDF content: ~ 7948 rows selected from… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes10downloads
5 commits on main
e60a3c18mo ago

Upload dataset

david-thrower
4f551708mo ago

Update README.md

david-thrower
364719c8mo ago

Update README.md

david-thrower
9eb3df38mo ago

Upload dataset

david-thrower
c4d78a28mo ago

initial commit

david-thrower