CoolFace
Datasetpublic

david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples

FIneMix Dataset: This is a 16,278,528 token, 15,897 text sample, 1094 sequence length (padded) subset of the curriculum used to train codelion/gpt-2-70m. It is designed for smoke testing, hyperparameter optimization, and baseline comparison on novel small language model architectures before scaling up and burning more resources. Data Sources and selection: 50% - FinePDFs (500M tokens): High-quality PDF content: ~ 7948 rows selected from… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes10downloads

david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples · main · files are served by the source, never re-hosted here