CoolFace
Datasetpublic

david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples

FIneMix Dataset: This is a 16,278,528 token, 15,897 text sample, 1094 sequence length (padded) subset of the curriculum used to train codelion/gpt-2-70m. It is designed for smoke testing, hyperparameter optimization, and baseline comparison on novel small language model architectures before scaling up and burning more resources. Data Sources and selection: 50% - FinePDFs (500M tokens): High-quality PDF content: ~ 7948 rows selected from… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes10downloads
Dataset Card

FIneMix Dataset:

  • —This is a 16,278,528 token, 15,897 text sample, 1094 sequence length (padded) subset of the curriculum used to train codelion/gpt-2-70m.
  • —It is designed for smoke testing, hyperparameter optimization, and baseline comparison on novel small language model architectures before scaling up and burning more resources.

Data Sources and selection:

  • —50% - FinePDFs (500M tokens): High-quality PDF content: ~ 7948 rows selected from https://huggingface.co/datasets/codelion/finepdfs-100M
  • —30% - DCLM Baseline (300M tokens): Filtered web content: ~ 4770 rows from https://huggingface.co/datasets/codelion/dclm-baseline-100M
  • —20% - FineWeb-Edu (200M tokens): Educational web content: ~ 7948 rows https://huggingface.co/datasets/codelion/fineweb-edu-100M

Selection criteria:

  • —Select samples < 1094 samples
  • —Quasi - Random selection (pandas.DataFrame.sample()) of the prescribed number of rows to meet the ratio recommended in https://huggingface.co/blog/codelion/optimal-dataset-mixing