david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples
FIneMix Dataset: This is a 16,278,528 token, 15,897 text sample, 1094 sequence length (padded) subset of the curriculum used to train codelion/gpt-2-70m. It is designed for smoke testing, hyperparameter optimization, and baseline comparison on novel small language model architectures before scaling up and burning more resources. Data Sources and selection: 50% - FinePDFs (500M tokens): High-quality PDF content: ~ 7948 rows selected from… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/codelion-finemix-pdf-dclm-edu-1024-seq-len-15897-samples.
010
FIneMix Dataset:
- This is a 16,278,528 token, 15,897 text sample, 1094 sequence length (padded) subset of the curriculum used to train codelion/gpt-2-70m.
- It is designed for smoke testing, hyperparameter optimization, and baseline comparison on novel small language model architectures before scaling up and burning more resources.
Data Sources and selection:
- 50% - FinePDFs (500M tokens): High-quality PDF content: ~ 7948 rows selected from https://huggingface.co/datasets/codelion/finepdfs-100M
- 30% - DCLM Baseline (300M tokens): Filtered web content: ~ 4770 rows from https://huggingface.co/datasets/codelion/dclm-baseline-100M
- 20% - FineWeb-Edu (200M tokens): Educational web content: ~ 7948 rows https://huggingface.co/datasets/codelion/fineweb-edu-100M
Selection criteria:
- Select samples < 1094 samples
- Quasi - Random selection (pandas.DataFrame.sample()) of the prescribed number of rows to meet the ratio recommended in https://huggingface.co/blog/codelion/optimal-dataset-mixing
