jdpressman/comma_v0.1_training_dataset_sample_1B
Comma v0.1 Training Dataset (1 Billion Token Sample) This is a 1 billion token subset of the Comma v0.1 Training Set intended as a convenience for small deep learning experiments. It is similar in spirit to the 1 billion token RedPajama sample which is no longer functioning with HuggingFace transformers due to involving the execution of arbitrary code at load time. Method The subset was created using a single item batch version of the following script which I no… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/comma_v0.1_training_dataset_sample_1B.
014
