CoolFace
Datasetpublic

jdpressman/comma_v0.1_training_dataset_sample_1B

Comma v0.1 Training Dataset (1 Billion Token Sample) This is a 1 billion token subset of the Comma v0.1 Training Set intended as a convenience for small deep learning experiments. It is similar in spirit to the 1 billion token RedPajama sample which is no longer functioning with HuggingFace transformers due to involving the execution of arbitrary code at load time. Method The subset was created using a single item batch version of the following script which I no… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/comma_v0.1_training_dataset_sample_1B.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes14downloads
filetrain-00000-of-00004.jsonl.gz271.3 MBdownload
filetrain-00001-of-00004.jsonl.gz275.7 MBdownload
filetrain-00002-of-00004.jsonl.gz276.1 MBdownload
filetrain-00003-of-00004.jsonl.gz271.9 MBdownload

jdpressman/comma_v0.1_training_dataset_sample_1B · main · files are served by the source, never re-hosted here