jdpressman/comma_v0.1_training_dataset_sample_1B
Comma v0.1 Training Dataset (1 Billion Token Sample) This is a 1 billion token subset of the Comma v0.1 Training Set intended as a convenience for small deep learning experiments. It is similar in spirit to the 1 billion token RedPajama sample which is no longer functioning with HuggingFace transformers due to involving the execution of arbitrary code at load time. Method The subset was created using a single item batch version of the following script which I no… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/comma_v0.1_training_dataset_sample_1B.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face