CoolFace
Datasetpublic

david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming

A corpus of high quality fine tuning data meant for fine tuning various HelixLM models Dataset Composition: A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ... Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning. Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.

sourceHugging Faceodc-byupdated 4mo agoView on Hugging Face
0likes173downloads
3 commits on main
af33e874mo ago

Update README.md

david-thrower
46aab054mo ago

Upload dataset

david-thrower
76a8a3f4mo ago

initial commit

david-thrower