CoolFace
Datasetpublic

HexQuant/smoltalk

SmolTalk Dataset description This is a synthetic dataset designed for supervised finetuning (SFT) of LLMs. It was used to build SmolLM2-Instruct family of models and contains 1M samples. More details in our paper https://arxiv.org/abs/2502.02737 During the development of SmolLM2, we observed that models finetuned on public SFT datasets underperformed compared to other models with proprietary instruction datasets. To address this gap, we created new synthetic… See the full description on the dataset page: https://huggingface.co/datasets/HexQuant/smoltalk.

sourceHugging Faceupdated 9mo agoView on Hugging Face
0likes235downloads
1 commits on main
c964cf69mo ago

Duplicate from HuggingFaceTB/smoltalk

HexQuant, loubnabnl