CoolFace
Datasetpublic

lenamerkli/distilled-web

Dataset Card for lenamerkli/distilled-web This dataset consists of web-scraped data using a custom crawler purpose-built for each website. Dataset Details Dataset Sources Repository: https://github.com/lenamerkli/distilled-web Uses This dataset is useful for training large language models. The train split provides instruction-following and chat data for supervised fine-tuning (SFT) and instruction tuning. The pretrain split… See the full description on the dataset page: https://huggingface.co/datasets/lenamerkli/distilled-web.

sourceHugging Faceupdated 7d agoView on Hugging Face
3likes1.1kdownloads

lenamerkli/distilled-web · main · files are served by the source, never re-hosted here