CoolFace
Datasetpublic

MiniLLM/pile-diff_samp-qwen_1.8B-qwen_104M-r0.5

This repository contains the refined pre-training corpus from the paper MiniPLM: Knowledge Distillation for Pre-Training Language Models. Code: https://github.com/thu-coai/MiniPLM

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes275downloads
Dataset Card

This repository contains the refined pre-training corpus from the paper MiniPLM: Knowledge Distillation for Pre-Training Language Models.

Code: https://github.com/thu-coai/MiniPLM