CoolFace
20 results

mini-llm

credi-net /CDB_DEC2024-CochranSampled_MiniLLMV2_Emb CrediBench Web Content Embeddings (December 2024) This repository contains MiniLLM_MultiLingual_V2 embeddings for the CrediBench WebContent (December 2024) dataset. Source dataset:https://huggingface.co/datasets/Hussein-Abdallah/CrediBench-WebContent-Dec2024_CochranSampled Overview The source dataset is a Cochran-sampled subset of the December 2024 CrediBench WebContent corpus. Documents are sampled independently for each web domain using Cochran's sampling… See the full description on the dataset page: https://huggingface.co/datasets/credi-net/CDB_DEC2024-CochranSampled_MiniLLMV2_Emb.10M<n<100M1 likes1k downloads2mo agoHugging FaceMiniLLM /pile-diff_samp-qwen_1.8B-qwen_104M-r0.5This repository contains the refined pre-training corpus from the paper MiniPLM: Knowledge Distillation for Pre-Training Language Models. Code: https://github.com/thu-coai/MiniPLM text-generation0 likes562 downloads1y agoHugging FaceMiniLLM /sinst sinst dataset Original version: Super-Natural-Instructions This dataset is used to evaluate MiniLLM. text1K<n<10K1 likes490 downloads2y agoHugging FaceMiniLLM /dolly Dolly Dataset Original version: aisquared/databricks-dolly-15k. This dataset is used to train MiniLLM. textn<1K0 likes430 downloads2y agoHugging FaceMiniLLM /dolly-processedtext100K<n<1M1 likes377 downloads2y agoHugging FaceMiniLLM /uinst uinst dataset Original version: Unnatural-Instructions This dataset is used to evaluate MiniLLM text10K<n<100K1 likes364 downloads2y agoHugging Face