CoolFace
Modelpublic

Mxode/NanoLM-365M-Base

sourceHugging Facegpl-3.0updated 2y agoView on Hugging Face
0likes292downloads
Model Card

NanoLM-365M-base

English | 简体中文

Introduction

Based on Qwen2-0.5B, the tokenizer has been replaced with BilingualTokenizer-8K to reduce the number of parameters. The total parameters have been reduced from 0.5B to 365M.

Details

To recover some performance and facilitate fine-tuning for downstream tasks, I chose to freeze the backbone parameters and only train the embedding part after replacing the tokenizer. Training was conducted for 40,000 steps on wikipedia-zh and cosmopedia-100k.

Value
Total Params365 M
Trainable Params< 10 M
Trainable Partsmodel.embed_tokens
Training Steps40,000
Training Datasetwikipedia-zh, cosmopedia-100k
Optimizeradamw_torch
Learning Rate2e-4
LR Schedulercosine
Weight Decay0.1
Warm-up Ratio0.03
Batch Size16
Gradient Accumulation Steps1
Seq Len4096
Dtypebf16
Peak GPU Memory< 48 GB
DeviceNVIDIA A100-SXM4-80GB

The specific training records are as follows: [image]