CoolFace
Modelpublic

stvlynn/Qwen-7B-Chat-Cantonese

sourceHugging Faceagpl-3.0updated 2y agoView on Hugging Face
23likes26downloads
Model Card

Qwen-7B-Chat-Cantonese (通义千问·粤语)

Intro

Qwen-7B-Chat-Cantonese is a fine-tuned version based on Qwen-7B-Chat, trained on a substantial amount of Cantonese language data.

Qwen-7B-Chat-Cantonese係基於Qwen-7B-Chat嘅微調版本,基於大量粵語數據進行訓練。

ModelScope(魔搭社区)

Usage

Requirements

  • —python 3.8 and above
  • —pytorch 1.12 and above, 2.0 and above are recommended
  • —CUDA 11.4 and above are recommended (this is for GPU users, flash-attention users, etc.)

Dependency

To run Qwen-7B-Chat-Cantonese, please make sure you meet the above requirements, and then execute the following pip commands to install the dependent libraries.

bash
pip install transformers==4.32.0 accelerate tiktoken einops scipy transformers_stream_generator==0.0.4 peft deepspeed

In addition, it is recommended to install the flash-attention library (we support flash attention 2 now.) for higher efficiency and lower memory usage.

bash
git clone https://github.com/Dao-AILab/flash-attention
cd flash-attention && pip install .

Quickstart

Pls turn to QwenLM/Qwen - Quickstart

Training Parameters

ParameterDescriptionValue
Learning RateAdamW optimizer learning rate7e-5
Weight DecayRegularization strength0.8
GammaLearning rate decay factor1.0
Batch SizeNumber of samples per batch1000
PrecisionFloating point precisionfp16
Learning PolicyLearning rate adjustment policycosine
Warmup StepsInitial steps without learning rate adjustment0
Total StepsTotal training steps1024
Gradient Accumulation StepsNumber of steps to accumulate gradients before updating8

loss

Demo

深水埗有哪些美食

鲁迅为什么打周树人

树上几只鸟

Special Note

This is my first fine-tuning LLM project. Pls forgive me if there's anything wrong.

If you have any questions or suggestions, feel free to contact me.

Twitter @stv_lynn

Telegram @stvlynn

email i@stv.pm