CoolFace
Modelpublic

RuiqianLi/malaya-speech_Mrbrown_finetune1

sourceHugging Faceupdated 4y agoView on Hugging Face
0likes19downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

malaya-speechMrbrownfinetune1

This model is a fine-tuned version of malay-huggingface/wav2vec2-xls-r-300m-mixed on the uob_singlish dataset.

This time use self-made dataset(cut the audio of "https://www.youtube.com/watch?v=a2ZOTD3R7JI" into slices and write the corresponding transcript, totally 4 mins), get really bad fine-tuning result, that may mean the training/fine-tuning dataset must be high quality/at least several hours? Or maybe is because the learning rate is set too high(0.01) ? Still searching for the important factors.

It achieves the following results on the evaluation set:

  • —Loss: 3.8458
  • —Wer: 1.01

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 0.01
  • —trainbatchsize: 2
  • —evalbatchsize: 8
  • —seed: 42
  • —gradientaccumulationsteps: 2
  • —totaltrainbatch_size: 4
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —lrschedulerwarmup_steps: 500
  • —num_epochs: 100
  • —mixedprecisiontraining: Native AMP

Training results

Training LossEpochStepValidation LossWer
0.318620.02004.22251.13
0.491140.04004.04270.99
0.901460.06005.32851.04
1.095580.08003.69221.02
0.7533100.010003.84581.01

Framework versions

  • —Transformers 4.11.3
  • —Pytorch 1.10.0+cu113
  • —Datasets 1.18.3
  • —Tokenizers 0.10.3