CoolFace
Modelpublic

farsipal/whisper-sm-el-intlv-xs

sourceHugging Faceapache-2.0updated 4y agoView on Hugging Face
0likes15downloads
Model Card

Whisper small (Greek) Trained on Interleaved Datasets

This model is a fine-tuned version of openai/whisper-small on interleaved mozilla-foundation/commonvoice110 (el) and google/fleurs (elgr) dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.4741
  • —Wer: 20.0687

Model description

The model was developed during the Whisper Fine-Tuning Event in December 2022. More details on the model can be found in the original paper

Intended uses & limitations

The model is fine-tuned for transcription in the Greek language.

Training and evaluation data

This model was trained by interleaving the training and evaluation splits from two different datasets:

  • —mozilla-foundation/commonvoice11_0 (el)
  • —google/fleurs (el_gr)

Training procedure

The python script used is a modified version of the script provided by Hugging Face and can be found here

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 1e-05
  • —trainbatchsize: 16
  • —evalbatchsize: 8
  • —seed: 42
  • —gradientaccumulationsteps: 4
  • —totaltrainbatch_size: 64
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —lrschedulerwarmup_steps: 500
  • —training_steps: 5000
  • —mixedprecisiontraining: Native AMP

Training results

Training LossEpochStepValidation LossWer
0.01864.9810000.361921.0067
0.00129.9520000.434720.3009
0.000514.9330000.474120.0687
0.000319.940000.497420.1152
0.000324.8850000.506620.2266

Framework versions

  • —Transformers 4.26.0.dev0
  • —Pytorch 1.13.0
  • —Datasets 2.7.1.dev0
  • —Tokenizers 0.12.1