CoolFace
Modelpublic

SakshiRathi77/wav2vec2-large-xlsr-300m-hi-kagglex

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
0likes12downloads
Model Card

datasets:

  • —mozilla-foundation/commonvoice15_0
  • —mozilla-foundation/commonvoice13_0 language:
  • —hi metrics:
  • —cer
  • —wer libraryname: transformers pipelinetag: automatic-speech-recognition model-index:
  • —name: whisper-small-hi-cv results:
  • —task: name: Automatic Speech Recognition type: automatic-speech-recognition dataset: name: Common Voice 15 type: mozilla-foundation/commonvoice15_0 args: hi metrics:
  • —name: Test WER type: wer value: 13.9913
  • —name: Test CER type: cer value: 5.8844
  • —task: name: Automatic Speech Recognition type: automatic-speech-recognition dataset: name: Common Voice 13 type: mozilla-foundation/commonvoice13_0 args: hi metrics:
  • —name: Test WER type: wer value: 23.3824
  • —name: Test CER type: cer value: 10.5288

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

Model Details

This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m on this dataset . It achieves the following results on the evaluation set:

  • —Loss: 0.3691
  • —Wer: 0.3285
  • —Cer: 0.0875

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 0.0001
  • —trainbatchsize: 32
  • —evalbatchsize: 8
  • —seed: 42
  • —gradientaccumulationsteps: 4
  • —totaltrainbatch_size: 128
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —lrschedulerwarmup_steps: 300
  • —num_epochs: 100

Training results

Training LossEpochStepValidation LossWerCer
7.31419.053003.46611.01.0
2.569838.16000.65770.52030.1466
0.611257.149000.40480.37230.1005
0.382676.1912000.37780.33860.0901
0.316895.2415000.36910.32850.0875

Framework versions

  • —Transformers 4.33.0
  • —Pytorch 2.0.0
  • —Datasets 2.1.0
  • —Tokenizers 0.13.3

SPACE

Automatic Speech Recognization in hindi