CoolFace
Modelpublic

anuragshas/wav2vec2-xls-r-300m-lv-cv8-with-lm

sourceHugging Faceapache-2.0updated 5y agoView on Hugging Face
0likes105downloads
Model Card

<!-- This model card has been generated automatically according to the information the Trainer had access to. You should probably proofread and complete it, then remove this comment. -->

XLS-R-300M - Latvian

This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m on the MOZILLA-FOUNDATION/COMMONVOICE8_0 - LV dataset. It achieves the following results on the evaluation set:

  • —Loss: 0.1660
  • —Wer: 0.1705

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 7.5e-05
  • —trainbatchsize: 32
  • —evalbatchsize: 16
  • —seed: 42
  • —optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • —lrschedulertype: linear
  • —lrschedulerwarmup_steps: 1000
  • —num_epochs: 50.0
  • —mixedprecisiontraining: Native AMP

Training results

Training LossEpochStepValidation LossWer
3.4892.564003.35901.0
2.99035.138002.97041.0001
1.67127.6912000.61790.6566
1.263510.2616000.31760.4531
1.081912.8220000.25170.3508
1.013615.3824000.22570.3124
0.962517.9528000.19750.2311
0.90120.5132000.19860.2097
0.884223.0836000.19040.2039
0.854225.6440000.18470.1981
0.824428.2144000.18050.1847
0.768930.7748000.17360.1832
0.782533.3352000.16980.1821
0.781735.956000.17580.1803
0.748838.4660000.16630.1760
0.717141.0364000.16360.1721
0.722243.5968000.16630.1729
0.715646.1572000.16330.1715
0.712148.7276000.16660.1718

Framework versions

  • —Transformers 4.17.0.dev0
  • —Pytorch 1.10.2+cu102
  • —Datasets 1.18.2.dev0
  • —Tokenizers 0.11.0
Evaluation Commands
  1. 1.To evaluate on mozilla-foundation/common_voice_8_0 with split test
bash
python eval.py --model_id anuragshas/wav2vec2-xls-r-300m-lv-cv8-with-lm --dataset mozilla-foundation/common_voice_8_0 --config lv --split test
  1. 1.To evaluate on speech-recognition-community-v2/dev_data
bash
python eval.py --model_id anuragshas/wav2vec2-xls-r-300m-lv-cv8-with-lm --dataset speech-recognition-community-v2/dev_data --config lv --split validation --chunk_length_s 5.0 --stride_length_s 1.0

Inference With LM

python
import torch
from datasets import load_dataset
from transformers import AutoModelForCTC, AutoProcessor
import torchaudio.functional as F
model_id = "anuragshas/wav2vec2-xls-r-300m-lv-cv8-with-lm"
sample_iter = iter(load_dataset("mozilla-foundation/common_voice_8_0", "lv", split="test", streaming=True, use_auth_token=True))
sample = next(sample_iter)
resampled_audio = F.resample(torch.tensor(sample["audio"]["array"]), 48_000, 16_000).numpy()
model = AutoModelForCTC.from_pretrained(model_id)
processor = AutoProcessor.from_pretrained(model_id)
input_values = processor(resampled_audio, return_tensors="pt").input_values
with torch.no_grad():
    logits = model(input_values).logits
transcription = processor.batch_decode(logits.numpy()).text
# => "domāju ka viņam viss labi"

Eval results on Common Voice 8 "test" (WER):

Without LMWith LM (run `./eval.py`)
16.9979.633