CoolFace
Modelpublic

JackyHoCL/whisper-large-v3-turbo-cantonese-yue-english

sourceHugging Facemitupdated 1y agoView on Hugging Face
14likes354downloads
Model Card

A version with noise detection is trained base on this model, to reduce hallucination during streaming:<br>

Name: JackyHoCL/whisper-large-v3-turbo-cantonese-noise-detection<br> https://huggingface.co/JackyHoCL/whisper-large-v3-turbo-cantonese-noise-detection <br/> <br/> transformers-4.49.0 <br/> For Cantonese + English, use 'yue', for Cantonese + Mandarin + English, use 'zh' <br/> --------------------------------------------------------------- TODO: <br/> 1.Improve zh-CN performance <br/> 2.Improve overall performance (yue+zh+en) with background noise (Please kindly suggest/provide dataset if possible, thx) <br/>

temperature=0.3, extra_body=dict( seed=4419, repetition_penalty=1.05, top_p=0.5 )

2025-07-21: CER: | Dataset | Lang | Split | CER(in %) | | -------- | ------- | ------- | ------- | |Training|yue|validation|8.05| |mozilla-foundation/commonvoice170|yue|test|**0.64**| |JackyHoCL/cleanedmixedcantoneseandenglishspeech|yue|test|8.3| |mozilla-foundation/commonvoice170|en|test(2k samples)|5.22| |mozilla-foundation/commonvoice161|zh-CN|test|11.89|

2025-07-19: CER: | Dataset | Lang | Split | CER(in %) | | -------- | ------- | ------- | ------- | |Training|yue|validation|8.94| |mozilla-foundation/commonvoice170|yue|test|1.29| |JackyHoCL/cleanedmixedcantoneseandenglishspeech|yue|test|8.00| |mozilla-foundation/commonvoice170|en|test|6.8| |mozilla-foundation/commonvoice161|zh-CN|test|50.9|

2025-07-06: CER: | Dataset | Lang | Split | CER(in %) | | -------- | ------- | ------- | ------- | |Training|yue|validation|8.92| |mozilla-foundation/commonvoice170|yue|test|8.86| |JackyHoCL/cleanedmixedcantoneseandenglishspeech|yue|test|7.96| |mozilla-foundation/commonvoice170|en|test|6.84| |mozilla-foundation/commonvoice161|zh-CN|test|43.0|

perdevicetrainbatchsize=32,<br/> learning_rate=1e-7,<br/>


2025-07-03: CER: | Dataset | Lang | Split | CER(in %) | | -------- | ------- | ------- | ------- | |Training|yue|validation|9.705| |mozilla-foundation/commonvoice170|yue|test|9.31| |JackyHoCL/cleanedmixedcantoneseandenglishspeech|yue|test|8.37|

perdevicetrainbatchsize=32,<br/> learning_rate=1e-5,<br/>


CER: 13.7% <br/>

Train Args:<br/> perdevicetrainbatchsize=16,<br/> gradientaccumulationsteps=1,<br/> learningrate=1e-5,<br/> gradientcheckpointing=True,<br/> perdeviceevalbatchsize=16,<br/> generationmaxlength=225,<br/>

Hardware:<br/> NVIDIA Tesla V100 16GB * 4<br/>

A Realtime Streaming application example is built on this model:<br/> https://github.com/JackyHoCL/whisper-realtime.git <br/>

FAQ:

  1. 1.If having tokenizer issue during inference, please update your transformers version to >= 4.49.0
bash
pip install --upgrade transformers