JackyHoCL/whisper-large-v3-turbo-cantonese-yue-english
A version with noise detection is trained base on this model, to reduce hallucination during streaming:<br>
Name: JackyHoCL/whisper-large-v3-turbo-cantonese-noise-detection<br> https://huggingface.co/JackyHoCL/whisper-large-v3-turbo-cantonese-noise-detection <br/> <br/> transformers-4.49.0 <br/> For Cantonese + English, use 'yue', for Cantonese + Mandarin + English, use 'zh' <br/> --------------------------------------------------------------- TODO: <br/> 1.Improve zh-CN performance <br/> 2.Improve overall performance (yue+zh+en) with background noise (Please kindly suggest/provide dataset if possible, thx) <br/>
temperature=0.3, extra_body=dict( seed=4419, repetition_penalty=1.05, top_p=0.5 )
2025-07-21: CER: | Dataset | Lang | Split | CER(in %) | | -------- | ------- | ------- | ------- | |Training|yue|validation|8.05| |mozilla-foundation/commonvoice170|yue|test|**0.64**| |JackyHoCL/cleanedmixedcantoneseandenglishspeech|yue|test|8.3| |mozilla-foundation/commonvoice170|en|test(2k samples)|5.22| |mozilla-foundation/commonvoice161|zh-CN|test|11.89|
2025-07-19: CER: | Dataset | Lang | Split | CER(in %) | | -------- | ------- | ------- | ------- | |Training|yue|validation|8.94| |mozilla-foundation/commonvoice170|yue|test|1.29| |JackyHoCL/cleanedmixedcantoneseandenglishspeech|yue|test|8.00| |mozilla-foundation/commonvoice170|en|test|6.8| |mozilla-foundation/commonvoice161|zh-CN|test|50.9|
2025-07-06: CER: | Dataset | Lang | Split | CER(in %) | | -------- | ------- | ------- | ------- | |Training|yue|validation|8.92| |mozilla-foundation/commonvoice170|yue|test|8.86| |JackyHoCL/cleanedmixedcantoneseandenglishspeech|yue|test|7.96| |mozilla-foundation/commonvoice170|en|test|6.84| |mozilla-foundation/commonvoice161|zh-CN|test|43.0|
perdevicetrainbatchsize=32,<br/> learning_rate=1e-7,<br/>
2025-07-03: CER: | Dataset | Lang | Split | CER(in %) | | -------- | ------- | ------- | ------- | |Training|yue|validation|9.705| |mozilla-foundation/commonvoice170|yue|test|9.31| |JackyHoCL/cleanedmixedcantoneseandenglishspeech|yue|test|8.37|
perdevicetrainbatchsize=32,<br/> learning_rate=1e-5,<br/>
CER: 13.7% <br/>
Train Args:<br/> perdevicetrainbatchsize=16,<br/> gradientaccumulationsteps=1,<br/> learningrate=1e-5,<br/> gradientcheckpointing=True,<br/> perdeviceevalbatchsize=16,<br/> generationmaxlength=225,<br/>
Hardware:<br/> NVIDIA Tesla V100 16GB * 4<br/>
A Realtime Streaming application example is built on this model:<br/> https://github.com/JackyHoCL/whisper-realtime.git <br/>
FAQ:
- If having tokenizer issue during inference, please update your transformers version to >= 4.49.0
pip install --upgrade transformers