CoolFace
Modelpublic

aahouzi/whisper-npu

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
1likes
Model Card

Whisper-medium OpenVINO IR

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

Whisper was proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al from OpenAI. The original code repository can be found here.

Disclaimer: Content for this model card has partly been copied and pasted from this model card.

Model details

Whisper is a Transformer based encoder-decoder model, also referred to as a sequence-to-sequence model.

[image]

Model TypeParametersn_audio_ctxn_audio_staten_audio_headn_audio_layern_text_ctxn_text_staten_text_headn_text_layern_melsn_vocab
whisper-tiny39 M150038464224384648051865
whisper-base74 M150051286224512868051865
whisper-small244 M1500768121222476812128051865
whisper-medium769 M150010241624224102416168051865
whisper-large-v11550 M150012802032224128020208051865
whisper-large-v21550 M150012802032224128020208051865
distil-whisper-large-v2756 M15001280203222412802028051865
whisper-large-v31550 M1500128020322241280202012851866
distil-whisper-large-v3756 M150012802032224128020212851866
whisper-large-v3-turbo809 M150012802032224128020412851866