CoolFace
Modelpublic

efficient-speech/lite-whisper-large-v3

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
8likes59downloads
Model Card

Model Card for Lite-Whisper large-v3

<!-- Provide a quick summary of what the model is/does. -->

Lite-Whisper is a compressed version of OpenAI Whisper with LiteASR. See our GitHub repository and paper for details.

Here's a code snippet to get started:

python
import librosa 
import torch
from transformers import AutoProcessor, AutoModel

device = "cuda:0"
dtype = torch.float16

# load the compressed Whisper model
model = AutoModel.from_pretrained(
    "efficient-speech/lite-whisper-large-v3-turbo", 
    trust_remote_code=True, 
)
model.to(dtype).to(device)

# we use the same processor as the original model
processor = AutoProcessor.from_pretrained("openai/whisper-large-v3")

# set the path to your audio file
path = "path/to/audio.wav"
audio, _ = librosa.load(path, sr=16000)

input_features = processor(audio, sampling_rate=16000, return_tensors="pt").input_features
input_features = input_features.to(dtype).to(device)

predicted_ids = model.generate(input_features)
transcription = processor.batch_decode(
    predicted_ids, 
    skip_special_tokens=True
)[0]

print(transcription)

Benchmark Results

Following is the average word error rate (WER) evaluated on the ESB datasets:

ModelAverage WER (↓)Encoder SizeDecoder Size
whisper-large-v310.1635M907M
lite-whisper-large-v3-acc10.1429M907M
lite-whisper-large-v310.2377M907M
lite-whisper-large-v3-fast11.3308M907M
&nbsp;&nbsp;&nbsp;&nbsp;
whisper-large-v3-turbo10.1635M172M
lite-whisper-large-v3-turbo-acc10.2421M172M
lite-whisper-large-v3-turbo12.6374M172M
lite-whisper-large-v3-turbo-fast20.1313M172M
&nbsp;&nbsp;&nbsp;&nbsp;
whisper-medium14.8306M457M