joaorura/distil-whisper-large-v3-ptbr-openvino
distil-whisper-large-v3-ptbr-openvino (FP32)
English
OpenVINO IR (FP32) conversion of freds0/distil-whisper-large-v3-ptbr, a Portuguese (pt-BR) fine-tune of distil-whisper/distil-large-v3 by freds0. All credit for the underlying model weights and training goes to the original author — this repository only provides an OpenVINO-optimized export for CPU/GPU/NPU inference on Intel hardware.
Conversion
Converted with optimum-intel 2.2.0, using openvino 2026.4.0 and openvino-genai 2026.4.0.0. The optimum-cli export openvino CLI command failed due to a version incompatibility with transformers 5.5.4, so the conversion was done directly through the optimum-intel Python API instead:
from optimum.intel.openvino import OVModelForSpeechSeq2Seq
from transformers import AutoProcessor
model_id = "freds0/distil-whisper-large-v3-ptbr"
model = OVModelForSpeechSeq2Seq.from_pretrained(
model_id,
export=True,
)
processor = AutoProcessor.from_pretrained(model_id)
model.save_pretrained("distil-whisper-large-v3-ptbr-ov")
processor.save_pretrained("distil-whisper-large-v3-ptbr-ov")The exported model uses a single merged decoder graph (openvino_decoder_model.xml/.bin, with use_cache: true in config.json) rather than separate decoder/decoder-with-past graphs, plus the OpenVINO tokenizer/detokenizer pair (openvino_tokenizer.*, openvino_detokenizer.*) generated for use with openvino_genai.
Usage with openvino_genai.WhisperPipeline
import openvino_genai as ov_genai
device = "CPU" # or "GPU", "NPU"
pipeline_config = {}
if device == "NPU":
# Avoids recompiling the model on every NPU run.
pipeline_config["CACHE_DIR"] = "./npu_cache"
pipe = ov_genai.WhisperPipeline("distil-whisper-large-v3-ptbr-ov", device=device, **pipeline_config)
gen_config = pipe.get_generation_config()
gen_config.language = "<|pt|>"
gen_config.task = "transcribe"
result = pipe.generate("audio.wav", gen_config)
print(result.texts[0])Measured performance
Hardware: Intel Core Ultra 7 265H. Test audio: 8.35 s.
Limitations
Testing was done with synthetic (TTS-generated) speech, not recorded human speech. Transcription quality depends entirely on the original freds0/distil-whisper-large-v3-ptbr model — this repository changes only the runtime format, not the weights' training or quality.
Português
Conversão para OpenVINO IR (FP32) do modelo freds0/distil-whisper-large-v3-ptbr, um fine-tune em português (pt-BR) do distil-whisper/distil-large-v3 feito por freds0. Todo o crédito pelos pesos e pelo treinamento do modelo original é do autor original — este repositório apenas fornece uma exportação otimizada para OpenVINO, para inferência em CPU/GPU/NPU em hardware Intel.
Conversão
Convertido com optimum-intel 2.2.0, usando openvino 2026.4.0 e openvino-genai 2026.4.0.0. O comando optimum-cli export openvino quebrou por incompatibilidade de versão com transformers 5.5.4, então a conversão foi feita diretamente pela API Python do optimum-intel (ver trecho de código acima, seção em inglês).
O modelo exportado usa um único grafo de decoder mesclado (openvino_decoder_model.xml/.bin, com use_cache: true no config.json), em vez de grafos separados de decoder/decoder-with-past, além do par tokenizer/detokenizer OpenVINO (openvino_tokenizer.*, openvino_detokenizer.*) gerado para uso com openvino_genai.
Uso com openvino_genai.WhisperPipeline
Ver o trecho de código na seção em inglês acima — é o mesmo para CPU, GPU e NPU, bastando trocar o valor de device. Para NPU, defina CACHE_DIR nas propriedades do pipeline para evitar recompilar o modelo a cada execução.
Desempenho medido
Hardware: Intel Core Ultra 7 265H. Áudio de teste: 8,35 s.
Limitações
O teste foi feito com fala sintética (gerada por TTS), não com fala humana real gravada. A qualidade da transcrição depende inteiramente do modelo original freds0/distil-whisper-large-v3-ptbr — este repositório muda apenas o formato de execução, não os pesos nem a qualidade do treinamento.
