FredrikKarlssonSpeech/wav2vec2-large-voxrex-swedish-onnx
026
wav2vec2-large-voxrex-swedish โ ONNX
ONNX export of KBLab/wav2vec2-large-voxrex-swedish, a 317M-parameter Wav2Vec2ForCTC model fine-tuned for Swedish automatic speech recognition using the VoxRex corpus.
Exported with ๐ค Optimum (v2.1.0), opset 17. FP16/INT8/UINT8/Q4 quantized variants included.
Available Variants
FP16 is the recommended quantized variant. Dynamic INT8/UINT8/Q4 show logit drift without Swedish calibration data; FP16 produces text-identical output to FP32.
Usage
With ๐ค Optimum (recommended)
from optimum.onnxruntime import ORTModelForCTC
from transformers import Wav2Vec2Processor
import soundfile as sf
import torch
model_id = "FredrikKarlssonSpeech/wav2vec2-large-voxrex-swedish-onnx"
processor = Wav2Vec2Processor.from_pretrained(model_id)
model = ORTModelForCTC.from_pretrained(model_id, subfolder="onnx", file_name="model.onnx")
audio, sr = sf.read("your_audio.wav")
assert sr == 16000, "Resample to 16kHz first"
inputs = processor(audio, sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
logits = model(**inputs).logits
predicted_ids = torch.argmax(logits, dim=-1)
transcription = processor.batch_decode(predicted_ids)[0]
print(transcription)With ONNX Runtime directly
import onnxruntime as ort
import numpy as np
import soundfile as sf
import torch
from transformers import Wav2Vec2Processor
from huggingface_hub import hf_hub_download
model_id = "FredrikKarlssonSpeech/wav2vec2-large-voxrex-swedish-onnx"
processor = Wav2Vec2Processor.from_pretrained(model_id)
# Download and load ONNX model (use model_fp16.onnx for half the size)
model_path = hf_hub_download(model_id, filename="onnx/model.onnx")
sess = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
audio, sr = sf.read("your_audio.wav")
inputs = processor(audio, sampling_rate=16000, return_tensors="np")
logits = sess.run(["logits"], {"input_values": inputs["input_values"]})[0]
predicted_ids = np.argmax(logits, axis=-1)
transcription = processor.batch_decode(torch.tensor(predicted_ids))[0]
print(transcription)Using FP16 variant (half the size, same accuracy)
from optimum.onnxruntime import ORTModelForCTC
from transformers import Wav2Vec2Processor
model_id = "FredrikKarlssonSpeech/wav2vec2-large-voxrex-swedish-onnx"
processor = Wav2Vec2Processor.from_pretrained(model_id)
# FP16: 632 MB, text-identical to FP32
model = ORTModelForCTC.from_pretrained(model_id, subfolder="onnx", file_name="model_fp16.onnx")Resampling audio to 16kHz
import torchaudio
waveform, sample_rate = torchaudio.load("your_audio.wav")
if sample_rate != 16000:
waveform = torchaudio.functional.resample(waveform, sample_rate, 16000)
audio = waveform.squeeze().numpy()Performance
Benchmarked on M1 Mac (Apple Silicon), 10 seconds of 16kHz audio, 5 runs after 2 warmup:
On Apple Silicon, PyTorch uses Apple's Accelerate framework (vecLib BLAS) which outperforms ONNX Runtime's generic CPU path. On x86/Linux, ONNX Runtime with INT8 typically achieves 2โ4x speedup over PyTorch.
Accuracy
FP32 ONNX output is numerically identical to PyTorch (max logit diff: 0.002, mean: 0.00006, text match: exact).
For better INT8 accuracy, static quantization with Swedish calibration data is recommended.
Export Details
- Optimum version: 2.1.0
- ONNX opset: 17 (native GroupNorm op)
- ORT version: 1.24.4
- Graph: 1838 nodes, single model (no encoder/decoder split)
- Input:
input_valuesโ float32, shape[batch, sequence], 16kHz raw waveform - Output:
logitsโ float32, shape[batch, sequence/320, vocab_size] - Vocab size: 46 (Swedish character set + CTC blank)
- CTC decoding: not included in graph โ greedy argmax or beam search applied in Python
- Note: FP16 model uses
keep_io_types=True; I/O remains float32, internal weights are float16
Original Model
See KBLab/wav2vec2-large-voxrex-swedish for training details, dataset, and full evaluation results.
Citation
If you use this model, please cite the original:
@misc{wav2vec2-large-voxrex-swedish,
author = {KBLab},
title = {wav2vec2-large-voxrex-swedish},
year = {2022},
publisher = {HuggingFace},
url = {https://huggingface.co/KBLab/wav2vec2-large-voxrex-swedish}
}License
Apache 2.0, same as the original model.
