phaxz10/belle-whisper-large-v3-turbo-zh_timestamped
Belle-whisper-large-v3-turbo-zh — ONNX (_timestamped)
ONNX weights for `BELLE-2/Belle-whisper-large-v3-turbo-zh`, a Mandarin fine-tune of openai/whisper-large-v3-turbo, for use with Transformers.js.
This is a `_timestamped` export: the decoder is exported with output_attentions=True so its graph emits cross_attentions.{0..3}, and generation_config.json carries alignment_heads. Without both, return_timestamps: 'word' throws "Model outputs must contain cross attentions".
Produced with the same recipe as `onnx-community/whisper-large-v3-turbo_timestamped` (transformers.js scripts/convert.py at tag 3.8.1, --output_attentions).
Files
No .onnx_data: every weight file is self-contained and under the 2 GB protobuf limit.
Usage
import { pipeline } from '@huggingface/transformers'
const asr = await pipeline(
'automatic-speech-recognition',
'phaxz10/belle-whisper-large-v3-turbo-zh_timestamped',
{ dtype: { encoder_model: 'fp16', decoder_model_merged: 'q4' }, device: 'webgpu' },
)
const out = await asr(audio, { language: 'chinese', task: 'transcribe', return_timestamps: true })Known limitation: word timestamps on short audio
return_timestamps: true (segment timings, read off Whisper's <|t|> tokens) is accurate at every clip length tested, from 0.5 s up.
return_timestamps: 'word' (cross-attention DTW) is accurate only from roughly 8 s of audio upward. Below that the DTW path latches onto the padded region of the 30 s mel window and returns times far outside the clip (e.g. [16.42, 24.14] for a 4 s clip, [29.82, 29.92] for a 1.6 s one). The stock onnx-community/whisper-large-v3-turbo_timestamped export is correct at all lengths on the same clips, so this comes from the Mandarin fine-tune, not from the conversion.
alignment_heads here are the stock whisper-large-v3-turbo heads, which the base model already ships and which scripts/extra/whisper.py also resolves to. Substituting other head sets (all of layer 3, all of layers 2+3, all 80 heads) was tested and none fixed the short-clip case; several were worse. Prefer return_timestamps: true unless your chunks are reliably longer than ~8 s.
