CoolFace
Modelpublic

phaxz10/belle-whisper-large-v3-turbo-zh_timestamped

sourceHugging Faceapache-2.0updated 18d agoView on Hugging Face
1likes42downloads
Model Card

Belle-whisper-large-v3-turbo-zh — ONNX (_timestamped)

ONNX weights for `BELLE-2/Belle-whisper-large-v3-turbo-zh`, a Mandarin fine-tune of openai/whisper-large-v3-turbo, for use with Transformers.js.

This is a `_timestamped` export: the decoder is exported with output_attentions=True so its graph emits cross_attentions.{0..3}, and generation_config.json carries alignment_heads. Without both, return_timestamps: 'word' throws "Model outputs must contain cross attentions".

Produced with the same recipe as `onnx-community/whisper-large-v3-turbo_timestamped` (transformers.js scripts/convert.py at tag 3.8.1, --output_attentions).

Files

filesize
onnx/encoder_model_fp16.onnx1274.40 MB
onnx/decoder_model_merged_q4.onnx334.05 MB
onnx/decoder_model_merged_quantized.onnx (q8)438.12 MB

No .onnx_data: every weight file is self-contained and under the 2 GB protobuf limit.

Usage

js
import { pipeline } from '@huggingface/transformers'

const asr = await pipeline(
  'automatic-speech-recognition',
  'phaxz10/belle-whisper-large-v3-turbo-zh_timestamped',
  { dtype: { encoder_model: 'fp16', decoder_model_merged: 'q4' }, device: 'webgpu' },
)

const out = await asr(audio, { language: 'chinese', task: 'transcribe', return_timestamps: true })

Known limitation: word timestamps on short audio

return_timestamps: true (segment timings, read off Whisper's <|t|> tokens) is accurate at every clip length tested, from 0.5 s up.

return_timestamps: 'word' (cross-attention DTW) is accurate only from roughly 8 s of audio upward. Below that the DTW path latches onto the padded region of the 30 s mel window and returns times far outside the clip (e.g. [16.42, 24.14] for a 4 s clip, [29.82, 29.92] for a 1.6 s one). The stock onnx-community/whisper-large-v3-turbo_timestamped export is correct at all lengths on the same clips, so this comes from the Mandarin fine-tune, not from the conversion.

alignment_heads here are the stock whisper-large-v3-turbo heads, which the base model already ships and which scripts/extra/whisper.py also resolves to. Substituting other head sets (all of layer 3, all of layers 2+3, all 80 heads) was tested and none fixed the short-clip case; several were worse. Prefer return_timestamps: true unless your chunks are reliably longer than ~8 s.

Base model quality (from the base model card)

benchmark`whisper-large-v3-turbo`Belle fine-tune
AISHELL-1 (CER)8.643.07
WenetSpeech-meeting (CER)20.3113.36