soniqo/MOSS-Transcribe-Diarize-0.9B-ONNX-INT8
MOSS-Transcribe-Diarize-0.9B-ONNX-INT8
ONNX conversion of OpenMOSS-Team/MOSS-Transcribe-Diarize, pinned to revision e6d68cdfcddbdad1a7e8454f0cb859cad76e2502. The model produces timestamped, speaker-attributed text in the form [start][Sxx]text[end].
Model
Files
Performance
Measured with greedy decoding on an Apple M5 Pro with 48 GB unified memory. Word error rate (WER), character error rate (CER), and real-time factor (RTF) are lower when better. Throughput is 1 / RTF, is higher when better, and reports how many seconds of audio are processed per wall-clock second. RTF excludes model loading. RSS is process memory sampled at clip boundaries; OS high-water RSS also includes transient peaks when the operating system reports it.
Aggregate inference phases across 80 clips: Processor 0.18 s, Audio encoder 5.57 s, Decoder prefill 3.90 s, Token decode 25.12 s, Other host work 0.01 s; 5.602 ms/generated token.
Multilingual precision check
Round trips
- 30-second overlapping AMI window: RTF 0.077; text=exact, speaker_sequence=exact, raw=drift. The source predicted 3 speakers where RTTM contains 4.
- 33.12-second two-chunk German clip: RTF 0.036; text=exact, speaker_sequence=exact, raw=exact.
The multilingual check covers only ten German, ten French, and ten Modern Standard Arabic clips. It is a conversion check, not proof of the upstream model's full 50-language quality. The meeting round trip checks transcript and speaker-sequence parity but is not a diarization error-rate benchmark.
Usage
Python
from pathlib import Path
import coremltools as ct
import onnxruntime as ort
import onnxruntime_ep_webgpu as webgpu_ep
from huggingface_hub import snapshot_download
bundle = Path(snapshot_download("soniqo/MOSS-Transcribe-Diarize-0.9B-ONNX-INT8"))
ort.register_execution_provider_library("webgpu", webgpu_ep.get_library_path())
devices = [
device for device in ort.get_ep_devices()
if device.ep_name == webgpu_ep.get_ep_name()
]
options = ort.SessionOptions()
options.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_BASIC
options.add_provider_for_devices(devices, {})
audio = ct.models.CompiledMLModel(str(bundle / "audio_encoder.mlmodelc"))
decoder = ort.InferenceSession(str(bundle / "decoder.onnx"), sess_options=options)Allocate each past_key_N / past_value_N as a 1,024-row WebGPU OrtValue. Use ONNX Runtime I/O binding to bind the same OrtValue to the corresponding present_key_N / present_value_N output. Pass the real seqlens_k and total_sequence_length; capacity is not the logical token position.
Command line
hf download soniqo/MOSS-Transcribe-Diarize-0.9B-ONNX-INT8 --local-dir ./moss-transcribe-diarizeThis downloads a complete low-level model bundle. SDK integration is tracked in speech-swift issue #388; until that integration lands, applications must implement the documented host contract around the exported graphs or weights.
Runtime contract
Decoder linear weights use homogeneous symmetric INT8 block-32 MatMulNBits. Q/K/V and gate/up projections are fused into wide calls. Its per-layer FP16 K/V buffers remain on the WebGPU device and are aliased as both past inputs and present outputs; the host passes the logical cache lengths directly. Prompt construction, audio chunking, generation, cache management, and transcript parsing remain host responsibilities. Machine-readable details and measured results are in config.json, export_config.json, and validation.json.
Source
No training was performed for this conversion. The source weights contain 908,513,280 parameters and are licensed under Apache 2.0.
Limitations
Speaker IDs are anonymous within each inference. The source model can emit malformed or overlapping timestamps and can miss speakers; conversion parity does not correct those behaviors. The WebGPU execution provider is a preview plugin and this profile requires macOS plus the bundled Core ML audio encoder. The fixed decoder cache limits total prompt plus generated tokens to 1,024. The validation here does not establish the upstream 90-minute claim for this deployment format.
Links
- speech-swift — Apple SDK
- Docs — Apple setup and CLI docs
- soniqo.audio — website
- blog — blog
License
Apache License 2.0, inherited from the upstream model.
