smdesai/canary-1b-v2-int8-coreml
Canary 1B v2 — CoreML (INT8 encoder, KV-cache)
Memory-optimized CoreML build of NVIDIA canary-1b-v2 for Apple silicon: the encoder's weights are quantized to INT8 per-channel (linear symmetric) and decompressed in flight on the Apple Neural Engine, so the encoder's resident memory halves with the same FP16 activations and Neural Engine placement as the FP16 reference build (canary-1b-v2-coreml). Decoder and cross-KV stay FP16. Accuracy is within noise of FP16.
Base model: nvidia/canary-1b-v2 (NVIDIA NeMo EncDecMultiTaskModel, 1B parameters, 25 European languages, ASR + speech translation). License: the base model is released under CC-BY-4.0; this conversion carries the same license. Please attribute NVIDIA for the model.
This build
Build family
All 1B v2 builds share the same preprocessor, tokenizer, package layout, and decode contract; only the weight format of the encoder and/or decoder differs. Download sizes are as hosted on the Hub; iOS RAM was measured in an app while transcribing a 6-minute file.
Files
Credits
Model: NVIDIA NeMo team, canary-1b-v2, CC-BY-4.0.
