vadimsuhanov/gigaam-multilingual-large-ctc-coreml
GigaAM Multilingual Large CTC for Core ML
Model description
This repository contains a native Core ML conversion of the multilingual GigaAM Large CTC speech-recognition model. It is intended for offline Russian and English transcription in the MeetingTranscriber macOS application.
Provenance
- Upstream repository:
ai-sage/GigaAM-Multilingual - Immutable upstream revision:
3905cd51c3ed4e88c8edf33f3302969ba480a327 - Upstream variant:
large_ctc - Conversion precision: FP32
The FP16 candidate was not published because it failed strict PyTorch/Core ML numeric and token-frame parity gates. The validated artifact uses FP32 weights.
Requirements
- Apple Silicon Mac
- macOS 26 or later
- Core ML ML Program runtime
- Mono 16 kHz PCM audio
Core ML contract
- Maximum input window: 25 seconds
features: float32[1, 64, 2499]log-mel featuresfeature_length: int32[1]logits: float32[1, 625, 71]encoded_length: int32[1]- Vocabulary tokens: 70
- CTC classes: 71
- Blank ID: 70
- Greedy CTC decoding is performed by the application
Languages and Auto detection
The application exposes Russian (ru), English (en), and Auto/Bilingual. Fixed language selection controls result metadata and does not filter characters from the shared vocabulary. Auto examines recognized Unicode letters: at least 80% Cyrillic selects Russian, at least 80% Latin selects English, and mixed or empty text remains Auto.
Validation
The FP32 Core ML output was compared with the pinned PyTorch source on private Russian, English, and mixed-alphabet fixtures. The published validation report requires finite outputs, exact greedy token sequences, exact emitted token frames, bounded logit error, bounded timestamp error, and no WER regression. Fixture audio, transcripts, paths, caches, and credentials are not included.
Limitations
- The model accepts one batch item and at most 25 seconds per inference.
- Audio resampling, VAD chunking, decoding, language metadata, and word timing construction are application responsibilities.
- Auto detection is alphabet-based; it is not a separate acoustic language model.
- Validation of mixed-alphabet input does not add a third fixed language mode.
License
The upstream model and this conversion are distributed under the MIT license. See LICENSE for the complete license text.
Files
model.mlpackage: FP32 Core ML ML Programmodel-config.json: fixed feature and encoder contractvocabulary.json: ordered 70-token vocabularyconversion-manifest.json: provenance, tensor declarations, and SHA-256 inventoryvalidation-report.json: public, path-free parity metricsLICENSE: MIT license text
