CoolFace
Modelpublic

cstr/gigaam-v3-GGUF

sourceHugging Facemitupdated 2mo agoView on Hugging Face
3likes1.2kdownloads
Model Card

GigaAM-v3 — GGUF (ggml conversions)

GGUF conversions of `ai-sage/GigaAM-v3` for use with the gigaam backend in [CrispStrobe/CrispASR](https://github.com/CrispStrobe/CrispASR).

GigaAM-v3 is a 220 M-parameter Conformer foundation model for Russian ASR, pretrained with a HuBERT-CTC objective on ~700 K hours of Russian speech. The upstream repo ships five checkpoints as git revisions; the four ASR ones are converted here (the ssl encoder has no head and produces no transcript).

FileSizeHeadVocabularyOutput
gigaam-v3-e2e-rnnt-{f16,q8_0,q4_k}.gguf452 / 249 / 154 MBRNN-TSentencePiece 1024punctuation + casing + ITN — best WER (8.4 % avg)
gigaam-v3-e2e-ctc-{f16,q8_0,q4_k}.gguf449 / 247 / 152 MBCTCSentencePiece 256punctuation + casing + ITN, faster decode
gigaam-v3-rnnt-{f16,q8_0,q4_k}.gguf449 / 246 / 152 MBRNN-T33 Cyrillic charslowercase, no punctuation
gigaam-v3-ctc-{f16,q8_0,q4_k}.gguf449 / 246 / 151 MBCTC33 Cyrillic charslowercase, no punctuation

Which one to pick

`gigaam-v3-e2e-rnnt-q8_0.gguf` unless you have a reason not to — it is the lowest-WER variant, emits punctuation and casing, and its transcript is identical to the PyTorch reference.

Usage

bash
crispasr --backend gigaam -m gigaam-v3-e2e-rnnt-q8_0.gguf -f audio.wav
# or let the registry fetch it:
crispasr --backend gigaam -m auto --auto-download -f audio.wav

Audio is 16 kHz mono. Long inputs are sliced by the CLI's VAD/chunking; the model itself has a ~25 s practical window (full attention, O(T²)).

Verification

Every file was checked against a per-stage PyTorch reference dumped from the upstream modeling_gigaam.py (crispasr-diff gigaam <model> <ref> <wav>), on GigaAM's own example.wav:

variantmelencoder (cos)transcript vs PyTorch
f16 (all four)1.0000001.000000byte-identical
q8_0 (all four)1.0000000.9974 – 0.9988byte-identical
q4_k ctc, rnnt1.0000000.95 – 0.99byte-identical
q4k `e2ectc`1.0000000.982one spurious trailing ,
q4k `e2ernnt`1.0000000.987content identical; 4 words lose their capital letter

So: q8_0 is the safe quant; q4_k is fine for the charwise models and costs a little casing/punctuation fidelity on the two SentencePiece ones.

In every quant the mel filterbank, Hann window, encoder.pre.* subsampling convs and the decode head (joint.* / decoder.* / head.ctc.*) are kept at source precision — the mel is un-normalized log-mel, so subsampling rounding error would otherwise cascade through all 16 conformer blocks, and the head is a blank-vs-token argmax where a flipped decision derails the greedy decode.

Conversion

bash
python models/convert-gigaam-to-gguf.py \
    --model ai-sage/GigaAM-v3 --revision e2e_rnnt \
    --output gigaam-v3-e2e-rnnt-f16.gguf
./build/bin/crispasr-quantize gigaam-v3-e2e-rnnt-f16.gguf \
    gigaam-v3-e2e-rnnt-q8_0.gguf q8_0

License

MIT, inherited from `ai-sage/GigaAM-v3`. Please cite the upstream model when you use these weights.

Provenance and EU AI Act Art. 53 note

  • —Upstream model: ai-sage/GigaAM-v3 — published by ai-sage.
  • —Upstream licence: mit. This repository redistributes under the same terms; it grants no rights the upstream licence does not.
  • —What was done here: format conversion and/or quantisation only (GGUF/GGML). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
  • —Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
  • —Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.