cstr/granite-speech-4.1-2b-plus-GGUF
granite-speech-4.1-2b-plus — GGUF
GGUF conversion of ibm-granite/granite-speech-4.1-2b-plus for use with CrispASR.
The PLUS variant adds two capabilities over the base 4.1-2b:
- Punctuated and capitalised transcripts by default — no special prompt required.
- Speaker labels and word-level timestamps in the model's structured output (full output parsing in CrispASR is the next step; raw text works today).
Architecturally PLUS is the base 4.1-2b plus a single change: the encoder's layer-3 hidden state is concatenated with the final layer output (config: cat_hidden_layers: [3]), producing a 2048-dim projector input instead of 1024. The Q-Former cross-attention K/V projection weights are correspondingly (1024, 2048).
Files
Cosine parity (vs PyTorch BF16 reference, JFK 11 s clip)
encoder_out is the 2048-dim concatenation of the layer-3 hidden state and the final encoder layer (the PLUS architectural delta). On the recommended and -f16enc files the encoder weights stay in F32/F16, so parity is essentially indistinguishable from the F16 reference. On the -mini file the encoder weights are Q4K — rounding error compounds across the 16-layer Conformer and shows up amplified after the concat, which is why `encoderout` cos_min drops to ~0.62 on PLUS where base-4.1 mini sits at ~0.93. End-to-end JFK transcription is still correct.
Tested with `crispasr-diff granite-4.1 <model.gguf> <ref.gguf> samples/jfk.wav`
Usage with CrispASR
# auto-download and transcribe
crispasr --backend granite-4.1-plus -m auto samples/audio.wav
# or with explicit path
crispasr --backend granite-4.1-plus \
-m granite-speech-4.1-2b-plus-f16.gguf \
samples/audio.wavEnd-to-end example on the JFK 11s clip:
$ crispasr --backend granite-4.1-plus -m auto samples/jfk.wav
And so my fellow Americans, ask not what your country can do for
you, ask what you can do for your country.(Note the punctuation + capitalisation that the base 4.1-2b only produces with an explicit --ask "transcribe with proper punctuation..." prompt.)
Architecture
Total ~2.2 B parameters. The "+" capability is encoded entirely in training — the architectural delta from base is just the layer concatenation.
Conversion
# Convert HF safetensors → GGUF F16
python models/convert-granite-speech-to-gguf.py \
--input /path/to/granite-speech-4.1-2b-plus \
--output granite-speech-4.1-2b-plus-f16.gguf
# Quantise F16 → Q4_K (encoder + projector preserved F32, LLM Q4_K)
crispasr-quantize granite-speech-4.1-2b-plus-f16.gguf \
granite-speech-4.1-2b-plus-q4_k.gguf q4_k
# Q4_K with F16 encoder/projector (smaller, no measurable parity loss)
CRISPASR_GRANITE_ENC_F16=1 \
crispasr-quantize granite-speech-4.1-2b-plus-f16.gguf \
granite-speech-4.1-2b-plus-q4_k-f16enc.gguf q4_k
# Aggressive Q4_K everywhere (encoder + projector + LLM)
CRISPASR_GRANITE_QUANT_ALL=1 \
crispasr-quantize granite-speech-4.1-2b-plus-f16.gguf \
granite-speech-4.1-2b-plus-q4_k-mini.gguf q4_kThe same converter handles base / 4.1-2b / 4.1-2b-plus from a single script — variant detection happens via config.json keys (cat_hidden_layers, encoder_hidden_size).
Licence
Apache 2.0 — same as the original ibm-granite/granite-speech-4.1-2b-plus.
Provenance and EU AI Act Art. 53 note
- Upstream model: ibm-granite/granite-speech-4.1-2b-plus — published by
ibm-granite. - Upstream licence:
apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
