cstr/granite-speech-4.1-2b-GGUF
granite-speech-4.1-2b — GGUF
GGUF conversions of ibm-granite/granite-speech-4.1-2b for use with CrispASR.
Files
Cosine parity (vs PyTorch BF16 reference, JFK 11 s clip)
The encoder is a 16-layer Conformer where Q4K rounding error compounds across layers; the recommended Q4K file pins those weights at F32 to preserve numerical fidelity. The -f16enc file relaxes that to F16 and ships ~1 GB smaller while keeping cosine essentially indistinguishable from F16 (every Whisper / Llama / parakeet GGUF in the wild already runs F16 weights). The -mini file applies Q4_K to every quantisable 2D weight including the encoder — useful when disk or download size matters more than transcript quality.
Tested with `crispasr-diff granite-4.1 <model.gguf> <ref.gguf> samples/jfk.wav`
Architecture
Granite Speech 4.1 2B is a speech-LLM with three components:
- Encoder: 16-layer Macaron Conformer (hidden 1024, 8 heads, 15-tap depthwise conv, dual CTC heads for characters + BPE). Input: 80-bin log-mel × 2-frame stacked = 160-dim, 10 ms hop.
- Projector: 2-layer BLIP-2 Q-Former with 3 learned queries per 15-frame window (5× temporal downsampling). Combined with encoder's 2× → 10 Hz acoustic token rate for the LLM.
- LLM: Granite 4.0-1B (40 layers, 2048 hidden, GQA 16/4, SwiGLU, RoPE θ=10000, μP multipliers).
Total ~2.2 B parameters. Named "2B" to reflect the full system size rather than the base LLM alone.
Usage with CrispASR
# auto-download and transcribe
crispasr --backend granite-4.1 -m auto samples/audio.wav
# or with explicit path
crispasr --backend granite-4.1 \
-m granite-speech-4.1-2b-q4_k.gguf \
samples/audio.wavSupported tasks via prompt (-p flag):
Supported languages: English, French, German, Spanish, Portuguese, Japanese.
Conversion
# Convert HF safetensors → GGUF F16
python models/convert-granite-speech-to-gguf.py \
--input /path/to/granite-speech-4.1-2b \
--output granite-speech-4.1-2b-f16.gguf
# Quantise F16 → Q4_K
crispasr-quantize granite-speech-4.1-2b-f16.gguf \
granite-speech-4.1-2b-q4_k.gguf q4_kThe converter handles all three Granite Speech 4.x releases (4.0-1b, 4.1-2b) from the same script; parameters are read from config.json at conversion time.
Licence
Apache 2.0 — same as the original ibm-granite/granite-speech-4.1-2b.
Provenance and EU AI Act Art. 53 note
- Upstream model: ibm-granite/granite-speech-4.1-2b — published by
ibm-granite. - Upstream licence:
apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
