cstr/mini-omni2-GGUF
0902
Mini-Omni2 GGUF
GGUF conversion of gpt-omni/mini-omni2 for use with CrispASR.
Architecture: Whisper-small encoder (80 mel, 12L, 768d) + whisperMLP adapter (SwiGLU 768→4864→896) + Qwen2-0.5B LLM (896d, 24L, GQA 14/2).
Supports ASR (audio→text), TTS (text→audio), and speech-to-speech (audio→audio). TTS/S2S require the SNAC 24kHz codec companion (cstr/snac-24khz-GGUF).
Files
Usage
# ASR
crispasr -m mini-omni2-q4_k.gguf -f audio.wav --backend mini-omni2
# TTS (needs SNAC codec)
crispasr -m mini-omni2-q4_k.gguf --tts "Hello world" \
--codec-model snac-24khz.gguf --tts-output out.wav --backend mini-omni2Provenance and EU AI Act Art. 53 note
- Upstream model: gpt-omni/mini-omni2 — published by
gpt-omni. - Upstream licence:
mit. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
