aufklarer/Qwen3-ForcedAligner-0.6B-4bit
14.9k
Qwen3-ForcedAligner-0.6B-4bit (MLX)
4-bit quantized version of Qwen/Qwen3-ForcedAligner-0.6B for Apple Silicon inference via MLX.
Predicts word-level timestamps for audio+text pairs in a single non-autoregressive forward pass.
Model Details
How It Works
Audio + Text → Audio Encoder → Text Decoder (single pass) → Classify Head → argmax at <timestamp> positions → word timestampsUnlike ASR (autoregressive, token-by-token), the forced aligner runs the entire sequence in one forward pass through the decoder. The classify head predicts a timestamp class (0–4999) at each <timestamp> token position, which maps to time via class_index × 80ms.
Usage with Swift (MLX)
This model is designed for use with speech-swift:
import Qwen3ASR
let aligner = try await Qwen3ForcedAligner.fromPretrained()
let aligned = aligner.align(
audio: audioSamples,
text: "Can you guarantee that the replacement part will be shipped tomorrow?",
sampleRate: 24000
)
for word in aligned {
print("[\(String(format: "%.2f", word.startTime))s - \(String(format: "%.2f", word.endTime))s] \(word.text)")
}CLI
# Align with provided text
qwen3-asr-cli --align --text "Hello world" audio.wav
# Transcribe first, then align
qwen3-asr-cli --align audio.wavOutput:
[0.12s - 0.45s] Can
[0.45s - 0.72s] you
[0.72s - 1.20s] guarantee
[1.20s - 1.48s] that
...Quantization
Text decoder (attention projections, MLP, embeddings) quantized to 4-bit using group quantization (group_size=64). Audio encoder and classify head kept as float16 for accuracy.
Converted with:
python scripts/convert_forced_aligner.py \
--source Qwen/Qwen3-ForcedAligner-0.6B \
--upload --repo-id aufklarer/Qwen3-ForcedAligner-0.6B-4bitLinks
- Swift library: soniqo/speech-swift — Swift Package for Qwen3-ASR, Qwen3-TTS, CosyVoice, PersonaPlex, and Forced Alignment on Apple Silicon
- Base model: Qwen/Qwen3-ForcedAligner-0.6B
- bf16 variant: mlx-community/Qwen3-ForcedAligner-0.6B-bf16
- Guide: soniqo.audio/guides/align
- Docs: soniqo.audio
- GitHub: soniqo/speech-swift
