CoolFace
Modelpublic

hiraki/parakeet-turntaking-stage1-libri-top4-ctc-spk-ds1

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes7downloads
Model Card

Stage1: LibriSpeech 960h (Top-4 Encoder + CTC + Speaker Kernel)

Speech-LLM checkpoint for ASR alignment, trained on LibriSpeech 960h with CTC auxiliary loss and SortFormer-based speaker kernel.

Architecture

ComponentModelParameters
Encodernvidia/parakeet-ctc-0.6b (FastConformer CTC)600M
ProjectorConv1d + Transformer (1024 → 896, stride=2)trainable
LLMQwen/Qwen2.5-0.5B500M
  • —Freeze strategy: projector_and_encoder_top — projector fully trainable, top-4 encoder layers unfrozen
  • —CTC auxiliary loss: λ=0.3
  • —Speaker kernel: SortFormer-based speaker activity injected into encoder

Training Data

DatasetSamplesDescription
LibriSpeech 960h~276kRead English speech (standard ASR benchmark)

Training Configuration

  • —Optimizer: AdamW (lr=5e-5, weight_decay=0.01)
  • —Scheduler: warmup (2000 steps) + cosine decay (min_lr=1e-6)
  • —Effective batch size: 512 (8 GPUs × batchsize=16 × gradaccum=4)
  • —Total steps: 100,000 (trained to ~40,000)
  • —DeepSpeed ZeRO-2, AMP bfloat16
  • —Gradient clipping: 5.0
  • —Dynamic batching: maxframesin_batch=400,000

Usage

python
import torch

checkpoint = torch.load("final.pt", map_location="cpu")
# Load into SpeechLLM model — see the parakeet-turntaking repository for details

Files

  • —final.pt — Full model checkpoint (encoder + projector + LLM state dicts)
  • —config.yaml — Training configuration