hiraki/parakeet-turntaking-stage1-libri-top4-ctc-spk-ds1
07
Stage1: LibriSpeech 960h (Top-4 Encoder + CTC + Speaker Kernel)
Speech-LLM checkpoint for ASR alignment, trained on LibriSpeech 960h with CTC auxiliary loss and SortFormer-based speaker kernel.
Architecture
- Freeze strategy:
projector_and_encoder_top— projector fully trainable, top-4 encoder layers unfrozen - CTC auxiliary loss: λ=0.3
- Speaker kernel: SortFormer-based speaker activity injected into encoder
Training Data
Training Configuration
- Optimizer: AdamW (lr=5e-5, weight_decay=0.01)
- Scheduler: warmup (2000 steps) + cosine decay (min_lr=1e-6)
- Effective batch size: 512 (8 GPUs × batchsize=16 × gradaccum=4)
- Total steps: 100,000 (trained to ~40,000)
- DeepSpeed ZeRO-2, AMP bfloat16
- Gradient clipping: 5.0
- Dynamic batching: maxframesin_batch=400,000
Usage
import torch
checkpoint = torch.load("final.pt", map_location="cpu")
# Load into SpeechLLM model — see the parakeet-turntaking repository for detailsFiles
final.pt— Full model checkpoint (encoder + projector + LLM state dicts)config.yaml— Training configuration
