tiagoCuervo/gslm-scaling-309m-3p0b
040
gslm-scaling-309m-3p0b
Model-only FP32 weights for a final-budget scaling-curve endpoint from Scaling Properties of Speech Language Models (EMNLP 2024). This checkpoint has 308,845,568 parameters and saw exactly 3,022,110,000 tokens at 516,600 tokens per update.
import torch
from slm import TransformerLM
model = TransformerLM.from_pretrained("tiagoCuervo/gslm-scaling-309m-3p0b", device="cuda")
prompt = torch.tensor([[12, 91, 204]], device="cuda")
units = model.generate(prompt, max_new_tokens=200, temperature=0.8, top_k=50)Evaluation scores are 57.58% sBLIMP phenomenon macro accuracy, 54.69% sBLIMP voice-pair accuracy, 54.20% sStoryCloze, and 71.19% tStoryCloze.
The model predicts consecutive-run-collapsed 25 Hz layer-11 mHuBERT K-means units (K=500; EOS=500; PAD=501). A vocoder is not bundled.
Speech-tokenizer and vocoder artifacts are credited in the SLM third-party notices. See provenance.json for exact hashes and evaluation metadata.
