CoolFace
Modelpublic

tiagoCuervo/gslm-scaling-20m-43p5b

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes39downloads
Model Card

gslm-scaling-20m-43p5b

Model-only FP32 weights for a final-budget scaling-curve endpoint from Scaling Properties of Speech Language Models (EMNLP 2024). This checkpoint has 20,710,912 parameters and saw exactly 43,545,280,000 tokens at 131,200 tokens per update.

python
import torch
from slm import TransformerLM

model = TransformerLM.from_pretrained("tiagoCuervo/gslm-scaling-20m-43p5b", device="cuda")
prompt = torch.tensor([[12, 91, 204]], device="cuda")
units = model.generate(prompt, max_new_tokens=200, temperature=0.8, top_k=50)

Evaluation scores are 56.20% sBLIMP phenomenon macro accuracy, 54.00% sBLIMP voice-pair accuracy, 52.65% sStoryCloze, and 70.66% tStoryCloze.

The model predicts consecutive-run-collapsed 25 Hz layer-11 mHuBERT K-means units (K=500; EOS=500; PAD=501). A vocoder is not bundled.

Speech-tokenizer and vocoder artifacts are credited in the SLM third-party notices. See provenance.json for exact hashes and evaluation metadata.