CoolFace
Modelpublic

poolside/Laguna-S-2.1-DFlash-FP8

sourceHugging Faceupdated 2mo agoView on Hugging Face
8likes1.9kdownloads
Model Card

<p align="center"> <img alt="poolside-banner" src="https://poolside.ai/assets/laguna/laguna-s-2-1-banner.svg" width="800px"> </p>

<p align="center"> <a href="https://openrouter.ai/poolside/laguna-s-2.1"><strong>Use on OpenRouter</strong></a> · <a href="https://vercel.com/ai-gateway/models/laguna-s-2.1"><strong>Use on Vercel AI Gateway</strong></a> · <a href="https://poolside.ai/blog/introducing-laguna-s-2-1"><strong>Release blog post</strong></a> </p>

<br>

poolside/Laguna-S-2.1-DFlash-FP8

DFlash speculator for the FP8 target poolside/Laguna-S-2.1-FP8. The speculator is a 6-layer Laguna-style draft model (BF16); pair it with the FP8 base for lower-latency serving via speculative decoding.

Trained: e0630_rhiemann_baseline SFT, DFlash Stage-2, 15k steps. Recommended serving setting: num_speculative_tokens=7. DFlash upstream support is in progress (vLLM #46853, SGLang #29446, TRT-LLM #15666). Use poolside/Laguna-S-2.1-FP8 as the target model.

Benchmarks

Measured with TP=2, temperature=0, and num_speculative_tokens=15.

Throughput speedup

ConcurrencyGSM8KMATH-500HumanEvalMBPPMT-Bench
13.179x2.938x3.269x2.380x2.603x
42.614x2.423x2.691x1.963x2.090x
82.666x2.410x2.803x1.962x2.230x
162.618x2.364x2.866x2.031x2.302x

Acceptance length

ConcurrencyGSM8KMATH-500HumanEvalMBPPMT-Bench
15.7485.1975.8894.2474.663
45.7655.2125.8824.2184.411
85.8635.1996.0944.1784.572
165.7875.1616.1444.2914.600