poolside/Laguna-S-2.1-DFlash-FP8
<p align="center"> <img alt="poolside-banner" src="https://poolside.ai/assets/laguna/laguna-s-2-1-banner.svg" width="800px"> </p>
<p align="center"> <a href="https://openrouter.ai/poolside/laguna-s-2.1"><strong>Use on OpenRouter</strong></a> · <a href="https://vercel.com/ai-gateway/models/laguna-s-2.1"><strong>Use on Vercel AI Gateway</strong></a> · <a href="https://poolside.ai/blog/introducing-laguna-s-2-1"><strong>Release blog post</strong></a> </p>
<br>
poolside/Laguna-S-2.1-DFlash-FP8
DFlash speculator for the FP8 target poolside/Laguna-S-2.1-FP8. The speculator is a 6-layer Laguna-style draft model (BF16); pair it with the FP8 base for lower-latency serving via speculative decoding.
Trained: e0630_rhiemann_baseline SFT, DFlash Stage-2, 15k steps. Recommended serving setting: num_speculative_tokens=7. DFlash upstream support is in progress (vLLM #46853, SGLang #29446, TRT-LLM #15666). Use poolside/Laguna-S-2.1-FP8 as the target model.
Benchmarks
Measured with TP=2, temperature=0, and num_speculative_tokens=15.
