CoolFace
Datasetpublic

Relativ3pa1n/dsv4-flash-sm86-8x3090

DeepSeek-V4-Flash on 8x RTX 3090 (SM86): 262K context, 120 tok/s aggregate Serving recipes, launch wrappers, and the measured throughput/context ladder for running a W4A16 DeepSeek-V4-Flash-class model on 8x RTX 3090 (SM 8.6, 24 GiB each) with CUDA graph decode, FlashInfer sparse MLA, Marlin MoE, and compressed hybrid KV. The ladder Concurrent sequences amortize the TP8 allreduce that dominates each decode step, so aggregate throughput scales near-linear while… See the full description on the dataset page: https://huggingface.co/datasets/Relativ3pa1n/dsv4-flash-sm86-8x3090.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes35downloads
5 commits on main
bd0c98a1mo ago

Add exploitbench doc: ladder results for the served checkpoint

Relativ3pa1n
094dd581mo ago

Fix punctuation from em-dash sweep

Relativ3pa1n
27109801mo ago

Reframe terminal-bench doc: coherence check, not leaderboard claim

Relativ3pa1n
62b29f81mo ago

Add terminal-bench head-to-head vs Laguna-S-2.1 (18 tasks, equal clocks)

Relativ3pa1n
b07ebdf1mo ago

DSV4-Flash SM86 8x3090 serving: ladder, runbook, roadmap, model prep, wrappers

Relativ3pa1n