Relativ3pa1n/dsv4-flash-sm86-8x3090
DeepSeek-V4-Flash on 8x RTX 3090 (SM86): 262K context, 120 tok/s aggregate Serving recipes, launch wrappers, and the measured throughput/context ladder for running a W4A16 DeepSeek-V4-Flash-class model on 8x RTX 3090 (SM 8.6, 24 GiB each) with CUDA graph decode, FlashInfer sparse MLA, Marlin MoE, and compressed hybrid KV. The ladder Concurrent sequences amortize the TP8 allreduce that dominates each decode step, so aggregate throughput scales near-linear while… See the full description on the dataset page: https://huggingface.co/datasets/Relativ3pa1n/dsv4-flash-sm86-8x3090.
035
