CoolFace
Datasetpublic

cogeanu-marius/qwen3.6-27b-h100-bf16-benchmark

Two GPUs do not mean twice the users This study answers a serving decision, not a hardware trivia question: when a 27B model already fits on one H100, should a second GPU shard the model or run a second independent replica? The answer Concurrency is a load-generator setting, not a user count and not a promise. Capacity is the highest real arrival rate that satisfies a declared service objective. That is why this study measures both a saturated concurrency curve… See the full description on the dataset page: https://huggingface.co/datasets/cogeanu-marius/qwen3.6-27b-h100-bf16-benchmark.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes20downloads

cogeanu-marius/qwen3.6-27b-h100-bf16-benchmark · main · files are served by the source, never re-hosted here